EDBT 2026 Demo / reviewers in the wild / expert
Stefan Scherer
dblp:60/5336
· DBLP profile ↗
94ranked-venue papers
18as first author
5since 2021 · last 2024
0000-0002-0280-5393ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 12 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 36 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Probabilistic and Bayesian machine learning · 27% Trustworthy machine learning · 24% Optimization for machine learning · 14% | |
| Human-computer interaction and pervasive computing
4 papers |
Human-AI interaction · 38% Health and well-being technologies · 16% Haptics and multimodal interaction · 16% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 97% Image and video processing · 3% |
Topics — the 19 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.8 | 1 | 2024 | Accelerating Look-ahead in Bayesian Optimization: Multilevel Monte Carlo is All you Need · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning
monte carlo methods |
0.8 | 1 | 2024 | Accelerating Look-ahead in Bayesian Optimization: Multilevel Monte Carlo is All you Need · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
multilevel monte carlo |
0.8 | 1 | 2024 | Accelerating Look-ahead in Bayesian Optimization: Multilevel Monte Carlo is All you Need · ICML 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
0.7 | 1 | 2023 | Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023 |
Machine learning › Trustworthy machine learning
uncertainty and abstention |
0.7 | 1 | 2023 | Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023 |
Computer vision › Vision and language
visual question answering |
0.7 | 1 | 2023 | Improving Selective Visual Question Answering by Learning from Your Peers · CVPR 2023 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.3 | 1 | 2018 | Modeling Temporality of Human Intentions by Domain Adaptation · EMNLP 2018 |
Natural language and speech › Question answering and dialogue systems
intention detection |
0.3 | 1 | 2018 | Modeling Temporality of Human Intentions by Domain Adaptation · EMNLP 2018 |
Natural language and speech › Language models and text generation
neural language model |
0.3 | 1 | 2017 | Affect-LM: A Neural Language Model for Customizable Affective Text Generation · ACL (1) 2017 |
Multimedia analysis and retrieval › multimedia analysis
multimodal behavior analysis |
0.2 | 1 | 2015 | A Multimodal Predictive Model of Successful Debaters or How I Learned to Sway Votes · ACM Multimedia 2015 |
Haptics and multimodal interaction
multimodal perception |
0.2 | 1 | 2015 | SimSensei Demonstration: A Perceptive Virtual Human Interviewer for Healthcare Applications · AAAI 2015 |
Learning and educational technologies › skill training
public speaking training |
0.2 | 1 | 2015 | Exploring feedback strategies to improve public speaking: an interactive virtual audience framework · UbiComp 2015 |
Human-AI interaction › virtual agents › virtual human
virtual audience |
0.2 | 1 | 2015 | Exploring feedback strategies to improve public speaking: an interactive virtual audience framework · UbiComp 2015 |
Human-AI interaction
affective computing |
0.1 | 1 | 2017 | Affect-LM: A Neural Language Model for Customizable Affective Text Generation · ACL (1) 2017 |
Computer vision › 3D vision › local feature descriptor
ordinal measures |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Computer vision › 3D vision › stereo vision
shape from stereo |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Computer vision › 3D vision
stereo vision |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Image and video processing
image matching |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Methods — techniques the papers use, named apart from their topics
multilevel monte carlo · 0.8bayesian optimization · 0.8multimodal selection function · 0.7learning from peers · 0.7neural language model · 0.6affect conditioning · 0.6multimodal machine learning · 0.4fusion · 0.4topic model · 0.3domain adaptation · 0.3BiLSTM · 0.3self-assessment questionnaire · 0.2multimodal sensing · 0.2expert assessment · 0.2behavioral analysis · 0.2ordinal correlation coefficient · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Accelerating Look-ahead in Bayesian Optimization: Multilevel Monte Carlo is All you NeedabstractWe leverage multilevel Monte Carlo (MLMC) to improve the performance of multi-step look- ahead Bayesian optimization (BO) methods that involve nested expectations and maximizations. Often these expectations must be computed by Monte Carlo (MC). The complexity rate of naive MC degrades for nested operations, whereas MLMC is capable of achieving the canonical MC convergence rate for this type of problem, independently of dimension and without any smoothness assumptions. Our theoretical study focuses on the approximation improvements for two- and three-step look-ahead acquisition functions, but, as we discuss, the approach is generalizable in various ways, including beyond the context of BO. Our findings are verified numerically and the benefits of MLMC for BO are illustrated on several benchmark examples. Code is available at https://github.com/Shangda-Yang/MLMCBO. Shangda Yang, Vitaly Zankin, Maximilian Balandat, Stefan Scherer, Kevin Carlberg, Neil S. Walton, Kody J. H. Law |
ICML | 4 |
| 2023 | Therapist Empathy Assessment in Motivational InterviewsabstractThe quality and effectiveness of psychotherapy sessions are highly influenced by the therapists’ ability to lead the conversation with empathy and acceptance. Manual assessment of the quality of therapy sessions is labor-intensive and difficult to scale. In this paper, we propose a method for estimating session-level therapist empathy ratings for Motivational Interviewing (MI) using therapist language, which has applications in clinical assessment and training. We analyze different stages within therapy sessions to investigate the importance of each stage and its topics of conversation in estimating session-level therapist empathy. We perform experiments on two datasets of MI therapy sessions for alcohol use disorder with session-level empathy scores provided by expert annotators. We achieve average CCC (Concordance Correlation Coefficient) scores of 0.596 and 0.408 for estimating therapist empathy under therapist-dependent and therapist-independent evaluation settings. Our results suggest that therapist responses to client’s discussions on activities and experiences around the problematic behavior (in this case, alcohol abuse) along with the therapist’s usage of in-depth reflections, are the most significant factors in the perception of therapists’ empathy. Leili Tavabi, Trang Tran 0001, Brian Borsari, Joannalyn Delacruz, Joshua Woolley, Stefan Scherer, Mohammad Soleymani 0001 |
ACII | 6 |
| 2023 | Improving Selective Visual Question Answering by Learning from Your PeersabstractDespite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains under-explored. Recent work has shown that VQA models, out-of-the-box, can have difficulties abstaining from answering when they are wrong. The option to abstain, also called Selective Prediction, is highly relevant when deploying systems to users who must trust the system's output (e.g., VQA assistants for users with visual impairments). For such scenarios, abstention can be especially important as users may provide out-of-distribution (OOD) or adversarial inputs that make incorrect answers more likely. In this work, we explore Selective VQA in both in-distribution (ID) and OOD scenarios, where models are presented with mixtures of ID and OOD data. The goal is to maximize the number of questions answered while minimizing the risk of error on those questions. We propose a simple yet effective Learning from Your Peers (LYP) approach for training multimodal selection functions for making abstention decisions. Our approach uses predictions from models trained on distinct subsets of the training data as targets for optimizing a Selective VQA model. It does not require additional manual labels or held-out data and provides a signal for identifying examples that are easy/difficult to generalize to. In our extensive evaluations, we show this benefits a number of models across different architectures and scales. Overall, for ID, we reach 32.92% in the selective prediction metric coverage at 1 % risk of error$(\mathcal{C} {@} 1\%)$which doubles the previous best coverage of 15.79% on this task. For mixed ID/OOD, using models' softmax confidences for abstention decisions performs very poorly, answering$\mathcal{C}$@1%. Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, Marcus Rohrbach |
CVPR | 5 |
| 2022 | Speech Behavioral Markers Align on Symptom Factors in Psychological DistressabstractAutomatic detection of psychological disorders has gained significant attention in recent years due to the rise in their prevalence. However, the majority of studies have overlooked the complexity of disorders in favor of a “present/not present” dichotomy in representing disorders. Recent psychological research challenges favors transdiagnostic approaches, moving beyond general disorder classifications to symptom level analysis, as symptoms are often not exclusive to individual disorder classes. In our study, we investigated the link between speech signals and psychological distress symptoms in a corpus of 333 screening interviews from the Distress Analysis Interview Corpus (DAIC). Given the semi-structured organization of interviews, we aggregated speech utterances from responses to shared questions across interviews. We employed deterministic sample selection in classification to rank salient questions for eliciting symptom-specific behaviors in order to predict symptom presence. Some questions include “Do you find therapy helpful?” and “When was the last time you felt happy?”. The prediction results align closely to the factor structure of psychological distress symptoms, linking speech behaviors primarily to somatic and affective alterations in both depression and PTSD. This lends support for the transdiagnostic validity of speech markers for detecting such symptoms. Surprisingly, we did not find a strong link between speech markers and cognitive or psychomotor alterations. This is surprising, given the complexity of motor and cognitive actions required in speech production. The results of our analysis highlight the importance of aligning affective computing research with psychological research to investigate the use of automatic behavioral sensing to assess psychiatric risk. Larry Zhang, Jacek Kolacz, Albert A. Rizzo, Stefan Scherer, Mohammad Soleymani 0001 |
ACII | 4 |
| 2022 | Development and Cross-Cultural Evaluation of a Scoring Algorithm for the Biometric Attachment Test: Overcoming the Challenges of Multimodal Fusion with "Small Data"abstractThe Biometric Attachment Test (BAT) is a recently developed psychometric assessment that exposes adults to standardized picture and music stimuli-sets while simultaneously capturing their linguistic, behavioral and physiological responses, with the goal of objectively measuring their psychological attachment characteristics. Within this work,(I)we describe a new version of the BAT (v2) that implements a remote photoplethysmography method to obtain physiological measures from video alone.(II)We discuss the specific challenges we found when trying to develop an automatic scoring algorithm for the BAT v2 using machine learning: practicing multimodal fusion over a high-dimensional feature space with a small learning sample. We propose and evaluate an original combination of methods, including a three-step hybrid multimodal fusion procedure, that overcomes these challenges.(III)Using the proposed methodology, we train a scoring algorithm for the BAT v2 on a francophone sample, using the Adult Attachment Questionnaire as ground-truth.(IV)We then validate the scoring algorithm cross-culturally, testing its performance on an independent anglophone sample, showing low error and high correlation and serving as the BAT v2's first convergent validity evidence. We believe this work constitutes a breakthrough in the development of the first objective and automatic measure for adult attachment, and we hope that our “small data” learning methodology could be useful for other machine learning projects involving small samples coming from psychological research. Federico Parra, Stefan Scherer, Yannick Benezeth, Plamena Tsvetanova, Susana Tereno |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Multimodal Automatic Coding of Client Behavior in Motivational InterviewingabstractMotivational Interviewing (MI) is defined as a collaborative conversation style that evokes the client's own intrinsic reasons for behavioral change. In MI research, the clients' attitude (willingness or resistance) toward change as expressed through language, has been identified as an important indicator of their subsequent behavior change. Automated coding of these indicators provides systematic and efficient means for the analysis and assessment of MI therapy sessions. In this paper, we study and analyze behavioral cues in client language and speech that bear indications of the client's behavior toward change during a therapy session, using a database of dyadic motivational interviews between therapists and clients with alcohol-related problems. Deep language and voice encoders, \ie BERT and VGGish, trained on large amounts of data are used to extract features from each utterance. We develop a neural network to automatically detect the MI codes using both the clients' and therapists' language and clients' voice, and demonstrate the importance of semantic context in such detection. Additionally, we develop machine learning models for predicting alcohol-use behavioral outcomes of clients through language and voice analysis. Our analysis demonstrates that we are able to estimate MI codes using clients' textual utterances along with preceding textual context from both the therapist and client, reaching an F1-score of 0.72 for a speaker-independent three-class classification. We also report initial results for using the clients' data for predicting behavioral outcomes, which outlines the direction for future work. Leili Tavabi, Kalin Stefanov, Larry Zhang, Brian Borsari, Joshua Woolley, Stefan Scherer, Mohammad Soleymani 0001 |
ICMI | 6 |
| 2018 | Modeling Temporality of Human Intentions by Domain AdaptationabstractCategorizing a patient's intentions during clinical interactions in general and within motivational interviewing specifically may improve decision making in clinical treatments.Within this paper, we propose a method that models the temporal flow of a conversation and the transition between topics by using domain adaptation on a clinical dialogue corpus comprising Motivational Interviewing (MI) sessions.We deploy Bi-LSTM and topic models jointly to learn theme shifts across different time stages within these hour-long MI sessions to assess the patient's intent to change their habits or to sustain them respectively.Our experiments show promising results and improvements after considering temporality in the classification task over our baseline.This result confirms and extends related literature that has manually identified that certain phases within MI sessions are more predictive of patient outcomes than others. Xiaolei Huang 0002, Lixing Liu, Kate B. Carey, Joshua Woolley, Stefan Scherer, Brian Borsari |
EMNLP | 5 |
| 2018 | Towards Learning Nuisance-Free Representations of SpeechabstractRepresentation learning methods, such as deep autoencoders, have received sustained attention due to their ability to effectively learn meaningful representations for a variety of applications. While these learning approaches are able to derive representations from any source signal (e.g., images, language, or voice signals) and encourage the separation in dominating factor domains, they broadly treat factors of variation pertaining to nuisances (e.g., recording conditions, gender of speaker, accent etc.) no different from often subtle more interesting factors, such as paralinguistic target variables (e.g., voice quality and phonetic vowels). In paralinguistic speech analyses, nuisance variables (e.g. gender and accent of speakers) often dominate acoustic subtleties that pertain for example to the affect or well-being of the speaker. In this work, we seek to capture nuisance-free embeddings by learning two separate orthogonal representations: one representation specialized to capture nuisance factors and one that improves the representation of the target. We propose unsupervised and (semi-) supervised orthogonal autoencoders that allow us to learn informative representations of paralinguistic and phonetic targets while removing the effect of the nuisance - gender. Overall, our proposed model outperforms state-of-the-art approaches and shows improved target representations. Lixing Liu, Sayan Ghosh 0004, Stefan Scherer |
ICASSP | 3 |
| 2018 | Multimodal Analysis of Client Behavioral Change Coding in Motivational InterviewingabstractMotivational Interviewing (MI) is a widely disseminated and effective therapeutic approach for behavioral disorder treatment. Over the past decade, MI research has identified client language as a central mediator between therapist skills and subsequent behavior change. Specifically, in-session client language referred to as change talk (CT; personal arguments for change) or sustain talk (ST; personal argument against changing the status quo) has been directly related to post-session behavior change. Despite the prevalent use of MI and extensive studies of MI underlying mechanisms, most existing studies focus on the linguistic aspect of MI, especially of client change talk and sustain talk and how they as a mediator influence the outcome of MI. In this study, we perform statistical analyses on acoustic behavior descriptors to test their discriminatory powers. Then we utilize multimodality by combining acoustic features with linguistic features to improve the accuracy of client change talk prediction. Lastly, we investigate into our trained model to understand what features inform the model about client utterance class and gain insights into the nature of MISC codes. Chanuwas Aswamenakul, Lixing Liu, Kate B. Carey, Joshua Woolley, Stefan Scherer, Brian Borsari |
ICMI | 5 |
| 2018 | Influence of Individual Differences when Training Public Speaking with Virtual AudiencesabstractMultimodal interaction technologies have enabled new applications for training interpersonal skills such as public speaking. Various training paradigms have been proposed, most of them relying on some form of graphical feedback provided to the trainee in real-time during their training or after training using an after-action review tool. Another paradigm consists of using virtual characters to provide feedback through their behavior during simulated social interactions. Preliminary studies have started to explore the effectiveness of these different training paradigms; however, these have not investigated the impact of individual differences on which interaction paradigm is more efficient or motivating for different populations of users. In this article, we explore the impact of personality, public speaking anxiety, and immersive tendencies on the experiences of users training public speaking with an interactive virtual audience system providing realtime feedback through virtual audience behavior as well as delayed feedback with an after-action review tool. We found that these three factors impacted different output measures of user experience and user ratings of the system's quality. Mathieu Chollet, Pranav Ghate, Catherine Neubauer, Stefan Scherer |
IVA | 4 |
| 2018 | NADiA: Neural Network Driven Virtual Human Conversation AgentsabstractAdvances in artificial intelligence and in particular machine learning and neural networks have given rise to a new generation of virtual assistants and chatbots. Within this work, we present NADiA - Neurally Animated Dialog Agent - that leverages both the user's verbal input as well as their facial expressions to respond in a meaningful way. NADiA combines a neural language model that generates appropriate responses to user prompts, a convolutional neural network for facial expression analysis, and virtual human technology that is deployed on a mobile phone. Here, we evaluate NADiA's anthropomorphic characteristics and its ability to understand the human interlocutor using both subjective as well as objective measures. We find that NADiA significantly outperforms state of the art chatbot technology and produces comparable behavior to human generated reference outputs. Jason Wu 0001, Sayan Ghosh 0004, Mathieu Chollet, Steven Ly, Sharon Mozgai, Stefan Scherer |
IVA | 6 |
| 2018 | Unfolding the External Behavior and Inner Affective State of Teammates through Ensemble Learning: Experimental Evidence from a Dyadic Team Corpus
Aggeliki Vlachostergiou, Mark Dennison, Catherine Neubauer, Stefan Scherer, Peter Khooshabeh, Andre Harrison |
LREC | 4 |
| 2017 | Manual and automatic measures confirm - Intranasal oxytocin increases facial expressivityabstractThe effects of oxytocin on facial emotional expressivity were investigated in individuals with schizophrenia and age-matched healthy controls during the completion of a Social Judgment Task (SJT) with a double-blind, placebo-controlled, cross-over design. Although pharmacological interventions exist to help alleviate some symptoms of schizophrenia, currently available agents are not effective at improving the severity of blunted facial affect. Participant facial expressivity was previously quantified from video recordings of the SJT using a well-validated manual approach (Facial Expression Coding System; FACES). We confirm these findings using an automated computer-based approach. Using both methods we found that the administration of oxytocin significantly increased total facial expressivity in individuals with schizophrenia and increased facial expressivity at trend level in healthy controls. Secondary analysis showed that oxytocin also significantly increased the frequency of negative valence facial expressions in individuals with schizophrenia but not in healthy controls and that oxytocin did not significantly increase positive valence facial expressions in either group. Both manual coding and automatic facial analysis revealed the same pattern of findings. Considering manual annotation can be expensive and time-consuming, these results suggest that automatic facial analysis may be an efficient and cost-effective alternative to currently utilized manual approaches and may be ready for use in clinical settings. Catherine Neubauer, Sharon Mozgai, Brandon Chuang, Joshua Woolley, Stefan Scherer |
ACII | 5 |
| 2017 | Comparing models for gesture recognition of children's bullying behaviorsabstractWe explored gesture recognition applied to the problem of classifying natural physical bullying behaviors by children. To capture natural bullying behavior data, we developed a humanoid robot that used hand-coded gesture recognition to identify basic physical bullying gestures and responded by explaining why the gestures were inappropriate. Children interacted with the robot by trying various bullying behaviors, thereby allowing us to collect a natural bullying behavior dataset for training the classifiers. We trained three different sequence classifiers using the collected data and compared their effectiveness at classifying different types of common physical bullying behaviors. Overall, Hidden Conditional Random Fields achieved the highest average F1 score (0.645) over all tested gesture classes. Michael Tsang, Vadim Korolik, Stefan Scherer, Maja J. Mataric |
ACII | 3 |
| 2017 | What really matters - An information gain analysis of questions and reactions in automated PTSD screeningsabstractPost-traumatic stress disorder (PTSD) is a mental health condition which severely affects society on many levels. However, PTSD remains often undiagnosed and untreated. To improve screening access and acceptance, a fully automated virtual human was designed to conduct many PTSD screening interviews (N=198). Here, we investigate which questions elicit the most indicative multimodal behaviors for PTSD. To this effect, we employ an information gain driven ranking procedure to identify the most informative questions. Further, we look into question dependent behaviors. To evaluate the question ranking, we investigate the discriminative faculty of the top questions using support vector machines. Our results reveal that only a subset of posed questions are required to robustly detect symptoms of PTSD in subject-independent machine learning experiments. We observe strong performance and can confirm that many of the identified behaviors correspond to commonly found behavioral indicators related to PTSD and depression. Torsten Wörtwein, Stefan Scherer |
ACII | 2 |
| 2017 | Affect-LM: A Neural Language Model for Customizable Affective Text GenerationabstractSayan Ghosh, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Sayan Ghosh 0004, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer |
ACL (1) | 5 |
| 2017 | Assessing Public Speaking Ability from Thin Slices of BehaviorabstractAn important aspect of public speaking is delivery, which consists of the appropriate use of non-verbal cues to strengthen the message. Recent works have successfully predicted ratings of public speaking delivery aspects using the entire presentations of speakers. However, in other contexts, such as the assessment of personality or the prediction of job interview outcomes, it has been shown that thin slices, brief excerpts of behavior, provide enough information for raters to make accurate predictions. In this paper, we consider the use of thin slices for predicting ratings of public speaking behavior. We use a publicly available corpus of public speaking presentations and obtain ratings of full videos and thin slices. We first study how thin slices ratings are related to full video ratings. Then, we use automatic audio-visual feature extraction methods and machine learning algorithms to create models for predicting public speaking ratings, and evaluate these models for predicting thin slices ratings and full videos ratings. Mathieu Chollet, Stefan Scherer |
FG | 2 |
| 2017 | Learning representations of emotional speech with deep convolutional generative adversarial networksabstractAutomatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker that are indicative of positive or negative emotional states are often “overshadowed” by voice characteristics relating to emotional intensity or emotional activation. In this work we explore a representation learning approach that automatically derives discriminative representations of emotional speech. In particular, we investigate two machine learning strategies to improve classifier performance: (1) utilization of unlabeled data using a deep convolutional generative adversarial network (DCGAN), and (2) multitask learning. Within our extensive experiments we leverage a multitask annotated emotional corpus as well as a large unlabeled meeting corpus (around 100 hours). Our speaker-independent classification experiments show that in particular the use of unlabeled data in our investigations improves performance of the classifiers and both fully supervised baseline approaches are outperformed considerably. We improve the classification of emotional valence on a discrete 5-point scale to 43.88% and on a 3-point scale to 49.80%, which is competitive to state-of-the-art performance. Stefan Scherer |
ICASSP | 2 |
| 2017 | The relationship between task-induced stress, vocal changes, and physiological state during a dyadic team taskabstractIt is commonly known that a relationship exists between the human voice and various emotional states. Past studies have demonstrated changes in a number of vocal features, such as fundamental frequency f0 and peakSlope, as a result of varying emotional state. These voice characteristics have been shown to relate to emotional load, vocal tension, and, in particular, stress. Although much research exists in the domain of voice analysis, few studies have assessed the relationship between stress and changes in the voice during a dyadic team interaction. The aim of the present study was to investigate the multimodal interplay between speech and physiology during a high-workload, high-stress team task. Specifically, we studied task-induced effects on participants' vocal signals, specifically, the f0 and peakSlope features, as well as participants' physiology, through cardiovascular measures. Further, we assessed the relationship between physiological states related to stress and changes in the speaker's voice. We recruited participants with the specific goal of working together to diffuse a simulated bomb. Half of our sample participated in an "Ice Breaker" scenario, during which they were allowed to converse and familiarize themselves with their teammate prior to the task, while the other half of the sample served as our "Control". Fundamental frequency (f0), peakSlope, physiological state, and subjective stress were measured during the task. Results indicated that f0 and peakSlope significantly increased from the beginning to the end of each task trial, and were highest in the last trial, which indicates an increase in emotional load and vocal tension. Finally, cardiovascular measures of stress indicated that the vocal and emotional load of speakers towards the end of the task mirrored a physiological state of psychological "threat". Catherine Neubauer, Mathieu Chollet, Sharon Mozgai, Mark Dennison, Peter Khooshabeh, Stefan Scherer |
ICMI | 6 |
| 2017 | OpenMM: An Open-Source Multimodal Feature Extraction Tool
Michelle Renee Morales, Stefan Scherer, Rivka Levitan |
INTERSPEECH | 2 |
| 2017 | Racing Heart and Sweaty Palms - What Influences Users' Self-Assessments and Physiological Signals When Interacting with Virtual Audiences?
Mathieu Chollet, Talie Massachi, Stefan Scherer |
IVA | 3 |
| 2017 | Guest Editorial: Towards Machines Able to Deal with LaughterabstractThe papers in this special section focus on the concept of laughter computing. Laughter is considered a significant feature of human-human communication. Laughter is characterized by a complex behavior that includes major modules: auditory, facial expressions, body movements, and postural attitudes, and physiological signals. The goal of this special section is to gather recent achievements in laughter computing in order to trigger new research directions in this area. Maurizio Mancini, Radoslaw Niewiadomski, Shuji Hashimoto, Mary Ellen Foster, Stefan Scherer, Gualtiero Volpe |
IEEE Trans. Affect. Comput. | 5 |
| 2017 | Adolescent Suicidal Risk Assessment in Clinician-Patient InteractionabstractYouth suicide is a major public health problem. It is the third leading cause of death in the United States for ages 13 through 18. Many adolescents that face suicidal thoughts or make a suicide plan never seek professional care or help. Within this work, we evaluate both verbal and nonverbal responses to a five-item ubiquitous questionnaire to identify and assess suicidal risk of adolescents. We utilize a machine learning approach to identify suicidal from non-suicidal speech as well as characterize adolescents that repeatedly attempted suicide in the past. Our findings investigate both verbal and nonverbal behavior information of the face-to-face clinician-patient interaction. We investigate 60 audio-recorded dyadic clinician-patient interviews of 30 suicidal (13 repeaters and 17 non-repeaters) and 30 non-suicidal adolescents. The interaction between clinician and adolescents is statistically analyzed to reveal differences between suicidal versus non-suicidal adolescents and to investigate suicidal repeaters’ behaviors in comparison to suicidal non-repeaters. By using a hierarchical classifier we were able to show that the verbal responses to the ubiquitous questions sections of the interviews were useful to discriminate suicidal and non-suicidal patients. However, to additionally classify suicidal repeaters and suicidal non-repeaters more information especially nonverbal information is required. Verena Venek, Stefan Scherer, Louis-Philippe Morency, Albert A. Rizzo, John Pestian |
IEEE Trans. Affect. Comput. | 2 |
| 2016 | Native vs. non-native language fluency implications on multimodal interaction for interpersonal skills trainingabstractNew technological developments in the field of multimodal interaction show great promise for the improvement and assessment of public speaking skills. However, it is unclear how the experience of non-native speakers interacting with such technologies differs from native speakers. In particular, non-native speakers could benefit less from training with multimodal systems compared to native speakers. Additionally, machine learning models trained for the automatic assessment of public speaking ability on data of native speakers might not be performing well for assessing the performance of non-native speakers. In this paper, we investigate two aspects related to the performance and evaluation of multimodal interaction technologies designed for the improvement and assessment of public speaking between a population of English native speakers and a population of non-native English speakers. Firstly, we compare the experiences and training outcomes of these two populations interacting with a virtual audience system designed for training public speaking ability, collecting a dataset of public speaking presentations in the process. Secondly, using this dataset, we build regression models for predicting public speaking performance on both populations and evaluate these models, both on the population they were trained on and on how they generalize to the second population. Mathieu Chollet, Helmut Prendinger, Stefan Scherer |
ICMI | 3 |
| 2016 | Getting to know you: a multimodal investigation of team behavior and resilience to stressabstractTeam cohesion has been suggested to be a critical factor in emotional resilience following periods of stress. Team cohesion may depend on several factors including emotional state, communication among team members and even psychophysiological response. The present study sought to employ several multimodal techniques designed to investigate team behavior as a means of understanding resilience to stress. We recruited 40 subjects to perform a cooperative-task in gender-matched, two-person teams. They were responsible for working together to meet a common goal, which was to successfully disarm a simulated bomb. This high-workload task requires successful cooperation and communication among members. We assessed several behaviors that relate to facial expression, word choice and physiological responses (i.e., heart rate variability) within this scenario. A manipulation of an â€oeice breakerâ€❝ condition was used to induce a level of comfort or familiarity within the team prior to the task. We found that individuals in the â€oeice breakerâ€❝ condition exhibited better resilience to subjective stress following the task. These individuals also exhibited more insight and cognitive speech, more positive facial expressions and were also able to better regulate their emotional expression during the task, compared to the control. Catherine Neubauer, Joshua Woolley, Peter Khooshabeh, Stefan Scherer |
ICMI | 4 |
| 2016 | Representation Learning for Speech Emotion Recognition
Sayan Ghosh 0004, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer |
INTERSPEECH | 4 |
| 2016 | An Architecture for Biologically Grounded Real-Time Reflexive Behavior
Ulysses Bernardet, Mathieu Chollet, Steve DiPaola, Stefan Scherer |
IVA | 4 |
| 2016 | Manipulating the Perception of Virtual Audiences Using Crowdsourced Behaviors
Mathieu Chollet, Nithin Chandrashekhar, Ari Shapiro, Louis-Philippe Morency, Stefan Scherer |
IVA | 5 |
| 2016 | A Multimodal Corpus for the Assessment of Public Speaking Ability and Anxiety
Mathieu Chollet, Torsten Wörtwein, Louis-Philippe Morency, Stefan Scherer |
LREC | 4 |
| 2016 | Self-Reported Symptoms of Depression and PTSD Are Associated with Reduced Vowel Space in Screening InterviewsabstractReduced frequency range in vowel production is a well documented speech characteristic of individuals with psychological and neurological disorders. Affective disorders such as depression and post-traumatic stress disorder (PTSD) are known to influence motor control and in particular speech production. The assessment and documentation of reduced vowel space and reduced expressivity often either rely on subjective assessments or on analysis of speech under constrained laboratory conditions (e.g. sustained vowel production, reading tasks). These constraints render the analysis of such measures expensive and impractical. Within this work, we investigate an automatic unsupervised machine learning based approach to assess a speaker's vowel space. Our experiments are based on recordings of 253 individuals. Symptoms of depression and PTSD are assessed using standard self-assessment questionnaires and their cut-off scores. The experiments show a significantly reduced vowel space in subjects that scored positively on the questionnaires. We show the measure's statistical robustness against varying demographics of individuals and articulation rate. The reduced vowel space for subjects with symptoms of depression can be explained by the common condition of psychomotor retardation influencing articulation and motor control. These findings could potentially support treatment of affective disorders, like depression and PTSD in the future. Stefan Scherer, Gale M. Lucas, Jonathan Gratch, Albert A. Rizzo, Louis-Philippe Morency |
IEEE Trans. Affect. Comput. | 1 |
| 2015 | SimSensei Demonstration: A Perceptive Virtual Human Interviewer for Healthcare ApplicationsabstractWe present the SimSensei system, a fully automatic virtual agent that conducts interviews to assess indicators of psychological distress. We emphasize on the perception part of the system, a multimodal framework which captures and analyzes user state for both behavioral understanding and interactional purposes. Louis-Philippe Morency, Giota Stratou, David DeVault, Arno Hartholt, Margot Lhommet, Gale M. Lucas, Fabrizio Morbini, Kallirroi Georgila, Stefan Scherer, Jonathan Gratch, Stacy Marsella, David R. Traum, Albert A. Rizzo |
AAAI | 9 |
| 2015 | A multi-label convolutional neural network approach to cross-domain action unit detectionabstractAction Unit (AU) detection from facial images is an important classification task in affective computing. However most existing approaches use carefully engineered feature extractors along with off-the-shelf classifiers. There has also been less focus on how well classifiers generalize when tested on different datasets. In our paper, we propose a multi-label convolutional neural network approach to learn a shared representation between multiple AUs directly from the input image. Experiments on three AU datasets- CK+, DISFA and BP4D indicate that our approach obtains competitive results on all datasets. Cross-dataset experiments also indicate that the network generalizes well to other datasets, even when under different training and testing conditions. Sayan Ghosh 0004, Eugene Laksana, Stefan Scherer, Louis-Philippe Morency |
ACII | 3 |
| 2015 | Towards an affective interface for assessment of psychological distressabstractEven with the rise in use of TeleMedicine for health care and mental health, research suggests that clinicians may have difficulty reading nonverbal cues in computer-mediated situations. However, the recent progress in tracking affective markers (i.e., displays of emotional expressions on face and in voice) has opened the door to new clinical applications that might help health care providers better read nonverbal behaviors when employing TeleMedicine. For example, an interface that automatically quantified affective markers could assist clinicians in their assessment of and treatment for psychological distress (i.e., symptoms of depression and Post-traumatic Stress Disorder (PTSD)). To move towards this prospect, we will show that clinicians' judgments of these nonverbal affective markers (e.g., smile, frown, eye contact, tense voice) could be informed by such technology. The results of our evaluation suggest that clinicians' ratings of nonverbal affective markers are less predictive of psychological distress than automatically quantified affective markers. Because such quantifications are more strongly associated with psychological distress than clinician ratings of these same nonverbal behaviors, an affective interface providing quantifications of nonverbal affective markers could potentially improve assessment of psychological distress.. Gale M. Lucas, Jonathan Gratch, Stefan Scherer, Jill Boberg, Giota Stratou |
ACII | 3 |
| 2015 | A demonstration of the perception system in SimSensei, a virtual human application for healthcare interviewsabstractWe present the SimSensei system, a fully automatic virtual agent that conducts interviews to assess indicators of psychological distress. With this demo, we focus our attention on the perception part of the system, a multimodal framework which captures and analyzes user state behavior for both behavioral understanding and interactional purposes. We will demonstrate real-time user state sensing as a part of the SimSensei architecture and discuss how this technology enabled automatic analysis of behaviors related to psychological distress. Giota Stratou, Louis-Philippe Morency, David DeVault, Arno Hartholt, Edward Fast, Margot Lhommet, Gale M. Lucas, Fabrizio Morbini, Kallirroi Georgila, Stefan Scherer, Jonathan Gratch, Stacy Marsella, David R. Traum, Albert A. Rizzo |
ACII | 10 |
| 2015 | Automatic assessment and analysis of public speaking anxiety: A virtual audience case studyabstractPublic speaking has become an integral part of many professions and is central to career building opportunities. Yet, public speaking anxiety is often referred to as the most common fear in everyday life and can hinder one's ability to speak in public severely. While virtual and real audiences have been successfully utilized to treat public speaking anxiety in the past, little work has been done on identifying behavioral characteristics of speakers suffering from anxiety. In this work, we focus on the characterization of behavioral indicators and the automatic assessment of public speaking anxiety. We identify several indicators for public speaking anxiety, among them are less eye contact with the audience, reduced variability in the voice, and more pauses. We automatically assess the public speaking anxiety as reported by the speakers through a self-assessment questionnaire using a speaker independent paradigm. Our approach using ensemble trees achieves a high correlation between ground truth and our estimation (r=0.825). Complementary to automatic measures of anxiety, we are also interested in speakers' perceptual differences when interacting with a virtual audience based on their level of anxiety in order to improve and further the development of virtual audiences for the training of public speaking and the reduction of anxiety. Torsten Wörtwein, Louis-Philippe Morency, Stefan Scherer |
ACII | 3 |
| 2015 | Exploring feedback strategies to improve public speaking: an interactive virtual audience frameworkabstractGood public speaking skills convey strong and effective communication, which is critical in many professions and used in everyday life. The ability to speak publicly requires a lot of training and practice. Recent technological developments enable new approaches for public speaking training that allow users to practice in a safe and engaging environment. We explore feedback strategies for public speaking training that are based on an interactive virtual audience paradigm. We investigate three study conditions: (1) a non-interactive virtual audience (control condition), (2) direct visual feedback, and (3) nonverbal feedback from an interactive virtual audience. We perform a threefold evaluation based on self-assessment questionnaires, expert assessments, and two objectively annotated measures of eye-contact and avoidance of pause fillers. Our experiments show that the interactive virtual audience brings together the best of both worlds: increased engagement and challenge as well as improved public speaking skills as judged by experts. Mathieu Chollet, Torsten Wörtwein, Louis-Philippe Morency, Ari Shapiro, Stefan Scherer |
UbiComp | 5 |
| 2015 | Reduced vowel space is a robust indicator of psychological distress: A cross-corpus analysisabstractReduced frequency range in vowel production is a well documented speech characteristic of individuals' with psychological and neurological disorders. Depression is known to influence motor control and in particular speech production. The assessment and documentation of reduced vowel space and associated perceived hypoarticulation and reduced expressivity often rely on subjective assessments. Within this work, we investigate an automatic unsupervised machine learning approach to assess a speaker's vowel space within three distinct speech corpora and compare observed vowel space measures of subjects with and without psychological conditions associated with psychological distress, namely depression, post-traumatic stress disorder (PTSD), and suicidality. Our experiments are based on recordings of over 300 individuals. The experiments show a significantly reduced vowel space in conversational speech for depression, PTSD, and suicidality. We further observe a similar trend of reduced vowel space for read speech. A possible explanation for a reduced vowel space is psychomotor retardation, a common symptom of depression that influences motor control and speech production. Stefan Scherer, Louis-Philippe Morency, Jonathan Gratch, John Pestian |
ICASSP | 1 |
| 2015 | Acoustic and para-verbal indicators of persuasiveness in social multimediaabstractPersuasive communication and interaction play an important and pervasive role in many aspects of our lives. With the rapid growth of social multimedia websites such as YouTube, it has become more important and useful to understand persuasiveness in the context of online social multimedia content. In this paper, we present our results of conducting various analyses of persuasiveness in speech with our multimedia corpus of 1,000 movie review videos obtained from ExpoTV.com, a popular social multimedia website. Our experiments firstly show that a speaker's level of persuasiveness can be predicted from acoustic characteristics and para-verbal cues related to speech fluency. Secondly, we show that taking acoustic cues in different time periods of a movie review can improve the performance of predicting a speaker's level of persuasiveness. Lastly, we show that a speaker's positive or negative attitude toward a topic influences the prediction performance as well. Han Suk Shim, Sunghyun Park 0001, Moitreya Chatterjee, Stefan Scherer, Kenji Sagae, Louis-Philippe Morency |
ICASSP | 4 |
| 2015 | Combining Two Perspectives on Classifying Multimodal Data for Recognizing Speaker TraitsabstractHuman communication involves conveying messages both through verbal and non-verbal channels (facial expression, gestures, prosody, etc.). Nonetheless, the task of learning these patterns for a computer by combining cues from multiple modalities is challenging because it requires effective representation of the signals and also taking into consideration the complex interactions between them. From the machine learning perspective this presents a two-fold challenge: a) Modeling the intermodal variations and dependencies; b) Representing the data using an apt number of features, such that the necessary patterns are captured but at the same time allaying concerns such as over-fitting. In this work we attempt to address these aspects of multimodal recognition, in the context of recognizing two essential speaker traits, namely passion and credibility of online movie reviewers. We propose a novel ensemble classification approach that combines two different perspectives on classifying multimodal data. Each of these perspectives attempts to independently address the two-fold challenge. In the first, we combine the features from multiple modalities but assume inter-modality conditional independence. In the other one, we explicitly capture the correlation between the modalities but in a space of few dimensions and explore a novel clustering based kernel similarity approach for recognition. Additionally, this work investigates a recent technique for encoding text data that captures semantic similarity of verbal content and preserves word-ordering. The experimental results on a recent public dataset shows significant improvement of our approach over multiple baselines. Finally, we also analyze the most discriminative elements of a speaker's non-verbal behavior that contribute to his/her perceived credibility/passionateness. Moitreya Chatterjee, Sunghyun Park 0001, Louis-Philippe Morency, Stefan Scherer |
ICMI | 4 |
| 2015 | Public Speaking Training with a Multimodal Interactive Virtual Audience FrameworkabstractWe have developed an interactive virtual audience platform for public speaking training. Users' public speaking behavior is automatically analyzed using multimodal sensors, and ultimodal feedback is produced by virtual characters and generic visual widgets depending on the user's behavior. The flexibility of our system allows to compare different interaction mediums (e.g. virtual reality vs normal interaction), social situations (e.g. one-on-one meetings vs large audiences) and trained behaviors (e.g. general public speaking performance vs specific behaviors). Mathieu Chollet, Kalin Stefanov, Helmut Prendinger, Stefan Scherer |
ICMI | 4 |
| 2015 | Exploring Behavior Representation for Learning AnalyticsabstractMultimodal analysis has long been an integral part of studying learning. Historically multimodal analyses of learning have been extremely laborious and time intensive. However, researchers have recently been exploring ways to use multimodal computational analysis in the service of studying how people learn in complex learning environments. In an effort to advance this research agenda, we present a comparative analysis of four different data segmentation techniques. In particular, we propose affect- and pose-based data segmentation, as alternatives to human-based segmentation, and fixed-window segmentation. In a study of ten dyads working on an open-ended engineering design task, we find that affect- and pose-based segmentation are more effective, than traditional approaches, for drawing correlations between learning-relevant constructs, and multimodal behaviors. We also find that pose-based segmentation outperforms the two more traditional segmentation strategies for predicting student success on the hands-on task. In this paper we discuss the algorithms used, our results, and the implications that this work may have in non-education-related contexts. Marcelo Worsley, Stefan Scherer, Louis-Philippe Morency, Paulo Blikstein |
ICMI | 2 |
| 2015 | Multimodal Public Speaking Performance AssessmentabstractThe ability to speak proficiently in public is essential for many professions and in everyday life. Public speaking skills are difficult to master and require extensive training. Recent developments in technology enable new approaches for public speaking training that allow users to practice in engaging and interactive environments. Here, we focus on the automatic assessment of nonverbal behavior and multimodal modeling of public speaking behavior. We automatically identify audiovisual nonverbal behaviors that are correlated to expert judges' opinions of key performance aspects. These automatic assessments enable a virtual audience to provide feedback that is essential for training during a public speaking performance. We utilize multimodal ensemble tree learners to automatically approximate expert judges' evaluations to provide post-hoc performance assessments to the speakers. Our automatic performance evaluation is highly correlated with the experts' opinions with r = 0.745 for the overall performance assessments. We compare multimodal approaches with single modalities and find that the multimodal ensembles consistently outperform single modalities. Torsten Wörtwein, Mathieu Chollet, Boris Schauerte, Louis-Philippe Morency, Rainer Stiefelhagen, Stefan Scherer |
ICMI | 6 |
| 2015 | A Multimodal Predictive Model of Successful Debaters or How I Learned to Sway VotesabstractInterpersonal skills such as public speaking are essential assets for a large variety of professions and in everyday life. The ability to communicate in social environments often greatly influences a person's career development, can help resolve conflict, gain the upper hand in negotiations, or sway the public opinion. We focus our investigations on a special form of public speaking, namely public debates of socioeconomic issues that affect us all. In particular, we analyze performances of expert debaters recorded through the Intelligence Squared U.S. (IQ2US) organization. IQ2US collects high-quality audiovisual recordings of these debates and publishes them online free of charge. We extract audiovisual nonverbal behavior descriptors, including facial expressions, voice quality characteristics, and surface level linguistic characteristics. Within our experiments we investigate if it is possible to automatically predict if a debater or his/her team are going to sway the most votes after the debate using multimodal machine learning and fusion approaches. We identify unimodal nonverbal behaviors that characterize successful debaters and our investigations reveal that multimodal machine learning approaches can reliably predict which individual (~75% accuracy) or team (85% accuracy) is going to win the most votes in the debate. We created a database consisting of over 30 debates with four speakers per debate suitable for public speaking skill analysis and plan to make this database publicly available for the research community. Maarten F. Brilman, Stefan Scherer |
ACM Multimedia | 2 |
| 2015 | Preface of pattern recognition in human computer interaction
Friedhelm Schwenker, Stefan Scherer, Louis-Philippe Morency |
Pattern Recognit. Lett. | 2 |
| 2015 | Emotion recognition from speech signals via a probabilistic echo-state network
Edmondo Trentin, Stefan Scherer, Friedhelm Schwenker |
Pattern Recognit. Lett. | 2 |
| 2015 | A review of depression and suicide risk assessment using speech analysis
Nicholas Cummins, Stefan Scherer, Jarek Krajewski, Sebastian Schnieder, Julien Epps, Thomas F. Quatieri |
Speech Commun. | 2 |
| 2015 | I Can Already Guess Your Answer: Predicting Respondent Reactions during Dyadic NegotiationabstractNegotiation is a component deeply ingrained in our daily lives, and it can be challenging for a person to predict the respondent's reaction (acceptance or rejection) to a negotiation offer. In this work, we focus on finding acoustic and visual behavioral cues that are predictive of the respondent's immediate reactions using a face-to-face negotiation dataset, which consists of 42 dyadic interactions in a simulated negotiation setting. We show our results of exploring four different sources of information, namely nonverbal behavior of the proposer, that of the respondent, mutual behavior between the interactants related to behavioral symmetry and asymmetry, and past negotiation history between the interactants. Firstly, we show that considering other sources of information (other than the nonverbal behavior of the respondent) can also have comparable performance in predicting respondent reactions. Secondly, we show that automatically extracted mutual behavioral cues of symmetry and asymmetry are predictive partially due to their capturing information of the nature of the interaction itself, whether it is cooperative or competitive. Lastly, we identify audio-visual behavioral cues that are most predictive of the respondent's immediate reactions. Sunghyun Park 0001, Stefan Scherer, Jonathan Gratch, Peter J. Carnevale, Louis-Philippe Morency |
IEEE Trans. Affect. Comput. | 2 |
| 2014 | Context-based signal descriptors of heart-rate variability for anxiety assessmentabstractIn this paper, we investigate the role of multiple context-based heart-rate variability descriptors for evaluating a person's psychological health, specifically anxiety disorders. The descriptors are extracted from visually sensed heart-rate signals obtained during the course of a semi-structured interview with a virtual human and can potentially integrate question context as well. The proposed descriptors are motivated by prior related work and are constructed based on histogram-based approaches, time and frequency domain analysis of heart-rate variability. In order to contextualize our descriptors, we use information about the polarity and intimacy levels of the questions asked. Our experiments reveal that the descriptors, both with and without context, perform far better than chance in predicting anxiety. Further on, we perform at-a-par with the state-of-the-art in predicting anxiety and other psychological disorders when we integrate the question context information into the descriptors. Moitreya Chatterjee, Giota Stratou, Stefan Scherer, Louis-Philippe Morency |
ICASSP | 3 |
| 2014 | COVAREP - A collaborative voice analysis repository for speech technologiesabstractSpeech processing algorithms are often developed demonstrating improvements over the state-of-the-art, but sometimes at the cost of high complexity. This makes algorithm reimplementations based on literature difficult, and thus reliable comparisons between published results and current work are hard to achieve. This paper presents a new collaborative and freely available repository for speech processing algorithms called COVAREP, which aims at fast and easy access to new speech processing algorithms and thus facilitating research in the field. We envisage that COVAREP will allow more reproducible research by strengthening complex implementations through shared contributions and openly available code which can be discussed, commented on and corrected by the community. Presently COVAREP contains contributions from five distinct laboratories and we encourage contributions from across the speech processing research field. In this paper, we provide an overview of the current offerings of COVAREP and also include a demonstration of the algorithms through an emotion classification experiment. Gilles Degottex, John Kane 0002, Thomas Drugman, Tuomo Raitio, Stefan Scherer |
ICASSP | 5 |
| 2014 | Dyadic Behavior Analysis in Depression Severity Assessment InterviewsabstractPrevious literature suggests that depression impacts vocal timing of both participants and clinical interviewers but is mixed with respect to acoustic features. To investigate further, 57 middle-aged adults (men and women) with Major Depression Disorder and their clinical interviewers (all women) were studied. Participants were interviewed for depression severity on up to four occasions over a 21 week period using the Hamilton Rating Scale for Depression (HRSD), which is a criterion measure for depression severity in clinical trials. Acoustic features were extracted for both participants and interviewers using COVAREP Toolbox. Missing data occurred due to missed appointments, technical problems, or insufficient vocal samples. Data from 36 participants and their interviewers met criteria and were included for analysis to compare between high and low depression severity. Acoustic features for participants varied between men and women as expected, and failed to vary with depression severity for participants. For interviewers, acoustic characteristics strongly varied with severity of the interviewee's depression. Accommodation - the tendency of interactants to adapt their communicative behavior to each other - between interviewers and interviewees was inversely related to depression severity. These findings suggest that interviewers modify their acoustic features in response to depression severity, and depression severity strongly impacts interpersonal accommodation. Stefan Scherer, Zakia Hammal, Ying Yang 0007, Louis-Philippe Morency, Jeffrey F. Cohn |
ICMI | 1 |
| 2014 | The Distress Analysis Interview Corpus of human and computer interviews
Jonathan Gratch, Ron Artstein, Gale M. Lucas, Giota Stratou, Stefan Scherer, Angela Nazarian, Rachel Wood, Jill Boberg, David DeVault, Stacy Marsella, David R. Traum, Albert A. Rizzo, Louis-Philippe Morency |
LREC | 5 |
| 2014 | Adolescent suicidal risk assessment in clinician-patient interaction: A study of verbal and acoustic behaviorsabstractSuicide among adolescents is a major public health problem: it is the third leading cause of death in the US for ages 13-18. Up to now, there is no objective ways to assess the suicidal risk, i.e. whether a patient is non-suicidal, suicidal re-attempter (i.e. repeater) or suicidal non-repeater (i.e. individuals with one suicide attempt or showing signs of suicidal gestures or ideation). Therefore, features of the conversation including verbal information and nonverbal acoustic information were investigated from 60 audio-recorded interviews of 30 suicidal (13 repeaters and 17 non-repeaters) and 30 non-suicidal adolescents interviewed by a social worker. The interaction between clinician and patients was statistically analyzed to reveal differences between suicidal vs. non-suicidal adolescents and to investigate suicidal repeaters' behaviors in comparison to suicidal non-repeaters. By using a hierarchical ensemble classifier we were able to successfully discriminate non-suicidal patients, suicidal repeaters and suicidal non-repeaters. Verena Venek, Stefan Scherer, Louis-Philippe Morency, Albert A. Rizzo, John Pestian |
SLT | 2 |
| 2014 | Automatic audiovisual behavior descriptors for psychological disorder analysis
Stefan Scherer, Giota Stratou, Gale M. Lucas, Marwa Mahmoud, Jill Boberg, Jonathan Gratch, Albert A. Rizzo, Louis-Philippe Morency |
Image Vis. Comput. | 1 |
| 2014 | Investigating automatic measurements of prosodic accommodation and its dynamics in social interaction
Céline De Looze, Stefan Scherer, Brian Vaughan, Nick Campbell 0001 |
Speech Commun. | 2 |
| 2013 | Mutual Behaviors during Dyadic Negotiation: Automatic Prediction of Respondent ReactionsabstractIn this paper, we analyze face-to-face negotiation interactions with the goal of predicting the respondent's immediate reaction (i.e., accept or reject) to a negotiation offer. Supported by the theory of social rapport, we focus on mutual behaviors which are defined as nonverbal characteristics that occur due to interactional influence. These patterns include behavioral symmetry (e.g., synchronized smiles) as well as asymmetry (e.g., opposite postures) between the two negotiators. In addition, we put emphasis on finding audio-visual mutual behaviors that can be extracted automatically, with the vision of a real-time decision support tool. We introduce a dyadic negotiation dataset consisting of 42 face-to-face interactions and show experiments confirming the importance of multimodal and mutual behaviors. Sunghyun Park 0001, Stefan Scherer, Jonathan Gratch, Peter J. Carnevale, Louis-Philippe Morency |
ACII | 2 |
| 2013 | Automatic Nonverbal Behavior Indicators of Depression and PTSD: Exploring Gender DifferencesabstractIn this paper, we show that gender plays an important role in the automatic assessment of psychological conditions such as depression and post-traumatic stress disorder (PTSD). We identify a directly interpretable and intuitive set of predictive indicators, selected from three general categories of nonverbal behaviors: affect, expression variability and motor variability. For the analysis, we introduce a semi-structured virtual human interview dataset which includes 53 video recorded interactions. Our experiments on automatic classification of psychological conditions show that a gender-dependent approach significantly improves the performance over a gender agnostic one. Giota Stratou, Stefan Scherer, Jonathan Gratch, Louis-Philippe Morency |
ACII | 2 |
| 2013 | Speaker and language independent voice quality classification applied to unlabelled corpora of expressive speechabstractVoice quality plays a pivotal role in speech style variation. Therefore, control and analysis of voice quality is critical for many areas of speech technology. Until now, most work has focused on small purpose built corpora. In this paper we apply state-of-the-art voice quality analysis to large speech corpora built for expressive speech synthesis. A fuzzy-input fuzzy-output support vector machine classifier is trained and validated using features extracted from these corpora. We then apply this classifier to freely available audiobook data and demonstrate a clustering of the voice qualities that approximates the performance of human perceptual ratings. The ability to detect voice quality variation in these widely available unlabelled audiobook corpora means that the proposed method may be used as a valuable resource in expressive speech synthesis. John Kane 0002, Stefan Scherer, Matthew P. Aylett, Louis-Philippe Morency, Christer Gobl |
ICASSP | 2 |
| 2013 | Investigating the speech characteristics of suicidal adolescentsabstractSuicide is a very serious problem. In the United states it ranks as the second most frequent cause of death among teenagers between the ages of 12 and 17. In this work, we investigate speech characteristics of prosody as well as voice quality in a dyadic interview corpus with suicidal and non-suicidal adolescents. In these interviews the adolescents answer specifically designed questions. Based on this limited dataset, we reveal statistically significant differences in the speech patterns of suicidal adolescents within the investigated interview corpus. Further, we investigate the classification capabilities of machine learning approaches both on an utterance as well as an interview level. The work shows promising results in a speaker-independent classification experiment based on only a dozen speech features. We believe that once the algorithms are refined and integrated with other methods, they may be of value to the clinician. Stefan Scherer, John Pestian, Louis-Philippe Morency |
ICASSP | 1 |
| 2013 | ICMI 2013 grand challenge workshop on multimodal learning analyticsabstractAdvances in learning analytics are contributing new empirical findings, theories, methods, and metrics for understanding how students learn. It also contributes to improving pedagogical support for students' learning through assessment of new digital tools, teaching strategies, and curricula. Multimodal learning analytics (MMLA)[1] is an extension of learning analytics and emphasizes the analysis of natural rich modalities of communication across a variety of learning contexts. This MMLA Grand Challenge combines expertise from the learning sciences and machine learning in order to highlight the rich opportunities that exist at the intersection of these disciplines. As part of the Grand Challenge, researchers were asked to predict: (1) which student in a group was the dominant domain expert, and (2) which problems that the group worked on would be solved correctly or not. Analyses were based on a combination of speech, digital pen and video data. This paper describes the motivation for the grand challenge, the publicly available data resources and results reported by the challenge participants. The results demonstrate that multimodal prediction of the challenge goals: (1) is surprisingly reliable using rich multimodal data sources, (2) can be accomplished using any of the three modalities explored, and (3) need not be based on content analysis. Louis-Philippe Morency, Sharon L. Oviatt, Stefan Scherer, Nadir Weibel, Marcelo Worsley |
ICMI | 3 |
| 2013 | Audiovisual behavior descriptors for depression assessmentabstractWe investigate audiovisual indicators, in particular measures of reduced emotional expressivity and psycho-motor retardation, for depression within semi-structured virtual human interviews. Based on a standard self-assessment depression scale we investigate the statistical discriminative strength of the audiovisual features on a depression/no-depression basis. Within subject-independent unimodal and multimodal classification experiments we find that early feature-level fusion yields promising results and confirms the statistical findings. We further correlate the behavior descriptors with the assessed depression severity and find considerable correlation. Lastly, a joint multimodal factor analysis reveals two prominent factors within the data that show both statistical discriminative power as well as strong linear correlation with the depression severity score. These preliminary results based on a standard factor analysis are promising and motivate us to investigate this approach further in the future, while incorporating additional modalities. Stefan Scherer, Giota Stratou, Louis-Philippe Morency |
ICMI | 1 |
| 2013 | A comparative study of glottal open quotient estimation techniquesabstractThe robust and efficient extraction of features related to the glottal excitation source has become increasingly important for speech technology. The glottal open quotient (OQ) is one rel-evant measurement which is known to significantly vary with changes in voice quality on a breathy to tense continuum. The extraction of OQ, however, is hampered in the time-domain by the difficulty in consistently locating the point of glottal open-ing as well the computational load of its measurement. De-termining OQ correlates in the frequency domain is an attrac-tive alternative, however the lower frequencies of glottal source spectrum are also affected by other aspects of the glottal pulse shape thereby precluding closed-form solutions and straightfor-ward mappings. The present study provides a comparison of three OQ estimation methods and shows a new method based on spectral features and artificial neural networks to outperform existing methods in terms of discrimination of voice quality, lower error values on a large volume of speech data and dra-matically reduced computation time. John Kane 0002, Stefan Scherer, Louis-Philippe Morency, Christer Gobl |
INTERSPEECH | 2 |
| 2013 | Prediction of strategy and outcome as negotiation unfolds by using basic verbal and behavioral featuresabstractNegotiations can be characterized by the strategy participants adopt to achieve their ends (e.g., individualistic strategies are based on self-interest, cooperative strategies are used when participants try to maximize the joint gain, while competitive strategies focus on maximizing each participant’s score against the other) and the outcomes that each participant achieves in the negotiation. This paper investigates the process and the result of predicting the outcome and strategy of participants throughout the progress of the negotiation by using basic, easy to extract, linguistic and acoustic features. We evaluate our approach on a face-to-face negotiation dataset consisting of 41 dyadic interactions and show that it’s possible to significantly improve over a majority-class baseline in tasks of predicting the strategy and outcome of the interaction by analyzing only basic low level features of the negotiation. Elnaz Nouri, Sunghyun Park 0001, Stefan Scherer, Jonathan Gratch, Peter J. Carnevale, Louis-Philippe Morency, David R. Traum |
INTERSPEECH | 3 |
| 2013 | Investigating voice quality as a speaker-independent indicator of depression and PTSDabstractWe seek to investigate voice quality characteristics, in particular on a breathy to tense dimension, as an indicator for psychological distress, i.e. depression and post-traumatic stress disorder (PTSD), within semi-structured virtual human interviews. Our evaluation identifies significant differences between the voice quality of psychologically distressed participants and not-distressed participants within this limited corpus. We investigate the capability of automatic algorithms to classify psychologically distressed speech in speaker-independent experiments. Additionally, we examine the impact of the posed questions’ affective polarity, as motivated by findings in the literature on positive stimulus attenuation and negative stimulus potentiation in emotional reactivity of psychologically distressed participants. The experiments yield promising results using standard machine learning algorithms and solely four distinct features capturing the tenseness of the speaker’s voice. Stefan Scherer, Giota Stratou, Jonathan Gratch, Louis-Philippe Morency |
INTERSPEECH | 1 |
| 2013 | Cicero - Towards a Multimodal Virtual Audience Platform for Public Speaking Training
Ligia Maria Batrinca, Giota Stratou, Ari Shapiro, Louis-Philippe Morency, Stefan Scherer |
IVA | 5 |
| 2013 | Verbal indicators of psychological distress in interactive dialogue with a virtual human
David DeVault, Kallirroi Georgila, Ron Artstein, Fabrizio Morbini, David R. Traum, Stefan Scherer, Albert A. Rizzo, Louis-Philippe Morency |
SIGDIAL Conference | 6 |
| 2013 | Investigating fuzzy-input fuzzy-output support vector machines for robust voice quality classification
Stefan Scherer, John Kane 0002, Christer Gobl, Friedhelm Schwenker |
Comput. Speech Lang. | 1 |
| 2012 | Detecting a targeted voice style in an audiobook using voice quality featuresabstractAudiobooks are known to contain a variety of expressive speaking styles that occur as a result of the narrator mimicking a character in a story, or expressing affect. An accurate modeling of this variety is essential for the purposes of speech synthesis from an audiobook. Voice quality differences are important features characterizing these different speaking styles, which are realized on a gradient and are often difficult to predict from the text. The present study uses a parameter characterizing breathy to tense voice qualities using features of the wavelet transform, and a measure for identifying creaky segments in an utterance. Based on these features, a combination of supervised and unsupervised classification is used to detect the regions in an audiobook, where the speaker changes his regular voice quality to a particular voice style. The target voice style candidates are selected based on the agreement of the supervised classifier ensemble output, and evaluated in a listening test. Éva Székely, John Kane 0002, Stefan Scherer, Christer Gobl, Julie Carson-Berndsen |
ICASSP | 3 |
| 2012 | Step-wise emotion recognition using concatenated-HMMabstractHuman emotion is an important part of human-human communication, since the emotional state of an individual often affects the way that he/she reacts to others. In this paper, we present a method based on concatenated Hidden Markov Model (co-HMM) to infer the dimensional and continuous emotion labels from audio-visual cues. Our method is based on the assumption that continuous emotion levels can be modeled by a set of discrete values. Based on this, we represent each emotional dimension by step-wise label classes, and learn the intrinsic and extrinsic dynamics using our co-HMM model. We evaluate our approach on the Audio-Visual Emotion Challenge (AVEC 2012) dataset. Our results show considerable improvement over the baseline regression model presented with the AVEC 2012. Derya Ozkan, Stefan Scherer, Louis-Philippe Morency |
ICMI | 2 |
| 2012 | 1st international workshop on multimodal learning analytics: extended abstractabstractThis summary describes the 1st International Workshop on Multimodal Learning Analytics. This area of study brings together the technologies of multimodal analysis with the learning sciences. The intersection of these domains should enable researchers to foster an improved understanding of student learning, lead to the creation of more natural and enriching learning interfaces, and motivate the development of novel techniques for tackling challenges that are specific of education. Stefan Scherer, Marcelo Worsley, Louis-Philippe Morency |
ICMI | 1 |
| 2012 | Perception Markup Language: Towards a Standardized Representation of Perceived Nonverbal Behaviors
Stefan Scherer, Stacy Marsella, Giota Stratou, Yuyu Xu, Fabrizio Morbini, Alesia Egan, Albert A. Rizzo, Louis-Philippe Morency |
IVA | 1 |
| 2012 | An audiovisual political speech analysis incorporating eye-tracking and perception data
Stefan Scherer, Georg Layher, John Kane 0002, Heiko Neumann, Nick Campbell 0001 |
LREC | 1 |
| 2012 | Spotting laughter in natural multiparty conversations: A comparison of automatic online and offline approaches using audiovisual dataabstractIt is essential for the advancement of human-centered multimodal interfaces to be able to infer the current user's state or communication state. In order to enable a system to do that, the recognition and interpretation of multimodal social signals (i.e., paralinguistic and nonverbal behavior) in real-time applications is required. Since we believe that laughs are one of the most important and widely understood social nonverbal signals indicating affect and discourse quality, we focus in this work on the detection of laughter in natural multiparty discourses. The conversations are recorded in a natural environment without any specific constraint on the discourses using unobtrusive recording devices. This setup ensures natural and unbiased behavior, which is one of the main foci of this work. To compare results of methods, namely Gaussian Mixture Model (GMM) supervectors as input to a Support Vector Machine (SVM), so-called Echo State Networks (ESN), and a Hidden Markov Model (HMM) approach, are utilized in online and offline detection experiments. The SVM approach proves very accurate in the offline classification task, but is outperformed by the ESN and HMM approach in the online detection (F1scores: GMM SVM 0.45, ESN 0.63, HMM 0.72). Further, we were able to utilize the proposed HMM approach in a cross-corpus experiment without any retraining with respectable generalization capability (F1score: 0.49). The results and possible reasons for these outcomes are shown and discussed in the article. The proposed methods may be directly utilized in practical tasks such as the labeling or the online detection of laughter in conversational data and affect-aware applications. Stefan Scherer, Michael Glodek, Friedhelm Schwenker, Nick Campbell 0001, Günther Palm |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2011 | Multiple Classifier Systems for the Classification of Audio-Visual Emotional States
Michael Glodek, Stephan Tschechne, Georg Layher, Martin Schels, Tobias Brosch, Stefan Scherer, Markus Kächele, Miriam Schmidt, Heiko Neumann, Günther Palm, Friedhelm Schwenker |
ACII (2) | 6 |
| 2011 | How Low Level Observations Can Help to Reveal the User's State in HCI
Stefan Scherer, Martin Schels, Günther Palm |
ACII (2) | 1 |
| 2011 | Conditioned Hidden Markov Model Fusion for Multimodal ClassificationabstractClassification using hidden Markov models (HMM) is in general done by comparing the model likelihoods and choosing the class more likely to have generated the data. This work investigates a conditioned HMM which additionally provides a probability for a class label and compares different fusion strategies. The notion is two-fold: on the one hand applications in affective computing might pass their uncertainty of the classification to the next processing unit, on the other hand different streams might be fused to increase the performance. The data set studied incorporates two modalities and is based on a naturalistic multiparty dialogue. The goal is to discriminate between laughter and utterances. It turned out that the conditioned HMM outperforms classical HMM using different late fusion approaches while additionally providing a certainty about class decision. Michael Glodek, Stefan Scherer, Friedhelm Schwenker |
INTERSPEECH | 2 |
| 2011 | On the Use of Multimodal Cues for the Prediction of Degrees of Involvement in Spontaneous ConversationabstractQuantifying the degree of involvement of a group of participants in a conversation is a task which humans accomplish every day, but it is something that, as of yet, machines are unable to do. In this study we first investigate the correlation between visual cues (gaze and blinking rate) and involvement. We then test the suitability of prosodic cues (acoustic model) as well as gaze and blinking (visual model) for the prediction of the degree of involvement by using a support vector machine (SVM). We also test whether the fusion of the acoustic and the visual model improves the prediction. We show that we are able to predict three classes of involvement with an reduction of error rate of 0.30 (accuracy =0.68). Catharine Oertel, Stefan Scherer, Nick Campbell 0001 |
INTERSPEECH | 2 |
| 2010 | Comparing measures of synchrony and alignment in dialogue speech timing with respect to turn-taking activity
Nick Campbell 0001, Stefan Scherer |
INTERSPEECH | 2 |
| 2010 | It takes two to tango - assessing the impact of delay on conversational interactivity on perceived speech quality
Sebastian Egger-Lampl, Raimund Schatz, Stefan Scherer |
INTERSPEECH | 3 |
| 2010 | An Open Source Process Engine Framework for Realtime Pattern Recognition and Information Fusion Tasks
Volker Fritzsch, Stefan Scherer, Friedhelm Schwenker |
LREC | 2 |
| 2010 | Developing an Expressive Speech Labeling Tool Incorporating the Temporal Characteristics of Emotion
Stefan Scherer, Ingo Siegert, Lutz Bigalke, Sascha Meudt |
LREC | 1 |
| 2010 | Evaluation of the PIT Corpus Or What a Difference a Face Makes?
Petra-Maria Strauß, Stefan Scherer, Georg Layher, Holger Hoffmann |
LREC | 2 |
| 2009 | The GMM-SVM Supervector Approach for the Recognition of the Emotional Status from Speech
Friedhelm Schwenker, Stefan Scherer, Yasmine M. Magdi, Günther Palm |
ICANN (1) | 2 |
| 2008 | Emotion Recognition from Speech: Stress Experiment
Stefan Scherer, Hansjörg Hofmann, Malte Lampmann, Martin Pfeil, Steffen Rhinow, Friedhelm Schwenker, Günther Palm |
LREC | 1 |
| 2008 | A Flexible Wizard of Oz Environment for Rapid Prototyping
Stefan Scherer, Petra-Maria Strauß |
LREC | 1 |
| 2008 | The PIT Corpus of German Multi-Party Dialogues
Petra-Maria Strauß, Holger Hoffmann, Wolfgang Minker, Heiko Neumann, Günther Palm, Stefan Scherer, Harald C. Traue, Ulrich Weidenbacher |
LREC | 6 |
| 2007 | A Novel Feature for Emotion Recognition in Voice Based Applications
Hari Krishna Maganti, Stefan Scherer, Günther Palm |
ACII | 2 |
| 2006 | Wizard-of-Oz Data Collection for Perception and Interaction in Multi-User Environments
Petra-Maria Strauß, Holger Hoffmann, Wolfgang Minker, Heiko Neumann, Günther Palm, Stefan Scherer, Friedhelm Schwenker, Harald C. Traue, Welf Walter, Ulrich Weidenbacher |
LREC | 6 |
| 2000 | 3D Model Based Pose Determination in Real-Time: Strategies, Convergence, AccuracabstractIn this paper, a new real-time model based pose determination system based on the Gauss-Newton method is described. Various model fitting strategies, which use gradient search, are discussed and compared. A novel strategy, the adaptive perpendicular search, is proposed for both matching ambiguity avoidance and convergence enforcement. The system is shown to be highly effective in practice, while keeping the computational cost low. A sufficient initial pose estimate as an input for the Gauss-Newton based methods is obtained from the parametric eigenspace. With certain restrictions, the overall pose determination system performs in real-time, as required in industrial applications. Thomas Auer, Gernot Bachler, Stefan Scherer, Axel Pinz |
ICPR | 4 |
| 2000 | A Novel Bidirectional Framework for Control and Refinement of Area Based Correlation TechniquesabstractA key issue in performing stereo reconstruction is to find stereo correspondences. For this purpose a novel unified framework for arbitrary area based correlation methods is presented. The commonly neglected problem of ambiguities is addressed and a solution based on relaxation of match candidates is introduced. The method comes with the major benefits of a reverse matching step at almost no extra computational cost. The results of one match direction is used to select match candidates. Then these candidates are correlated backward with a different method. The proposed framework does not require any initial estimations. Experiments are performed on ground truth under varying geometric and varying noise conditions. Finally, the method is compared to a number of well known techniques. Peter Werth, Stefan Scherer |
ICPR | 2 |
| 2000 | Vision Guided Bin Picking and Mounting in a Flexible Assembly Cell
Gernot Bachler, Stefan Scherer |
IEA/AIE | 3 |
| 1999 | A Vision Driven Automatic Assembly Unit
Gernot Bachler, Reinhard Röhrer, Stefan Scherer, Axel Pinz |
CAIP | 4 |
| 1999 | Subpixel Stereo Matching by Robust Estimation of Local Distortion Using Gabor Filters
Peter Werth, Stefan Scherer, Axel Pinz |
CAIP | 2 |
| 1999 | The Discriminatory Power of Ordinal Measures - Towards a New CoefficientabstractPerspective distortion, occlusion and specular reflection are challenging problems in shape-from-stereo. In this paper we review one recently published area-based stereo matching algorithm (Bhat and Nayar, 1998) designed to be robust in these cases. Although the algorithm is an important contribution to stereo-matching, we show that its coefficient has a low discriminatory power, which leads to a significant number of multiple best matches. In order to cope with this drawback we introduce a new normalized ordinal correlation coefficient. Experiments showing the behavior of the proposed coefficient are performed on various datasets including real data with ground truth. The new coefficient reduces the occurrence of multiple best matches to almost zero per cent. It also shows a more robust and equally accurate behavior. These benefits are achieved at almost no additional computational costs. Stefan Scherer, Axel Pinz, Peter Werth |
CVPR | 1 |
| 1998 | Robust adaptive window matching by homogeneity constraint and integration of descriptionsabstractThe crucial element of existing binocular stereo algorithms is to establish point to point correspondences in the two images. In order to overcome the problem of distortions, a promising approach is to locally adapt the correlation window size. In this paper we derive two major simplifications of an existing adaptive window approach. These simplifications are justified for images of objects with constant reflection properties. For this purpose the homogeneity constraint is introduced. In order to increase the robustness of the simplified algorithm the integration with a local description matching approach is proposed. The new algorithm is extended by subpixel approximation. Experiments are performed on real data and compared with manual measurements. A numerical analysis of the results demonstrates the achieved accuracy anal robustness. Stefan Scherer, Wilfried Andexer, Axel Pinz |
ICPR | 1 |