EDBT 2026 Demo / reviewers in the wild / expert
Wolfgang Minker
dblp:60/2193
· DBLP profile ↗
141ranked-venue papers
14as first author
23since 2021 · last 2026
0000-0003-4531-0662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 117 · 10 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 16 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent AgentsabstractLarge language models (LLMs) have advanced the field of artificial intelligence (AI) and are a powerful enabler for interactive systems. However, they still face challenges in long-term interactions that require adaptation towards the user as well as contextual knowledge and an understanding of the ever-changing environment. To overcome these challenges, holistic memory modeling is required to efficiently retrieve and store relevant information across sessions for accurate responses. Cognitive AI, which aims to simulate the human thought process in a computerized model, highlights interesting aspects, such as thoughts, memory mechanisms, and decision making, that can contribute towards improved memory modeling for LLMs. Inspired by these principles, we propose CAIM, a cognitive AI memory framework that models key aspects of human memory through a multi-agent architecture. Specialized LLM-based agents handle memory-related functions such as retrieval, contextual relevance evaluation, and memory maintenance. We compare CAIM against existing approaches, focusing on metrics such as retrieval accuracy, response correctness, and contextual coherence. The results demonstrate that CAIM outperforms baseline frameworks in different metrics, highlighting its context awareness and demonstrating its contribution to memory mechanisms for improving long-term human-AI interactions. Rebecca Westhäußer, Wolfgang Minker, Sebastian Zepf |
IUI | 2 |
| 2026 | MUDiC: A Dataset for Multi-User Dialogue and Collaboration in Chatbot Interaction
Nicolas Wagner 0001, Cristina Luna Jiménez, Elisabeth André, Wolfgang Minker, Stefan Ultes |
LREC | 4 |
| 2025 | Expanding the Singular Channel - MODALINK: A Generalized Automated Multimodal Dataset Generation Workflow for Emotion Recognition
Amany H. Abou El-Naga, May Hussien, Wolfgang Minker, Mohammed A.-M. Salem, Nada Sharaf |
DATA | 3 |
| 2025 | A Multilingual Telegram Chatbot for Mental Health Data Collection
Danila Mamontov, Alexey Karpov 0001, Wolfgang Minker |
ICMI | 3 |
| 2024 | Exploring the Impact of Non-Verbal Virtual Agent Behavior on User Engagement in Argumentative DialoguesabstractEngaging in discussions that involve diverse perspectives and exchanging arguments on a controversial issue is a natural way for humans to form opinions. In this process, the way arguments are presented plays a crucial role in determining how engaged users are, whether the interaction takes place solely among humans or within human-agent teams. This is of great importance as user engagement plays a crucial role in determining the success or failure of cooperative argumentative discussions. One main goal is to maintain the user’s motivation to participate in a reflective opinion-building process, even when addressing contradicting viewpoints. This work investigates how non-verbal agent behavior, specifically co-speech gestures, influences the user’s engagement and interest during an ongoing argumentative interaction. The results of a laboratory study conducted with 56 participants demonstrate that the agent’s co-speech gestures have a substantial impact on user engagement and interest and the overall perception of the system. Therefore, this research offers valuable insights for the design of future cooperative argumentative virtual agents. Annalena Aicher, Yuki Matsuda 0001, Keiichi Yasumoto, Wolfgang Minker, Elisabeth André, Stefan Ultes |
HAI | 4 |
| 2024 | Enhancing Model Transparency: A Dialogue System Approach to XAI with Domain KnowledgeabstractExplainable artificial intelligence (XAI) is a rapidly evolving field that seeks to create AI systems that can provide humanunderstandable explanations for their decisionmaking processes.However, these explanations rely on model and data-specific information only.To support better human decisionmaking, integrating domain knowledge into AI systems is expected to enhance understanding and transparency.In this paper, we present an approach for combining XAI explanations with domain knowledge within a dialogue system.We concentrate on techniques derived from the field of computational argumentation to incorporate domain knowledge and corresponding explanations into human-machine dialogue.We implement the approach in a prototype system for an initial user evaluation, where users interacted with the dialogue system to receive predictions from an underlying AI model.The participants were able to explore different types of explanations and domain knowledge.Our results indicate that users tend to more effectively evaluate model performance when domain knowledge is integrated.On the other hand, we found that domain knowledge was not frequently requested by the user during dialogue interactions. Isabel Feustel, Niklas Rach, Wolfgang Minker, Stefan Ultes |
SIGDIAL | 3 |
| 2023 | System-Initiated Transitions from Chit-Chat to Task-Oriented Dialogues with Transition Info Extractor and Transition Sentence GeneratorabstractIn this work, we study dialogue scenarios that start from chit-chat but eventually switch to task-related services, and investigate how a unified dialogue model, which can engage in both chit-chat and task-oriented dialogues, takes the initiative during the dialogue mode transition from chit-chat to task-oriented in a coherent and cooperative manner.We firstly build a transition info extractor (TIE) that keeps track of the preceding chit-chat interaction and detects the potential user intention to switch to a taskoriented service.Meanwhile, in the unified model, a transition sentence generator (TSG) is extended through efficient Adapter tuning and transition prompt learning.When the TIE successfully finds task-related information from the preceding chit-chat, such as a transition domain ("train" in Figure 1), then the TSG is activated automatically in the unified model to initiate this transition by generating a transition sentence under the guidance of transition information extracted by TIE.The experimental results show promising performance regarding the proactive transitions.We achieve an additional large improvement on TIE model by utilizing Conditional Random Fields (CRF).The TSG can flexibly generate transition sentences while maintaining the unified capabilities of normal chit-chat and task-oriented response generation. Ye Liu 0009, Stefan Ultes, Wolfgang Minker, Wolfgang Maier 0001 |
INLG | 3 |
| 2023 | The Influence of Avatar Interfaces on Argumentative DialoguesabstractHumans form opinions and justify different points of view by exchanging arguments and knowledge. Likewise to human-human interaction, the way arguments are presented influence the user's willingness to engage into a critical reflection. Especially when interacting with conversational agents the user's engagement and motivation are important factors and highly influence the success or failure of such a mixed team. To maintain the users' trust and satisfaction, the users' perception of the respective system is an important indicator. Thus, this work investigates the design of a cooperative argumentative dialogue system using a virtual avatar compared to a non-avatar interface by evaluating a crowdsourcing study conducted with 84 participants. The results indicate, that the avatar system is perceived as significantly more appealing and natural and thus, engaging which also influences the acceptance and perception of the quality of presented arguments. Furthermore, we found that the presence of the avatar often led to an increase in the anticipated level of conversational proficiency similar to that of a human interlocutor. Therefore, this work provides important insights for the design of future cooperative argumentative virtual avatar interfaces. Annalena Aicher, Klaus Weber 0001, Elisabeth André, Wolfgang Minker, Stefan Ultes |
IVA | 4 |
| 2023 | Does It Affect You? Social and Learning Implications of Using Cognitive-Affective State Recognition for Proactive Human-Robot TutoringabstractRobotic technology has proven to be advantageous for student learning and social development in educational settings. However, in order to enhance their effectiveness and provide a more human-like tutoring experience, robots must be capable of adapting to the user and exhibiting proactivity. By acting proactively, these intelligent robotic tutors can anticipate potential obstacles and take preventative measures to avoid negative outcomes. However, determining when and how to behave proactively remains an open question. This study investigates how a robotic tutor can utilize a student’s cognitive-affective states to trigger proactive tutoring dialogue and improve the learning experience. Specifically, we observed a concept learning task scenario where a robotic assistant proactively assisted the user when negative states, such as frustration and confusion, were detected. In an empirical study involving 40 undergraduate and doctoral students, we evaluated whether the initiation of proactive behavior after the detection of signs of confusion and frustration improves the student’s concentration and trust in the robot. We also examined which level of proactive dialogue is most effective for promoting concentration and trust. The results indicate that high levels of proactive behavior can harm trust, especially when triggered during negative cognitive-affective states. However, this behavior does contribute to keeping the student focused on the task when triggered during these states. Based on our findings, we discuss potential future steps for improving the proactive assistance of robotic tutoring systems. Matthias Kraus 0001, Diana Betancourt, Wolfgang Minker |
RO-MAN | 3 |
| 2023 | Towards Breaking the Self-imposed Filter Bubble in Argumentative DialoguesabstractHuman users tend to selectively ignore information that contradicts their pre-existing beliefs or opinions in their process of information seeking.These "self-imposed filter bubbles" (SFB) pose a significant challenge for cooperative argumentative dialogue systems aiming to build an unbiased opinion and a better understanding of the topic at hand.To address this issue, we develop a strategy for overcoming users' SFB within the course of the interaction.By continuously modeling the user's position in relation to the SFB, we are able to identify the respective arguments which maximize the probability to get outside the SFB and present them to the user.We implemented this approach in an argumentative dialogue system and evaluated in a laboratory user study with 60 participants to show its validity and applicability.The findings suggest that the strategy was successful in breaking users' SFBs and promoting a more reflective and comprehensive discussion of the topic. Annalena Aicher, Daniel Kornmüller, Yuki Matsuda 0001, Stefan Ultes, Wolfgang Minker, Keiichi Yasumoto |
SIGDIAL | 5 |
| 2023 | Improving Proactive Dialog Agents Using Socially-Aware Reinforcement LearningabstractThe next step for intelligent dialog agents is to escape their role as silent bystanders and become proactive. Well-defined proactive behavior may improve human-machine cooperation, as the agent takes a more active role during interaction and takes off responsibility from the user. However, proactivity is a double-edged sword because poorly executed pre-emptive actions may have a devastating effect on the task outcome and the relationship with the user. For designing adequate proactive dialog strategies, we propose a novel approach including both social and task-relevant features in the dialog. Here, the primary goal is to optimize proactive behavior so that it is task-oriented - this implies high task success and efficiency - while also being socially effective by fostering user trust. Including both aspects in the reward function for training a proactive dialog agent using reinforcement learning showed the benefit of our approach for more successful human-machine cooperation. Matthias Kraus 0001, Nicolas Wagner 0003, Ron Riekenbrauck, Wolfgang Minker |
UMAP | 4 |
| 2022 | KURT: A Household Assistance Robot Capable of Proactive DialogueabstractIn this work, we present a robot-dialogue framework to handle sophisticated robot-initiated interaction. We introduce a robotic assistant equipped with a dialogue system in a household assistance context. To become a truly collaborative companion, the assistant is able to engage in a proactive conversation for task assistance. The system actions are triggered by the recognition of persons or specific objects. To evaluate our system, we conducted a user study with 17 participants in a laboratory environment where users were able to interact with the system via natural language. The results showed that the behaviour of the system was accepted and perceived as trustworthy by the users. Matthias Kraus 0001, Nicolas Wagner 0003, Wolfgang Minker, Ankita Agrawal, Artur Schmidt, Pranav Krishna Prasad, Wolfgang Ertel |
HRI | 3 |
| 2022 | Towards Building a Spoken Dialogue System for Argument ExplorationabstractSpeech interfaces for argumentative dialogue systems (ADS) are rather scarce. The complex task they pursue hinders the application of common natural language understanding (NLU) approaches in this domain. To address this issue we include an adaption of a recently introduced NLU framework tailored to argumentative tasks into a complete ADS. We evaluate the likeability and motivation of users to interact with the new system in a user study. Therefore, we compare it to a solid baseline utilizing a drop-down menu. The results indicate that the integration of a flexible NLU framework enables a far more natural and satisfying interaction with human users in real-time. Even though the drop-down menu convinces regarding its robustness, the willingness to use the new system is significantly higher. Hence, the featured NLU framework provides a sound basis to build an intuitive interface which can be extended to adapt its behavior to the individual user. Annalena Aicher, Nadine Gerstenlauer, Isabel Feustel, Wolfgang Minker, Stefan Ultes |
LREC | 4 |
| 2022 | Towards Speech-only Opinion-level Sentiment AnalysisabstractThe growing popularity of various forms of Spoken Dialogue Systems (SDS) raises the demand for their capability of implicitly assessing the speaker’s sentiment from speech only. Mapping the latter on user preferences enables to adapt to the user and individualize the requested information while increasing user satisfaction. In this paper, we explore the integration of rank consistent ordinal regression into a speech-only sentiment prediction task performed by ResNet-like systems. Furthermore, we use speaker verification extractors trained on larger datasets as low-level feature extractors. An improvement of performance is shown by fusing sentiment and pre-extracted speaker embeddings reducing the speaker bias of sentiment predictions. Numerous experiments on Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) databases show that we beat the baselines of state-of-the-art unimodal approaches. Using speech as the only modality combined with optimizing an order-sensitive objective function gets significantly closer to the sentiment analysis results of state-of-the-art multimodal systems. Annalena Aicher, Alisa Gazizullina, Aleksei Gusev, Yuri Matveev, Wolfgang Minker |
LREC | 5 |
| 2022 | User Interest Modelling in Argumentative Dialogue SystemsabstractMost systems helping to provide structured information and support opinion building, discuss with users without considering their individual interest. The scarce existing research on user interest in dialogue systems depends on explicit user feedback. Such systems require user responses that are not content-related and thus, tend to disturb the dialogue flow. In this paper, we present a novel model for implicitly estimating user interest during argumentative dialogues based on semantically clustered data. Therefore, an online user study was conducted to acquire training data which was used to train a binary neural network classifier in order to predict whether or not users are still interested in the content of the ongoing dialogue. We achieved a classification accuracy of 74.9% and furthermore investigated with different Artificial Neural Networks (ANN) which new argument would fit the user interest best. Annalena Aicher, Nadine Gerstenlauer, Wolfgang Minker, Stefan Ultes |
LREC | 3 |
| 2022 | Towards Modelling Self-imposed Filter Bubbles in Argumentative Dialogue SystemsabstractTo build a well-founded opinion it is natural for humans to gather and exchange new arguments. Especially when being confronted with an overwhelming amount of information, people tend to focus on only the part of the available information that fits into their current beliefs or convenient opinions. To overcome this “self-imposed filter bubble” (SFB) in the information seeking process, it is crucial to identify influential indicators for the former. Within this paper we propose and investigate indicators for the the user’s SFB, mainly their Reflective User Engagement (RUE), their Personal Relevance (PR) ranking of content-related subtopics as well as their False (FK) and True Knowledge (TK) on the topic. Therefore, we analysed the answers of 202 participants of an online conducted user study, who interacted with our argumentative dialogue system BEA (“Building Engaging Argumentation”). Moreover, also the influence of different input/output modalities (speech/speech and drop-down menu/text) on the interaction with regard to the suggested indicators was investigated. Annalena Aicher, Wolfgang Minker, Stefan Ultes |
LREC | 2 |
| 2022 | ProDial - An Annotated Proactive Dialogue Act Corpus for Conversational Assistants using CrowdsourcingabstractRobots will eventually enter our daily lives and assist with a variety of tasks. Especially in the household domain, robots may become indispensable helpers by overtaking tedious tasks, e.g. keeping the place tidy. Their effectiveness and efficiency, however, depend on their ability to adapt to our needs, routines, and personal characteristics. Otherwise, they may not be accepted and trusted in our private domain. For enabling adaptation, the interaction between a human and a robot needs to be personalized. Therefore, the robot needs to collect personal information from the user. However, it is unclear how such sensitive data can be collected in an understandable way without losing a user’s trust in the system. In this paper, we present a conversational approach for explicitly collecting personal user information using natural dialogue. For creating a sound interactive personalization, we have developed an empathy-augmented dialogue strategy. In an online study, the empathy-augmented strategy was compared to a baseline dialogue strategy for interactive personalization. We have found the empathy-augmented strategy to perform notably friendlier. Overall, using dialogue for interactive personalization has generally shown positive user reception. Matthias Kraus 0001, Nicolas Wagner 0003, Wolfgang Minker |
LREC | 3 |
| 2022 | Including Social Expectations for Trustworthy Proactive Human-Robot DialogueabstractTrust forms an important factor in human-robot interaction and is highly influencing the success or failure of a mixed team of humans and machines. Similarly, to human-human teamwork, communication and proactivity are one of the keys to task success and efficiency. However, the level of proactive robot behaviour needs to be adapted to a dynamically changing social environment. Otherwise, it may be perceived as counterproductive and the robot’s assistance may not be accepted. For this reason, this work investigates the design of a socially-adaptive proactive dialogue strategy and its effects on humans’ trust and acceptance towards the robot. The strategy is implemented in a human-like household assistance robot that helps in the execution of domestic tasks, such as tidying up or fetch-and-carry tasks. For evaluation of the strategy, users interact with the robot while watching interactive videos of the robots in six different task scenarios. Here, the adaptive proactive behaviour of the robot is compared to four different levels of static proactivity: None, Notification, Suggestion, and Intervention. The results show that proactive robot behaviour that adapts to the social expectations of a user has a significant effect on the perceived trust in the system. Here, it is shown that a robot expressing socially-adaptive proactivity is perceived as more competent and reliable than a non-adaptive robot. Based on these results, important implications for the design of future robotic assistants at home are described. Matthias Kraus 0001, Nicolas Wagner 0003, Nico Untereiner, Wolfgang Minker |
UMAP | 4 |
| 2022 | Natural language understanding for argumentative dialogue systems in the opinion building domain
Waheed Ahmed Abro, Annalena Aicher, Niklas Rach, Stefan Ultes, Wolfgang Minker, Guilin Qi |
Knowl. Based Syst. | 5 |
| 2021 | Modelling and Predicting Trust for Developing Proactive Dialogue Strategies in Mixed-Initiative InteractionabstractIn mixed-initiative user interactions, a user and an autonomous agent collaborate for solving tasks by taking interleaving actions. However, this shift of control towards the agent requires a formation of trust for the user, otherwise the assistance possibly will be rejected and becomes obsolete. One approach for fostering a trustworthy interaction is to equip an agent with proactive dialogue capabilities. However, the development of adequate proactive dialogue strategies is complex and highly user- as well as context-dependent. Inappropriate usage of proactive conversation may even do more harm than good and corrupt the human-computer trust relationship. In order to alleviate this problem, modelling and predicting a proactive system’s perceived trustworthiness during an ongoing interaction is essential. Therefore, this paper presents novel work on the development of a user model for live prediction of trust during proactive interaction, incorporating user-, system-, and context-dependent features. For predicting trust, three machine-learning algorithms – support vector machine, eXtreme Gradient Boost, gated recurrent unit network – are trained and tested on a proactive dialogue corpus. The experimental results show that among the classifiers the support vector machine showed the most well-rounded performance, while the gated recurrent unit had the best accuracy. The results prove the developed user model to be reliable for predicting trust in proactive dialogue. Based on the outcomes, the usability of the proposed method in real-life scenarios is discussed and implications for developing user-adaptive proactive dialogue strategies are described. Matthias Kraus 0001, Nicolas Wagner 0003, Wolfgang Minker |
ICMI | 3 |
| 2021 | Ensemble-Within-Ensemble Classification for Escalation Prediction from Speech
Oxana Verkholyak, Denis Dresvyanskiy, Anastasia Dvoynikova, Denis Kotov, Elena Ryumina, Alena Velichko, Danila Mamontov, Wolfgang Minker, Alexey Karpov 0001 |
Interspeech | 8 |
| 2021 | Towards the Development of a Trustworthy Chatbot for Mental Health Applications
Matthias Kraus 0001, Philip Seldschopf, Wolfgang Minker |
MMM (2) | 3 |
| 2021 | From Argument Search to Argumentative Dialogue: A Topic-independent Approach to Argument Acquisition for Dialogue SystemsabstractNiklas Rach, Carolin Schindler, Isabel Feustel, Johannes Daxenberger, Wolfgang Minker, Stefan Ultes. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Niklas Rach, Carolin Schindler, Isabel Feustel, Johannes Daxenberger, Wolfgang Minker, Stefan Ultes |
SIGDIAL | 5 |
| 2020 | Increasing the Naturalness of an Argumentative Dialogue System Through Argument ChainsabstractThis work introduces chained arguments into a dialogue game for argumentation to allow a more natural and intuitive interaction with a respective system. Thus, the turn taking rules of the game are improved while still preserving the general consistency that is ensured by the framework. The improved system is used to generate artificial dialogues between two virtual agents which are assessed in a user study. The results show a significant improvement in the perceived naturalness without violating the logical consistency. Niklas Rach, Wolfgang Minker, Stefan Ultes |
COMMA | 2 |
| 2020 | "Was that successful?" On Integrating Proactive Meta-Dialogue in a DIY-Assistant using Multimodal CuesabstractEffectively supporting novices during performance of complex tasks, e.g. do-it-yourself (DIY) projects, requires intelligent assistants to be more than mere instructors. In order to be accepted as a competent and trustworthy cooperation partner, they need to be able to actively participate in the project and engage in helpful conversations with users when assistance is necessary. Therefore, a new proactive version of the DIY-assistant Robert is presented in this paper. It extends the previous prototype by including the capability to initiate reflective meta-dialogues using multimodal cues. Two different strategies for reflective dialogue are implemented: A progress-based strategy initiates a reflective dialogue about previous experience with the assistance for encouraging the self-appraisal of the user. An activity-based strategy is applied for providing timely, task-dependent support. Therefore, user activities with a connected drill driver are tracked that trigger dialogues in order to reflect on the current task and to prevent task failure. An experimental study comparing the proactive assistant against the baseline version shows that proactive meta-dialogue is able to build user trust significantly better than a solely reactive system. Besides, the results provide interesting insights for the development of proactive dialogue assistants. Matthias Kraus 0001, Marvin R. G. Schiller, Gregor Behnke, Pascal Bercher, Michael Dorna, Michael Dambier, Birte Glimm, Susanne Biundo-Stephan, Wolfgang Minker |
ICMI | 9 |
| 2020 | Ensembling End-to-End Deep Models for Computational Paralinguistics Tasks: ComParE 2020 Mask and Breathing Sub-ChallengesabstractThis paper describes deep learning approaches for the Mask and Breathing Sub-Challenges (SCs), which are addressed by the INTERSPEECH 2020 Computational Paralinguistics Challenge. Motivated by outstanding performance of state-of-the-art end-to-end (E2E) approaches, we explore and compare effectiveness of different deep Convolutional Neural Network (CNN) architectures on raw data, log Mel-spectrograms, and Mel-Frequency Cepstral Coefficients. We apply a transfer learning approach to improve model’s efficiency and convergence speed. In the Mask SC, we conduct experiments with several pretrained CNN architectures on log-Mel spectrograms, as well as Support Vector Machines on baseline features. For the Breathing SC, we propose an ensemble deep learning system that exploits E2E learning and sequence prediction. The E2E model is based on 1D CNN operating on raw speech signals and is coupled with Long Short-Term Memory layers for sequence modeling. The second model works with log-Mel features and is based on a pretrained 2D CNN model stacked to Gated Recurrent Unit layers. To increase performance of our models in both SCs, we use ensembles of the best deep neural models obtained from N-fold cross-validation on combined challenge training and development datasets. Our results markedly outperform the challenge test set baselines in both SCs. Maxim Markitantov, Denis Dresvyanskiy, Danila Mamontov, Heysem Kaya, Wolfgang Minker, Alexey Karpov 0001 |
INTERSPEECH | 5 |
| 2020 | A Comparison of Explicit and Implicit Proactive Dialogue Strategies for Conversational RecommendationabstractRecommendation systems aim at facilitating information retrieval for users by taking into account their preferences. Based on previous user behaviour, such a system suggests items or provides information that a user might like or find useful. Nonetheless, how to provide suggestions is still an open question. Depending on the way a recommendation is communicated influences the user’s perception of the system. This paper presents an empirical study on the effects of proactive dialogue strategies on user acceptance. Therefore, an explicit strategy based on user preferences provided directly by the user, and an implicit proactive strategy, using autonomously gathered information, are compared. The results show that proactive dialogue systems significantly affect the perception of human-computer interaction. Although no significant differences are found between implicit and explicit strategies, proactivity significantly influences the user experience compared to reactive system behaviour. The study contributes new insights to the human-agent interaction and the voice user interface design. Furthermore, we discover interesting tendencies that motivate futurework. Matthias Kraus 0001, Fabian Fischbach, Pascal Jansen, Wolfgang Minker |
LREC | 4 |
| 2020 | Estimating User Communication Styles for Spoken Dialogue SystemsabstractWe present a neural network approach to estimate the communication style of spoken interaction, namely the stylistic variations elaborateness and directness, and investigate which type of input features to the estimator are necessary to achive good performance. First, we describe our annotated corpus of recordings in the health care domain and analyse the corpus statistics in terms of agreement, correlation and reliability of the ratings. We use this corpus to estimate the elaborateness and the directness of each utterance. We test different feature sets consisting of dialogue act features, grammatical features and linguistic features as input for our classifier and perform classification in two and three classes. Our classifiers use only features that can be automatically derived during an ongoing interaction in any spoken dialogue system without any prior annotation. Our results show that the elaborateness can be classified by only using the dialogue act and the amount of words contained in the corresponding utterance. The directness is a more difficult classification task and additional linguistic features in form of word embeddings improve the classification results. Afterwards, we run a comparison with a support vector machine and a recurrent neural network classifier. Juliana Miehle, Isabel Feustel, Julia Hornauer, Wolfgang Minker, Stefan Ultes |
LREC | 4 |
| 2020 | Comparative Study of Sentence Embeddings for Contextual ParaphrasingabstractParaphrasing is an important aspect of natural-language generation that can produce more variety in the way specific content is presented. Traditionally, paraphrasing has been focused on finding different words that convey the same meaning. However, in human-human interaction, we regularly express our intention with phrases that are vastly different regarding both word content and syntactic structure. Instead of exchanging only individual words, the complete surface realisation of a sentences is altered while still preserving its meaning and function in a conversation. This kind of contextual paraphrasing did not yet receive a lot of attention from the scientific community despite its potential for the creation of more varied dialogues. In this work, we evaluate several existing approaches to sentence encoding with regard to their ability to capture such context-dependent paraphrasing. To this end, we define a paraphrase classification task that incorporates contextual paraphrases, perform dialogue act clustering, and determine the performance of the sentence embeddings in a sentence swapping task. Louisa Pragst, Wolfgang Minker, Stefan Ultes |
LREC | 2 |
| 2020 | Evaluation of Argument Search Approaches in the Context of Argumentative Dialogue SystemsabstractWe present an approach to evaluate argument search techniques in view of their use in argumentative dialogue systems by assessing quality aspects of the retrieved arguments. To this end, we introduce a dialogue system that presents arguments by means of a virtual avatar and synthetic speech to users and allows them to rate the presented content in four different categories (Interesting, Convincing, Comprehensible, Relation). The approach is applied in a user study in order to compare two state of the art argument search engines to each other and with a system based on traditional web search. The results show a significant advantage of the two search engines over the baseline. Moreover, the two search engines show significant advantages over each other in different categories, thereby reflecting strengths and weaknesses of the different underlying techniques. Niklas Rach, Yuki Matsuda 0001, Johannes Daxenberger, Stefan Ultes, Keiichi Yasumoto, Wolfgang Minker |
LREC | 6 |
| 2020 | How Users React to Proactive Voice Assistant Behavior While DrivingabstractNowadays Personal Assistants (PAs) are available in multiple environments and become increasingly popular to use via voice. Therefore, we aim to provide proactive PA suggestions to car drivers via speech. These suggestions should be neither obtrusive nor increase the drivers’ cognitive load, while enhancing user experience. To assess these factors, we conducted a usability study in which 42 participants perceive proactive voice output in a Wizard-of-Oz study in a driving simulator. Traffic density was varied during a highway drive and it included six in-car-specific use cases. The latter were presented by a proactive voice assistant and in a non-proactive control condition. We assessed the users’ subjective cognitive load and their satisfaction in different questionnaires during the interaction with both PA variants. Furthermore, we analyze the user reactions: both regarding their content and the elapsed response times to PA actions. The results show that proactive assistant behavior is rated similarly positive as non-proactive behavior. Furthermore, the participants agreed to 73.8% of proactive suggestions. In line with previous research, driving-relevant use cases receive the best ratings, here we reach 82.5% acceptance. Finally, the users reacted significantly faster to proactive PA actions, which we interpret as less cognitive load compared to non-proactive behavior. Maria Schmidt, Wolfgang Minker, Steffen Werner |
LREC | 2 |
| 2020 | Design and evaluation on task allocation interfaces in gamified participatory sensing for tourismabstractIn the tourism sector, user-generated information and communication among tourists are perceived to be more effective and reliable contents. In addition, the collection of dynamic tourism information with high spatio-temporal resolution is required to provide comfortable tourism in response to the changing tourism style with the advancement of information technology. Participatory sensing, which can collect various types of information, is a useful method by which to collect these contents. However, continuous participation of users is essential in participatory sensing, and it is one of the most important points to stimulate participation motivation. In the tourism situation, we also need to pay attention to the total tourist satisfaction of participants. In this study, we investigate the effects of task allocation interfaces and user types on the efficiency of tourism information collection, tourism behavior and satisfaction in gamified participatory sensing. Two types of task allocation interfaces (free selection and agent interaction) were designed and implemented, and a sightseeing experiment was conducted with 10 participants at an actual sightseeing spot (Nara, Japan). As a result, we found that there was no difference in the effect of each interface on sightseeing satisfaction, but the characteristics of the collected data that, free selection allows for the collection of quantitative data and agent interaction allows for the efficient collection of data needed by the system, were different. In addition, we found that different user types had different tendencies for their contribution to sensing and their interface preferences. Shogo Kawanaka, Juliana Miehle, Yuki Matsuda 0001, Hirohiko Suwa, Keiichi Yasumoto, Wolfgang Minker |
MobiQuitous | 6 |
| 2020 | Effects of Proactive Dialogue Strategies on Human-Computer TrustabstractIntelligent computer systems aim at providing user-assistance for challenging tasks, like decision-making, planning, or learning. For offering optimal assistance, it is essential for such systems to decide when to be reactive or proactive and how active system behaviour should be designed. Especially, as this decision may greatly influence the user's trust in the system. Therefore, we conducted a mixed-factorial study which examines how different levels of proactivity (none, notification, suggestion, and intervention) as well as timing strategies (fixed-timing and insecurity-based) are trusted by subjects while performing a planning task. The results showed, that proactive system behaviour is perceived trustworthy in insecure situations independent of the timing. However, proactive dialogue showed strong effects on cognition-based trust (system's perceived competence and reliability) depending on task difficulty. Furthermore, fully autonomous system behaviour fails to establish an adequate human-computer trust relationship, in contrast to conservative strategies. Matthias Kraus 0001, Nicolas Wagner 0003, Wolfgang Minker |
UMAP | 3 |
| 2019 | A Test Collection for Passage Retrieval Evaluation of Spanish Health-Related Resources
Eleni Kamateri, Theodora Tsikrika, Spyridon Symeonidis, Stefanos Vrochidis, Wolfgang Minker, Ioannis Kompatsiaris |
ECIR (2) | 5 |
| 2019 | Cross-Corpus Data Augmentation for Acoustic Addressee DetectionabstractAcoustic addressee detection (AD) is a modern paralinguistic and dialogue challenge that especially arises in voice assistants.In the present study, we distinguish addressees in two settings (a conversation between several people and a spoken dialogue system, and a conversation between several adults and a child) and introduce the first competitive baseline (unweighted average recall equals 0.891) for the Voice Assistant Conversation Corpus that models the first setting.We jointly solve both classification problems, using three models: a linear support vector machine dealing with acoustic functionals and two neural networks utilising raw waveforms alongside with acoustic low-level descriptors.We investigate how different corpora influence each other, applying the mixup approach to data augmentation.We also study the influence of various acoustic context lengths on AD.Two-second speech fragments turn out to be sufficient for reliable AD.Mixup is shown to be beneficial for merging acoustic data (extracted features but not raw waveforms) from different domains that allows us to reach a higher classification performance on human-machine AD and also for training a multipurpose neural network that is capable of solving both human-machine and adult-child AD problems. Oleg Akhtiamov, Ingo Siegert, Alexey Karpov 0001, Wolfgang Minker |
SIGdial | 4 |
| 2018 | Markov Games for Persuasive DialogueabstractThis work discusses the formulation of argumentative dialogue as Markov game. We show how formal systems for persuasive dialogues that adhere to a certain structure can be reformulated as Markov games and thus be addressed as Reinforcement Learning task in a multi-agent setting. We validate our approach on an implementation of a proof of principle scenario where we show that the optimal policy can be learned. Niklas Rach, Wolfgang Minker, Stefan Ultes |
COMMA | 2 |
| 2018 | EVA: A Multimodal Argumentative Dialogue SystemabstractThis work introduces EVA, a multimodal argumentative Dialogue System that is capable of discussing controversial topics with the user. The interaction is structured as an argument game in which the user and the system select respective moves in order to convince their opponent. EVA's response is presented as a natural language utterance by a virtual agent that supports the respective content using characteristic gestures and mimic. Niklas Rach, Klaus Weber 0001, Louisa Pragst, Elisabeth André, Wolfgang Minker, Stefan Ultes |
ICMI | 5 |
| 2018 | Instructing Novice Users on How to Use Tools in DIY ProjectsabstractNovice users require assistance when performing handicraft tasks. Adequate instruction ensures task completion and conveys knowledge and abilities required to perform the task. We present an assistant teaching novice users how to operate electronic tools, such as drills, saws, and sanders, in the context of Do-It-Yourself (DIY) home improvement projects. First, the actions that need to be performed for the project are determined by a planner. Second, a dialogue manager capable of natural language interaction presents these actions as instructions to the user. Third, questions on these actions and involved objects are answered by generating appropriate ontology-based explanations. Gregor Behnke, Marvin R. G. Schiller, Matthias Kraus 0001, Pascal Bercher, Mario Schmautz, Michael Dorna, Wolfgang Minker, Birte Glimm, Susanne Biundo-Stephan |
IJCAI | 7 |
| 2018 | Contextual Dependencies in Time-Continuous Multidimensional Affect Recognition
Dmitrii Fedotov, Denis Ivanko, Maxim Sidorov, Wolfgang Minker |
LREC | 4 |
| 2018 | Effects of Gender Stereotypes on Trust and Likability in Spoken Human-Robot Interaction
Matthias Kraus 0001, Johannes Kraus 0002, Martin Baumann 0001, Wolfgang Minker |
LREC | 4 |
| 2018 | Expert Evaluation of a Spoken Dialogue System in a Clinical Operating Room
Juliana Miehle, Nadine Gerstenlauer, Daniel Ostler, Hubertus Feußner, Wolfgang Minker, Stefan Ultes |
LREC | 5 |
| 2018 | What Causes the Differences in Communication Styles? A Multicultural Study on Directness and Elaborateness
Juliana Miehle, Wolfgang Minker, Stefan Ultes |
LREC | 2 |
| 2018 | On the Vector Representation of Utterances in Dialogue Context
Louisa Pragst, Niklas Rach, Wolfgang Minker, Stefan Ultes |
LREC | 3 |
| 2018 | Towards Estimating Emotions and Satisfaction Level of Tourist Based on Eye Gaze and Head MovementabstractFollowing the increase in demand for "smart tourism," various tourist information becomes available. Current tourist guidance systems can provide many possible routes around the city, which are usually based on an optimal distance and time or popularity of the place. To design more enjoyable tourist routes, we should be aware of tourist's perception of the urban environment. In this paper, we propose the methodology of estimating the tourist emotions and satisfaction level by analysing physiological features, such as head movement and eye gaze. Features derived from raw sensor data have the correlation up to 0.58 with emotion and satisfaction labels. This study also shows the differences in feature/label dependencies between different touristic areas and highlights the challenges of tourist satisfaction estimation. Dmitrii Fedotov, Yuki Matsuda 0001, Yutaka Arakawa, Keiichi Yasumoto, Wolfgang Minker |
SMARTCOMP | 6 |
| 2017 | Acquisition and Assessment of Semantic Content for the Generation of Elaborateness and Indirectness in Spoken Dialogue SystemsabstractIn a dialogue system, the dialogue manager selects one of several system actions and thereby determines the system’s behaviour. Defining all possible system actions in a dialogue system by hand is a tedious work. While efforts have been made to automatically generate such system actions, those approaches are mostly focused on providing functional system behaviour. Adapting the system behaviour to the user becomes a difficult task due to the limited amount of system actions available. We aim to increase the adaptability of a dialogue system by automatically generating variants of system actions. In this work, we introduce an approach to automatically generate action variants for elaborateness and indirectness. Our proposed algorithm extracts RDF triplets from a knowledge base and rates their relevance to the original system action to find suitable content. We show that the results of our algorithm are mostly perceived similarly to human generated elaborateness and indirectness and can be used to adapt a conversation to the current user and situation. We also discuss where the results of our algorithm are still lacking and how this could be improved: Taking into account the conversation topic as well as the culture of the user is likely to have beneficial effect on the user’s perception. Louisa Pragst, Koichiro Yoshino, Wolfgang Minker, Satoshi Nakamura 0001, Stefan Ultes |
IJCNLP(1) | 3 |
| 2017 | Speech and Text Analysis for Multimodal Addressee Detection in Human-Human-Computer Interaction
Oleg Akhtiamov, Maxim Sidorov, Alexey Karpov 0001, Wolfgang Minker |
INTERSPEECH | 4 |
| 2017 | Interaction Quality Estimation Using Long Short-Term MemoriesabstractFor estimating the Interaction Quality (IQ) in Spoken Dialogue Systems (SDS), the dialogue history is of significant importance.Previous works included this information manually in the form of precomputed temporal features into the classification process.Here, we employ a deep learning architecture based on Long Short-Term Memories (LSTM) to extract this information automatically from the data, thus estimating IQ solely by using current exchange features.We show that it is thereby possible to achieve competitive results as in a scenario where manually optimized temporal features have been included. Niklas Rach, Wolfgang Minker, Stefan Ultes |
SIGDIAL Conference | 2 |
| 2016 | An Approach to Off-talk Detection based on Text Classification within an Automatic Spoken Dialogue SystemabstractThis paper describes the problem of the off-talk detection within an automatic spoken dialogue system. The considered corpus contains realistic conversations between two users and an SDS. A two- (on-talk and off-talk) and a three-class (on-talk, problem-related off-talk, and irrelevant off-talk) problem statement are investigated using a speaker-independent approach to cross-validation. A novel off-talk detection approach based on text classification is proposed. Seven different term weighting methods and two classification algorithms are considered. As a dimensionality reduction method, a feature transformation based on term belonging to classes is applied. The comparative analysis of the proposed approach and a baseline one is performed; as a result, the best combinations of the text pre-processing methods and classification algorithms are defined for both problem statements. The novel approach demonstrates significantly better classification effectiveness in comparison with the baseline for the same task. Oleg Akhtiamov, Roman B. Sergienko, Wolfgang Minker |
ICINCO (2) | 3 |
| 2016 | A Comparative Study of Text Preprocessing Approaches for Topic Detection of User Utterances
Roman B. Sergienko, Muhammad Shan, Wolfgang Minker |
LREC | 3 |
| 2016 | Could Speaker, Gender or Age Awareness be beneficial in Speech-based Emotion Recognition?
Maxim Sidorov, Alexander Schmitt, Eugene Semenkin, Wolfgang Minker |
LREC | 4 |
| 2016 | Cultural Communication Idiosyncrasies in Human-Computer InteractionabstractIn this work, we investigate whether the cultural idiosyncrasies found in humanhuman interaction may be transferred to human-computer interaction.With the aim of designing a culture-sensitive dialogue system, we designed a user study creating a dialogue in a domain that has the potential capacity to reveal cultural differences.The dialogue contains different options for the system output according to cultural differences.We conducted a survey among Germans and Japanese to investigate whether the supposed differences may be applied in human-computer interaction.Our results show that there are indeed differences, but not all results are consistent with the cultural models. Juliana Miehle, Koichiro Yoshino, Louisa Pragst, Stefan Ultes, Satoshi Nakamura 0001, Wolfgang Minker |
SIGDIAL Conference | 6 |
| 2015 | A Planning-Based Assistance System for Setting Up a Home TheaterabstractModern technical devices are often too complex for many users to be able to use them to their full extent. Based on planning technology, we are able to provide advanced user assistance for operating technical devices. We present a system that assists a human user in setting up a complex home theater consisting of several HiFi devices. For a human user, the task is rather challenging due to a large number of different ports of the devices and the variety of available cables. The system supports the user by giving detailed instructions how to assemble the theater. Its performance is based on advanced user-centered planning capabilities including the generation, repair, and explanation of plans. Pascal Bercher, Felix Richter 0001, Thilo Hoernle, Thomas Geier, Daniel Höller, Gregor Behnke, Florian Nothdurft, Frank Honold, Wolfgang Minker, Michael Weber 0001, Susanne Biundo-Stephan |
AAAI | 9 |
| 2015 | Feature and Decision Level Audio-visual Data Fusion in Emotion Recognition ProblemabstractThe speech-based emotion recognition problem has already been investigated by many authors, and reasonable results have been achieved. This article focuses on applying audio-visual data fusion approach to emotion recognition. Two state-of-the-art classification algorithms were applied to one audio and three visual feature datasets. Feature level data fusion was applied to build a multimodal emotion classification system, which helped increase emotion classification accuracy by 4% compared to the best accuracy achieved by unimodal systems. The class precisions achieved by applying algorithms on unimodal and multimodal datasets helped to reveal that different data-classifier combinations are good at recognizing certain emotions. These data-classifier combinations were fused on the decision level using several approaches, which still helped increase the accuracy by 3% compared to the best accuracy achieved by feature level fusion. Maxim Sidorov, Evgenii Sopov, Ilia Ivanov, Wolfgang Minker |
ICINCO (2) | 4 |
| 2015 | Unconstrained Global Optimization: A Benchmark Comparison of Population-based AlgorithmsabstractIn this paper we provide a systematic comparison of the following population-based optimization techniques: Genetic Algorithm (GA), Evolution Strategy (ES), Cuckoo Search (CS), Differential Evolution (DE), and Particle Swarm Optimization (PSO). The considered techniques have been implemented and evaluated on a set of 67 multivariate functions. We carefully selected the tested optimization functions which have different features and gave exactly the same number of objective function evaluations for all of the algorithms. This study has revealed that the DE algorithm is preferable in the majority of cases of the tested functions. The results of numerical evaluations and parameter optimization are presented in this paper. Maxim Sidorov, Eugene Semenkin, Wolfgang Minker |
ICINCO (1) | 3 |
| 2015 | The Interplay of User-Centered Dialog Systems and AI PlanningabstractTechnical systems evolve from simple dedicated task solvers to cooperative and competent assistants, helping the user with increasingly complex and demanding tasks.For this, they may proactively take over some of the users responsibilities and help to find or reach a solution for the user's task at hand, using e.g., Artificial Intelligence (AI) Planning techniques.However, this intertwining of user-centered dialog and AI planning systems, often called mixed-initiative planning (MIP), does not only facilitate more intelligent and competent systems, but does also raise new questions related to the alignment of AI and human problem solving.In this paper, we describe our approach on integrating AI Planning techniques into a dialog system, explain reasons and effects of arising problems, and provide at the same time our solutions resulting in a coherent, userfriendly and efficient mixed-initiative system.Finally, we evaluate our MIP system and provide remarks on the use of explanations in MIP-related phenomena. Florian Nothdurft, Gregor Behnke, Pascal Bercher, Susanne Biundo-Stephan, Wolfgang Minker |
SIGDIAL Conference | 5 |
| 2015 | Quality-adaptive Spoken Dialogue Initiative Selection And Implications On Reward ModellingabstractAdapting Spoken Dialogue Systems to the user is supposed to result in more efficient and successful dialogues.In this work, we present an evaluation of a quality-adaptive strategy with a user simulator adapting the dialogue initiative dynamically during the ongoing interaction and show that it outperforms conventional non-adaptive strategies and a random strategy.Furthermore, we indicate a correlation between Interaction Quality and dialogue completion rate, task success rate, and average dialogue length.Finally, we analyze the correlation between task success and interaction quality in more detail identifying the usefulness of interaction quality for modelling the reward of reinforcement learning strategy optimization. Stefan Ultes, Matthias Kraus 0001, Alexander Schmitt, Wolfgang Minker |
SIGDIAL Conference | 4 |
| 2015 | Fusion paradigms in cognitive technical systems for human-computer interactionabstractRecent trends in human–computer interaction (HCI) show a development towards cognitive technical systems (CTS) to provide natural and efficient operating principles. To do so, a CTS has to rely on data from multiple sensors which must be processed and combined by fusion algorithms. Furthermore, additional sources of knowledge have to be integrated, to put the observations made into the correct context. Research in this field often focuses on optimizing the performance of the individual algorithms, rather than reflecting the requirements of CTS. This paper presents the information fusion principles in CTS architectures we developed for Companion Technologies. Combination of information generally goes along with the level of abstractness, time granularity and robustness, such that large CTS architectures must perform fusion gradually on different levels — starting from sensor-based recognitions to highly abstract logical inferences. In our CTS application we sectioned information fusion approaches into three categories: perception-level fusion, knowledge-based fusion and application-level fusion. For each category, we introduce examples of characteristic algorithms. In addition, we provide a detailed protocol on the implementation performed in order to study the interplay of the developed algorithms. Michael Glodek, Frank Honold, Thomas Geier, Gerald Krell, Florian Nothdurft, Stephan Reuter, Felix Schüssel, Thilo Hoernle, Klaus Dietmayer, Wolfgang Minker, Susanne Biundo-Stephan, Michael Weber 0001, Günther Palm, Friedhelm Schwenker |
Neurocomputing | 10 |
| 2014 | Dimension Reduction with Coevolutionary Genetic Algorithm for Text ClassificationabstractText classification of large-size corpora is time-consuming for implementation of classification algorithms. For this reason, it is important to reduce dimension of text classification problems. We propose a method for dimension reduction based on hierarchical agglomerative clustering of terms and cluster weight optimization using cooperative coevolutionary genetic algorithm. The method was applied on 5 different corpora using several classification methods with different text preprocessing. The method reduces dimension of text classification problem significantly. Classification efficiency increases or decreases non-significantly after clustering with optimization of cluster weights. Tatiana Gasanova, Roman B. Sergienko, Eugene Semenkin, Wolfgang Minker |
ICINCO (1) | 4 |
| 2014 | Text Categorization Methods Application for Natural Language Call RoutingabstractNatural language call routing can be treated as an instance of topic categorization of documents after speech recognition of calls. This categorization consists of two important parts. The first one is text preprocessing for numerical data extraction and the second one is classification with machine learning methods. This paper focuses on different text preprocessing methods applied for call routing. Different machine learning algorithms with several text representations have been applied for this problem. A novel text preprocessing technique has been applied and investigated. Numerical experiments have shown computational and classification effectiveness of the proposed method in comparison with standard techniques. Also a novel features selection method was proposed. The novel features selection method has demonstrated some advantages in comparison with standard techniques. Roman B. Sergienko, Tatiana Gasanova, Eugene Semenkin, Wolfgang Minker |
ICINCO (2) | 4 |
| 2014 | Speaker State Recognition with Neural Network-based Classification and Self-adaptive Heuristic Feature SelectionabstractWhile the implementation of existing feature sets and methods for automatic speaker state analysis has already achieved reasonable results, there is still much to be done for further improvement. In our research, we tried to carry out speech analysis with the self-adaptive multi-objective genetic algorithm as a feature selection technique and with a neural network as a classifier. The proposed approach was evaluated using a number of multi-language speech databases (English, German and Japanese). According to the obtained results, the developed technique allows an increase in emotion recognition performance by up to 6.2% relative improvement in average F-measure, up to 112.0% for the speaker identification task and up to 6.4% for the speech-based gender recognition, having approximately half as many features. Maxim Sidorov, Christina Brester, Eugene Semenkin, Wolfgang Minker |
ICINCO (1) | 4 |
| 2014 | Multi-agent Cooperative Algorithms of Global OptimizationabstractIn this paper we present multi-agent cooperative algorithms of global optimization based on a genetic algorithm, an evolution strategy and particle swarm optimization. Island and co-evolution approaches have been selected as a main scheme of cooperation. The proposed techniques have been implemented and evaluated on a set of 22 multivariate functions. We assert that the proposed techniques could achieve much higher results in terms of reliability and speed criteria than the performance of corresponding conventional algorithms (without cooperative schemes) with average parameters on 18 functions from the 22 selected for the evaluation procedure. Such advantages are much more observable with increasing dimensionality of functions. Furthermore, the performance of the suggested algorithms was even higher than the performance of conventional algorithms with the best parameters for 5 functions. Maxim Sidorov, Eugene Semenkin, Wolfgang Minker |
ICINCO (1) | 3 |
| 2014 | Emotion Recognition in Real-world Conditions with Acoustic and Visual FeaturesabstractThere is an enormous number of potential applications of the system which is capable to recognize human emotions. Such opportunity can be useful in various applications, e.g., improvement of Spoken Dialogue Systems (SDSs) or monitoring agents in call-centers. Therefore, the Emotion Recognition In The Wild Challenge 2014 (EmotiW 2014) is focused on estimating emotions in real-world situations. This study presents the results of multimodal emotion recognition based on support vector classifier. The described approach results in 41.77% of overall classification accuracy in the multimodal case. The obtained result is more than 17% higher than the baseline result for multimodal approach. Maxim Sidorov, Wolfgang Minker |
ICMI | 2 |
| 2014 | Companion-Technology: Towards User- and Situation-Adaptive Functionality of Technical SystemsabstractThe properties of multimodality, individuality, adaptability, availability, cooperativeness and trustworthiness are at the focus of the investigation of Companion Systems. In this article, we describe the involved key components of such a system and the way they interact with each other. Along with the article comes a video, in which we demonstrate a fully functional prototypical implementation and explain the involved scientific contributions in a simplified manner. The realized technology considers the entire situation of the user and the environment in current and past states. The gained knowledge reflects the context of use and serves as basis for decision-making in the presented adaptive system. Frank Honold, Pascal Bercher, Felix Richter 0001, Florian Nothdurft, Thomas Geier, Roland Barth, Thilo Hoernle, Felix Schüssel, Stephan Reuter, Matthias Rau, Gregor Bertrand, Bastian Seegebarth, Peter Kurzok, Bernd Schattenberg, Wolfgang Minker, Michael Weber 0001, Susanne Biundo-Stephan |
Intelligent Environments | 15 |
| 2014 | Application of Verbal Intelligence in Dialog Systems for Multimodal InteractionabstractIn this paper we present a prototypical dialog system adaptive to verbal user intelligence. Verbal intelligence (VI) is the ability to analyze information and to solve problems using language-based reasoning. VI can be analyzed by the number of reused words, lemmas, n-grams, cosine similarity and other features. Here we concentrate on the application of VI in a human-computer interaction (HCI) and how this value can be used by the dialog management to adapt the dialog flow, and complexity at run-time. In our work complexity as well as informative value of presented information can be reduced or increased when encountering human-computer interaction by individually adapting to a lower or higher verbally intelligent user. Especially in intelligent environments, where user's may rely on speech as their primary interaction modality, the adaptation of system instructions to the the user's VI could prove helpful. Florian Nothdurft, Frank Honold, Kseniya Zablotskaya, Amr Diab, Wolfgang Minker |
Intelligent Environments | 5 |
| 2014 | Probabilistic Explanation Dialog AugmentationabstractHuman-computer trust (HCT) is an important factor influencing the complexity and frequency of interaction in technical systems. Especially incomprehensible situations in human-computer interaction (HCI) may decrease the user's trust and through that the way of interaction. However, analogous to human-human interaction (HHI), providing explanations in these situations can help to remedy negative effects. In this paper, we present our approach of augmenting task-oriented dialogs with selected explanation dialogs to stabilize the HCT relationship. We conducted a study comparing the effects of different explanations on HCT. These results were used in a probabilistic trust handling architecture to augment pre-defined task-oriented dialogs. Florian Nothdurft, Felix Richter 0001, Wolfgang Minker |
Intelligent Environments | 3 |
| 2014 | Speech-Based Emotion Recognition: Feature Selection by Self-Adaptive Multi-Criteria Genetic Algorithm
Maxim Sidorov, Christina Brester, Wolfgang Minker, Eugene Semenkin |
LREC | 3 |
| 2014 | First Insight into Quality-Adaptive Dialogue
Stefan Ultes, Hüseyin Dikme, Wolfgang Minker |
LREC | 3 |
| 2014 | Probabilistic Human-Computer Trust HandlingabstractHuman-computer trust has shown to be a critical factor in influencing the complex-ity and frequency of interaction in techni-cal systems. Particularly incomprehensi-ble situations in human-computer interac-tion may lead to a reduced users trust in the system and by that influence the style of interaction. Analogous to human-human interaction, explaining these situations can help to remedy negative effects. In this pa-per we present our approach of augment-ing task-oriented dialogs with selected ex-planation dialogs to foster the human-computer trust relationship in those kinds of situations. We have conducted a web-based study testing the effects of different goals of explanations on the components of human-computer trust. Subsequently, we show how these results can be used in our probabilistic trust handling architec-ture to augment pre-defined task-oriented dialogs. 1 Florian Nothdurft, Felix Richter 0001, Wolfgang Minker |
SIGDIAL Conference | 3 |
| 2014 | Interaction Quality Estimation in Spoken Dialogue Systems Using Hybrid-HMMsabstractResearch trends on SDS evaluation are recently focusing on objective assessment methods. Most existing methods, which derive quality for each systemuser-exchange, do not consider temporal dependencies on the quality of previous exchanges. In this work, we investigate an approach for determining Interaction Quality for human-machine dialogue based on methods modeling the sequential characteristics using HMM modeling. Our approach significantly outperforms conventional approaches by up to 4.5% relative improvement based on Unweighted Average Recall metrics. Stefan Ultes, Wolfgang Minker |
SIGDIAL Conference | 2 |
| 2013 | JaCHMM: A Java-based conditioned Hidden Markov Model libraryabstractWe present JaCHMM, a Java implementation of a conditioned Hidden Markov Model (CHMM), which is made available under BSD license. It is based on the open source library “Jahmm” and provides implementations of the Viterbi, Forward-Backward, Baum-Welch and K-Means algorithms, all adapted for the CHMM. Like the Hidden Markov Model (HMM), the CHMM may be applied to a wide range of uni- and multimodal classification problems. The library is intended for academic and scientific purposes but may be also used in commercial systems. As a proof of concept, the JaCHMM library is successfully applied to speech-based emotion recognition outperforming HMM- and SVM-based approaches. Stefan Ultes, Robert ElChabb, Alexander Schmitt, Wolfgang Minker |
ICASSP | 4 |
| 2013 | On Quality Ratings for Spoken Dialogue Systems - Experts vs. Users
Stefan Ultes, Alexander Schmitt, Wolfgang Minker |
HLT-NAACL | 3 |
| 2013 | A Semi-supervised Approach for Natural Language Call Routing
Tatiana Gasanova, Eugene Zhukov, Roman B. Sergienko, Eugene Semenkin, Wolfgang Minker |
SIGDIAL Conference | 5 |
| 2013 | Improving Interaction Quality Recognition Using Error Correction
Stefan Ultes, Wolfgang Minker |
SIGDIAL Conference | 2 |
| 2012 | Adaptive Explanation Architecture for Maintaining Human-Computer TrustabstractOne of the most important challenges in the field of human-computer interaction is maintaining and enhancing the willingness of a user to interact with a technical system. This cooperativeness provides a solid basis for a real dialogue between user and technical system. The goals and tasks of an intelligent technical system seem unrealistic without it. In particular intelligent technical systems, which are continually assisting the user in his everyday life, degenerate without this willingness for dialogue to helpers for quite simple tasks and cannot fulfil their original purpose as intelligent assistants for complex tasks. Trust has shown to be an important factor influencing the frequency and kind of usage. If the user does not understand system actions or instructions, the trust of the user in the system will decrease and this can lead to a reduced frequency or in the worst case to a total cease of usage. Therefore, the intelligibility of a technical system should be upheld. This paper is concerned with how the intelligibility of an intelligent technical system can be upheld by providing explanations to the user. Providing explanations may prevent or at least decrease the loss of trust. However, trust is a complex construct consisting of different bases. We show why and how these bases of trust should be treated by giving individual kinds of explanations. Florian Nothdurft, Gregor Bertrand, Helmut Lang, Wolfgang Minker |
COMPSAC | 4 |
| 2012 | "What Do You Want to Do Next?" Providing the User with More Freedom in Adaptive Spoken Dialogue SystemsabstractSpoken dialogue is a suitable form for user interaction in intelligent environments because of its low demands to presentation hardware (e.g. microphone and speaker set) and its low cognitive demand. That means that the user can still do something else while communicating with a spoken dialogue system (SDS). In this paper, we investigate the demands to modern SDSs in intelligent environments and how to provide the user with more freedom by using constraint programming techniques for dialogue management. For this we present an architecture which allows for adaptivity and freedom in SDSs. Gregor Bertrand, Florian Nothdurft, Wolfgang Minker |
Intelligent Environments | 3 |
| 2012 | Speech Interaction with the Internet - A User StudyabstractThe arrival of smart phones significantly impacts the automotive environment. People tend to use their mobile Internet connection manually while driving which distracts and endangers the driver's safety. In order to reduce driver distraction a speech based interface to Internet services is essential. However, before developing a speech dialog system in a new domain, a data collection from real users is needed. In this report, a web-based user study is conducted, which aims at getting knowledge about how users would interact with Internet services by speech. The user study was separated in a questionnaire and audio recordings based on graphically depicted scenarios the subjects had to solve orally. The results show that the users are willing to use and trust in speech dialog systems. The speaking styles occurring in the audio data were classified into natural, command and keyword style. The occurrence differed depending on the web task category whereas natural speaking style was most frequently used by the subjects. Hansjörg Hofmann, Ute Ehrlich, André Berton, Wolfgang Minker |
Intelligent Environments | 4 |
| 2012 | Adaptive Speech Understanding for Intuitive Model-based Spoken Dialogues
Tobias Heinroth, Maximilian Grotz, Florian Nothdurft, Wolfgang Minker |
LREC | 4 |
| 2012 | Using multimodal resources for explanation approaches in intelligent systems
Florian Nothdurft, Wolfgang Minker |
LREC | 2 |
| 2012 | A Parameterized and Annotated Spoken Dialog Corpus of the CMU Let's Go Bus Information System
Alexander Schmitt, Stefan Ultes, Wolfgang Minker |
LREC | 3 |
| 2012 | Investigating Verbal Intelligence Using the TF-IDF Approach
Kseniya Zablotskaya, Fernando Fernández Martínez, Wolfgang Minker |
LREC | 3 |
| 2012 | Relating Dominance of Dialogue Participants with their Verbal Intelligence Scores
Kseniya Zablotskaya, Umair Rahim, Fernando Fernández Martínez, Wolfgang Minker |
LREC | 4 |
| 2012 | Speech and Language Resources for LVCSR of Russian
Sergey Zablotskiy, Alexander V. Shvets, Maxim Sidorov, Eugene Semenkin, Wolfgang Minker |
LREC | 5 |
| 2012 | Estimating Adaptation of Dialogue Partners with Different Verbal Intelligence
Kseniya Zablotskaya, Fernando Fernández Martínez, Wolfgang Minker |
SIGDIAL Conference | 3 |
| 2012 | Self-learning speaker identification for enhanced speech recognition
Tobias Herbig, Franz Gerl, Wolfgang Minker |
Comput. Speech Lang. | 3 |
| 2012 | Text categorization methods for automatic estimation of verbal intelligence
Fernando Fernández Martínez, Kseniya Zablotskaya, Wolfgang Minker |
Expert Syst. Appl. | 3 |
| 2011 | An Approach to Semi-supervised Classification using the Hungarian Algorithm
Amparo Albalate, Aparna Suchindranath, Wolfgang Minker |
ICAART (1) | 3 |
| 2011 | Rapid phonetic transcription using everyday life natural Chat Alphabet orthography for dialectal Arabic speech recognitionabstractWe propose the Arabic Chat Alphabet (ACA) as naturally written in everyday life for dialectal Arabic speech transcription. Our assumption is that ACA is a natural language that includes short vowels that are missing in traditional Arabic orthography. Furthermore, ACA transcriptions can be rapidly prepared. Egyptian Colloquial Arabic was chosen as a typical dialect. Two speech recognition baselines were built: phonemic and graphemic. Original transcriptions were re-written in ACA by different transcribers. Ambiguous ACA sequences were handled by automatically generating all possible variants. ACA variations across transcribers were modeled by phonemes normalization and merging. Results show that the ACA-based approach outperforms the graphemic baseline while it performs as accurate as the phoneme-based baseline with a slight increase in WER. Mohamed Elmahdy 0001, Rainer Gruhn, Slim Abdennadher, Wolfgang Minker |
ICASSP | 4 |
| 2011 | Measuring Verbal Intelligence Using Linguistic AnalysisabstractIn this paper we present a study on language use of people with different verbal intelligence. We asked test persons of different ages and educational background to describe the same event. Verbal intelligence was measured using the Hamburg Wechsler Intelligence Test for Adults. The transcribed monologues were analyzed using the DeLite readability checker and different linguistic features used for readability calculations were extracted. The test persons were then divided into two groups according to the results of the intelligence test using the SEM algorithm. For each group the averaged values of the features extracted from the monologues were compared using a one-way analysis of variance. Kseniya Zablotskaya, Mohsin Abbas, Sergey Zablotskiy, Steffen Walter 0001, Wolfgang Minker |
Intelligent Environments | 5 |
| 2011 | Optimal Operators of Hybrid Genetic Algorithm for GMM Parameter EstimationabstractA genetic algorithm is an evolutionary algorithm that is widely used for solving global optimization problems. It generates the solution in the form of encoded binary chromosome using operators inspired by a natural evolution process: selection, crossover and mutation. In this paper, a hybrid genetic algorithm is applied to the emission probability estimation task of a continuous Hidden Markov Model which is one of the common optimization problems in speech recognition. Three backbone operators of the genetic algorithm are investigated in order to find the optimal Gaussian parameters that result in the best mixture model. Sergey Zablotskiy, Teerat Pitakrat, Kseniya Zablotskaya, Wolfgang Minker |
Intelligent Environments | 4 |
| 2011 | Topic Switching Strategies for Spoken Dialogue Systems
Tobias Heinroth, Savina Koleva, Wolfgang Minker |
INTERSPEECH | 3 |
| 2011 | Tackling a Shilly-Shally Classifier for Predicting Task Success in Spoken Dialogue Interaction
Alexander Schmitt, Alexander Zgorzelski, Wolfgang Minker |
INTERSPEECH | 3 |
| 2011 | Attention, Sobriety Checkpoint! Can Humans Determine by Means of Voice, if Someone is Drunk... and Can Automatic Classifiers Compete?abstractThis paper analyzes the human performance of recognizing drunk speakers merely by voice and compares the results with the performance of an automatic statistical classifier. The study is carried out within the Interspeech 2011 Speaker State Challenge [1] employing the Alcohol Language Corpus (ALC) [2]. The 79 subjects yielded an average performance of 55.8% unweighted accuracy on a balanced intoxicated/non-intoxicated sample set. The statistical classifier developed in this study reaches a performance of 66.6% unweighted accuracy on the test set. In comparison, the subject with the highest performance yielded 70.0%. Our classifier is based on 4368 acoustic and prosodic features. Incorporating linguistic features along with feature selection using Information Gain Ratio (IGR) ranking added 0.7% absolute improvement with resulting in a 29% smaller feature space size. Copyright © 2011 ISCA. Stefan Ultes, Alexander Schmitt, Wolfgang Minker |
INTERSPEECH | 3 |
| 2011 | Modeling and Predicting Quality in Spoken Human-Computer Interaction
Alexander Schmitt, Benjamin Schatz, Wolfgang Minker |
SIGDIAL Conference | 3 |
| 2010 | Fast Adaptation of Speech and Speaker Characteristics for Enhanced Speech Recognition in Adverse Intelligent EnvironmentsabstractIn this paper we present a technique for fast adaptation of speech and speaker related information. Fast learning is particularly useful for automatic personalization of speech-controlled devices. Such a personalization of human-computer interfaces to be used in intelligent environments represents an important research issue. Speech recognition is enhanced by speaker specific profiles which are continuously adapted. A fast but robust tracking of speaker characteristics and optimal long-term adaptation are investigated to avoid an extensive enrollment of new speakers. We present an implementation suitable for speaker specific speech recognition in adverse intelligent environments. Exemplarily, in-car applications such as speech controlled navigation, hands-free telephony or infotainment systems are investigated for embedded systems. Results for a subset of the SPEECON database are presented. They validate the benefit of the presented speaker adaptation scheme for speech recognition. Speaker characteristics are captured after very few utterances. In the long run speaker characteristics are accurately represented. This adaptation scheme might be used to develop an unsupervised speech controlled system comprising speech recognition and speaker identification. A unified modeling of speech and speaker characteristics is proposed. Tobias Herbig, Franz Gerl, Wolfgang Minker |
Intelligent Environments | 3 |
| 2010 | GEEDI - Guards for Emotional and Explanatory DIaloguesabstractIn this paper, we describe the development of a dialogue model that integrates emotional dialogue strategies and explanations in a simple hence powerful way. As intelligent environments make inroads into the market, the need for user-friendly interaction with these systems grows. Pro-active reaction to user knowledge and emotions is one of the key points in user-friendly adaption of dialogue systems and therefore one of the main topics of research. As intelligent environments grow in complexity and field of application, the knowledge requirements for the user grow as well. Therefore it is vitally important to impart knowledge and information in an emotionally sensitive and user-aware way. In our dialog model we consider the natural structure of a nontrivial dialogue as a structure divided into several goals. These goals are protected by so called guards which represent preconditions which have to be fulfilled in order to tackle the related goal. Florian Nothdurft, Gregor Bertrand, Tobias Heinroth, Wolfgang Minker |
Intelligent Environments | 4 |
| 2010 | Inter-labeler Agreement for Anger Detection in Interactive Voice Response SystemsabstractAnger detection in speech-based automated telephone applications is a growing field of research. In this work we report on inter-labeler agreement in a “real-life” anger detection task for Interactive Voice Response (IVR) systems. The presented study is based on a corpus of 1.911 calls containing 22.711 utterances and describes considerations prior to the rating process. We point out difficulties we faced when annotating the corpus and present statistics and agreement values obtained after rating. The 3 raters that were asked to annotate angry user utterances agreed on the nature of “non-angry” utterances, but had difficulties to find an agreement on how an angry user utterance should sound. Alexander Schmitt, Ulrich Tschaffon, Wolfgang Minker |
Intelligent Environments | 3 |
| 2010 | Non-parametric Regression and Random Balance Method Modification for Determination of the Most Informative FeaturesabstractIn this paper we present a new method which allows us to detect the most informative features out of all data extracted from a certain data corpus. Widely used Pearson's coefficient is not reliable if the dependency between extracted features (input variables) and the objective function (output) is not linear. This approach is based on a modified random balance method (RBM) combined with non-parametric kernel regression for modeling the dependency between output and input variables. The standard random balance method stochastically determines the most important features of a process, but it requires the values of the objective function at the certain assigned points. If there is no possibility to calculate these values, it is necessary to approximate them. Since we assume that the dependency between stochastic variables can be non-linear, it is necessary to take an appropriate model. We used non-parametric kernel regression because knowledge about the parametric structure of the dependency is not needed. Moreover, we modified the random balance method to handle the non-linearity of the data. Kseniya Zablotskaya, Mumtaz Ahmed, Sergey Zablotskiy, Wolfgang Minker |
Intelligent Environments | 4 |
| 2010 | Some Approaches for Russian Speech RecognitionabstractIn this paper we present an overview of the state-of-the-art approaches for speech recognition of the Russian language. Since Russian is a highly inflective language with a complex mechanism of word formation, the main approaches for English speech recognition are not optimally applicable for Russian speech, at least directly and without appropriate adaptation. We make here an overview of some modern methods designed specially for Russian and other languages with similar word formation mechanism and grammar basis. Among them there are morpheme based speech recognition and compound language modeling. We also describe here our direction for possible subsequent Russian speech recognizer improvements. Sergey Zablotskiy, Kseniya Zablotskaya, Wolfgang Minker |
Intelligent Environments | 3 |
| 2010 | A semi-supervised cluster-and-label approach for utterance classification
Amparo Albalate, Aparna Suchindranath, David Suendermann-Oeft, Wolfgang Minker |
INTERSPEECH | 4 |
| 2010 | Cross-lingual acoustic modeling for dialectal Arabic speech recognitionabstractAmajor problem with dialectal Arabic acoustic modeling is due to the very sparse available speech resources. In this paper, we have chosen Egyptian Colloquial Arabic (ECA) as a typical di-alect. In order to benefit from existing Modern Standard Arabic (MSA) resources, a cross-lingual acoustic modeling approach is proposed that is based on supervised model adaptation. MSA acoustic models were adapted using MLLR and MAP with an in-house collected ECA corpus. Phoneme-based and grapheme-based acoustic modeling were investigated. To make phoneme-based adaptation feasible, we have normalized the phoneme sets of MSA and ECA. Since dialectal Arabic is mainly spo-ken, graphemic form usually does not match actual spelling as in MSA, a graphemic MSA acoustic model was used to force align and to choose the correct ECA spelling from a set of au-tomatically generated spelling variants lexicon. Results show that the adapted MSA acoustic models outperformed acoustic models trained with only ECA data. Mohamed Elmahdy 0001, Rainer Gruhn, Wolfgang Minker, Slim Abdennadher |
INTERSPEECH | 3 |
| 2010 | Speaker tracking in an unsupervised speech controlled system
Tobias Herbig, Franz Gerl, Wolfgang Minker |
INTERSPEECH | 3 |
| 2010 | Is it possible to predict task completion in automated troubleshooters?abstractThede online prediction of task success in Interactive Voice Response (IVR) systems is a comparatively new field of research. It helps to identify problemantic calls and enables the dialog system to react before the caller gets overly frustrated. This publication investigates, to which extent it is possible to predict task completion and how existing approaches generalize for long dialogs. We compare the performance of two different modeling techniques: linear modeling and n-gram modeling. We show that n-gram modeling outperforms linear modeling significantly at later prediction points. From a comprehensive set of interaction parameters, we identify the relevant ones using the Information Gain Ratio. New interaction parameters are presented and evaluated. The study is based on 41,422 calls from an automated Internet troubleshooter with an average of 21.4 turns per call. Alexander Schmitt, Wolfgang Minker, Jackson Liscombe, David Suendermann-Oeft |
INTERSPEECH | 3 |
| 2010 | Repair strategies on trial: which error recovery do users like best?
Alexander Zgorzelski, Alexander Schmitt, Tobias Heinroth, Wolfgang Minker |
INTERSPEECH | 4 |
| 2010 | Towards Investigating Effective Affective Dialogue Strategies
Gregor Bertrand, Florian Nothdurft, Steffen Walter 0001, Andreas Scheck, Henrik Kessler, Wolfgang Minker |
LREC | 6 |
| 2010 | Efficient Spoken Dialogue Domain Representation and Interpretation
Tobias Heinroth, Dan Denich, Alexander Schmitt, Wolfgang Minker |
LREC | 4 |
| 2010 | WITcHCRafT: A Workbench for Intelligent exploraTion of Human ComputeR conversaTions
Alexander Schmitt, Gregor Bertrand, Tobias Heinroth, Wolfgang Minker, Jackson Liscombe |
LREC | 4 |
| 2010 | The Influence of the Utterance Length on the Recognition of Aged Voices
Alexander Schmitt, Tim Polzehl, Wolfgang Minker, Jackson Liscombe |
LREC | 3 |
| 2010 | Speech Data Corpus for Verbal Intelligence Estimation
Kseniya Zablotskaya, Steffen Walter 0001, Wolfgang Minker |
LREC | 3 |
| 2010 | Advances in the Witchcraft Workbench Project
Alexander Schmitt, Wolfgang Minker, Nada Sharaf |
SIGDIAL Conference | 2 |
| 2010 | Simultaneous speech recognition and speaker identificationabstractIn this paper we present a self-learning speech controlled system comprising speech recognition, speaker identification and speaker adaptation for a small number of users, e.g. five recurring speakers. A compact representation of speech and speaker characteristics is discussed. It is combined with a technique for efficient information retrieval to capture individual speech characteristics allowing robust speaker identification with limited training data. Speech recognition is enhanced by applying speaker specific profiles which are incrementally adapted. However, the computational load and memory consumption are essential design parameters for an embedded system. Such a personalization of human-computer interfaces represents an important research issue. In this paper in-car applications such as speech controlled navigation, hands-free telephony or infotainment systems are investigated. Results for a subset of the SPEECON database are presented. They validate the benefit of the unified modeling of speech and speaker characteristics. Tobias Herbig, Franz Gerl, Wolfgang Minker |
SLT | 3 |
| 2009 | A hierarchical structure for modeling inter and intra phonetic information for phoneme recognitionabstractIn this paper, we present a two-layer hierarchical structure based on neural networks for phoneme recognition. The proposed structure attempts to model only the characteristics within a phoneme, i.e., intra-phonetic information. This differs from other state-of-the-art hierarchical structures where the first layer typically models the intra-phonetic information while the second layer focuses on modeling the contextual (inter-phonetic) information. An advantage of the proposed model is that it can be added to another layer that focuses on the inter-phonetic information. In this paper, we also show that the categorization between intra- and inter-phonetic information also allows to extend other state-of-the-art hierarchical approaches. A phoneme accuracy of 77.89% is achieved on the TIMIT database, which compares favorably to the best results obtained on this database. Daniel Vásquez, Guillermo Aradilla, Rainer Gruhn, Wolfgang Minker |
ASRU | 4 |
| 2009 | On speeding phoneme recognition in a hierarchical MLP structureabstractIn this paper, we propose a technique for speeding phoneme recognition in a hierarchical structure involving Multilayered Perceptrons (MLPs). The hierarchical structure consists of two MLP-based layers, where the output of the first layer is used as input for the second layer. In this paper, we efficiently speed up the system by removing the redundant information contained at the output of the first layer. Several techniques are investigated for removing this redundant information based on temporal and phonetic criteria. The best approach reduces the computational time by 57% while keeping a system accuracy comparable to the standard hierarchical approach. This scheme favors the implementation of such hierarchical structures in real-time applications. Daniel Vásquez, Guillermo Aradilla, Rainer Gruhn, Wolfgang Minker |
ASRU | 4 |
| 2009 | Multidimensional Pervasive Adaptation into Ambient Intelligent EnvironmentsabstractIn this paper we describe the ATRACO (adaptive and trusted ambient ecologies) approach towards next generation ambient intelligent environments. Several agents, such as a fuzzy task agent with learning capabilities and an interaction agent collaborate in a goal-related activity sphere and adapt heterogeneous artifacts within the sphere in order to support the user to fulfill tasks. All components work on a dynamic sphere ontology, which forms the main knowledge base of the ecology. The presented prototype is able to realize the goal ¿feel comfortable at home after work¿ and was implemented in an existing intelligent environment. Yacine Bellik, Gaëtan Pruvost, Achilles Kameas, Christos Goumopoulos, Hani Hagras, Michael Gardner, Tobias Heinroth, Wolfgang Minker |
DASC | 8 |
| 2009 | Creating an Ambient Intelligent Environment with an Emotion-Aware SystemabstractIn this paper, we describe a novel approach of combining an emotion voice aware system with a fuzzy logic system to develop embedded agents aimed at creating an Ambient Intelligent Environment (AIE). This combined approach was evaluated in the GUC intelligent Classroom (iClass) which a is real world AIE testbed. The developed system aims to control the user environment on his behalf while taking into account the user's emotions. The agent learns online the fuzzy membership functions and rules from the user monitored actions to generate an agent that models the user behavior in an educational AIE like the iClass. The agent then operates in a life long learning mode where the agent can be adapted over long time intervals to the user's changing desires and preferences. Sherief Mowafey, Alexander Schmitt, Hani Hagras, Wolfgang Minker |
Intelligent Environments | 4 |
| 2009 | Perceptual Evaluation of Radio Signal Quality DegradationabstractIn this paper we present an approach, which avoids expansive and time consuming subjective assessment of audio quality degradation, caused by different nature disturbances during the transmitting and receiving of stereo audio signal through the radio channel. This approach is based on the basic version of PEAQ (Perceptual Evaluation of Audio Quality) method, optimized generally for audio codec testing. The MOV (Model Output Variables) vector of the PEAQ method is transformed to the transmitted audio quality degradation scale, using the combination of neural networks and genetic algorithms. Sergey Zablotskiy, Wolfgang Minker |
Intelligent Environments | 3 |
| 2009 | Combining Agents and Ontologies to Support Task-Centred Interoperability in Ambient Intelligent EnvironmentsabstractThis article describes our approach towards the specification and realization of interoperability within Next Generation Ambient Intelligent Environments (NGAIE). These are populated with numerous devices and multiple occupants or users exhibit increasingly intelligent behaviour, provide optimized resource usage and support consistent functionality and human-centric operation. In NGAIEs users will interact with their environments using the devices therein complemented with adaptive multimodal dialogue. This requires the definition of the local and global information which is relevant to the interaction and mechanisms to share this knowledge among entities. In our approach, knowledge is represented as a set of heterogeneous ontologies which have to be aligned in order to provide a uniform and consistent knowledge representation. The combination of heterogeneous ontologies and ontology matching algorithms allows for semantically rich information exchange. Based on a combination of agent-based and service-oriented architectures, the proposed approach adopts a task-based model to maximize the use of available heterogeneous resources. Gaëtan Pruvost, Achilles Kameas, Tobias Heinroth, Lambrini Seremeti, Wolfgang Minker |
ISDA | 5 |
| 2008 | The PIT Corpus of German Multi-Party Dialogues
Petra-Maria Strauß, Holger Hoffmann, Wolfgang Minker, Heiko Neumann, Günther Palm, Stefan Scherer, Harald C. Traue, Ulrich Weidenbacher |
LREC | 3 |
| 2007 | BECAM tool - a semi-automatic tool for bootstrapping emotion corpus annotation and managementabstractCorpus annotation is an important aspect in speech applications where stochastic models need to be trained and evaluated. Multimodal corpora are also annotated. Moreover, corpus annotation is an essential phase in the construction of emotion recognizer engines. Large corpora, as they are essential to construct representative knowledge bases, have been a problem for corpus annotators. Time consumed for labeling such corpora is very significant. Furthermore, manageability becomes more arduous and tedious. In this paper, we propose a semi-automatic tool, called BECAM tool, that will help corpus annotators in managing and annotating large sample emotion corpora. Index Terms: Corpus annotation, emotion recognition, bootstrap 1. Slim Abdennadher, Dirk Bühler, Wolfgang Minker, Johannes Pittermann |
INTERSPEECH | 4 |
| 2006 | Stochastic Spoken Natural Language Parsing in the Framework of the French MEDIA Evaluation Campaign
Dirk Bühler, Wolfgang Minker |
LREC | 2 |
| 2006 | Wizard-of-Oz Data Collection for Perception and Interaction in Multi-User Environments
Petra-Maria Strauß, Holger Hoffmann, Wolfgang Minker, Heiko Neumann, Günther Palm, Stefan Scherer, Friedhelm Schwenker, Harald C. Traue, Welf Walter, Ulrich Weidenbacher |
LREC | 3 |
| 2005 | Speech and Human'Machine Dialog
Wolfgang Minker, Samir Bennacef |
Comput. Linguistics | 1 |
| 2004 | New challenges in usability evaluation - beyond task-oriented spoken dialogue systemsabstractThere is a fairly good baseline for usability evaluation of taskoriented unimodal spoken dialogue systems (SDSs) but much is still unknown regarding the usability of multimodal and non-task-oriented SDSs. This paper reviews and discusses approaches to usability evaluation of these kinds of SDSs. Laila Dybkjær, Niels Ole Bernsen, Wolfgang Minker |
INTERSPEECH | 3 |
| 2004 | Strategies for optimizing a stochastic spoken natural language parser
Wolfgang Minker, Dirk Bühler, Christiane Beuschel |
INTERSPEECH | 1 |
| 2004 | Usability Evaluation of Multimodal and Domain-Oriented Spoken Language Dialogue Systems
Laila Dybkjær, Niels Ole Bernsen, Wolfgang Minker |
LREC | 3 |
| 2004 | Comparative Evaluation of a Stochastic Parser on Semantic and Syntactic-semantic Labels
Wolfgang Minker |
LREC | 1 |
| 2004 | Evaluation and usability of multimodal spoken language dialogue systems
Laila Dybkjær, Niels Ole Bernsen, Wolfgang Minker |
Speech Commun. | 3 |
| 2004 | The SENECA spoken language dialogue system
Wolfgang Minker, Udo Haiber, Paul Heisterkamp, Sven Scheible |
Speech Commun. | 1 |
| 2003 | Safety and operating issues for mobile human-machine interfacesabstractIn this paper we present recent research and development efforts carried out at DaimlerChrysler to integrate speech technology for use in mobile environments, notably in cars. Speech undeniably has the potential to considerably improve the safety and user friendliness of Human-machine interfaces, especially when complex technical functionalities and devices need to be accessed. As an example, we describe Linguatronic, a commercially available in-vehicle Command&Control dialog system. In addition, the SmartKom project demonstrates advanced concepts for intuitive multimodal computer interfaces in three different application scenarios. Dirk Bühler, Sébastien Vignier, Paul Heisterkamp, Wolfgang Minker |
IUI | 4 |
| 2003 | Intelligent dialog overcomes speech technology limitations: the SENECa exampleabstractWe present a primarily speech-based user interface to a wide range of entertainment, navigation and communication applications for use in vehicles. The multimodal dialog en ables the system to uniquely identify one of 79,000 place name variants using an active vocabulary of only 3,000 words at any given time. Low confidence in speech recog nition and word-level ambiguities are compensated for in flexible clarification dialogs with the user. The underlying dialog concept was developed in the framework of the EU-project SENECa. Some recent evalua tion results of the SENECa system demonstrator are discussed in the paper Wolfgang Minker, Udo Haiber, Paul Heisterkamp, Sven Scheible |
IUI | 1 |
| 2002 | Flexible multimodal human-machine interaction in mobile environmentsabstractThis article describes requirements and a prototype system for a flexible multimodal human-machine interaction in two substantially different mobile environments, namely pedestrian and car. The system allows an integrated trip planning using multimodal input and output. Motivated by the specific safety and privacy requirements in both environments, we present a framework for flexible modality control. A characteristic feature of our framework is the insight that both user and system may independently and asynchronously initiate a modality transition. We conclude with a brief discussion of further issues and research questions. Dirk Bühler, Wolfgang Minker, Jochen Häußler, Sven Krger |
INTERSPEECH | 2 |
| 2002 | Overview on recent activities in speech understanding and dialogue systems evaluation
Wolfgang Minker |
INTERSPEECH | 1 |
| 2002 | Overview on recent activities in speech understanding and dialogue systems evaluation
Wolfgang Minker |
INTERSPEECH | 1 |
| 2001 | Robustness and Portability Issues in Multilingual Speech Processing
Wolfgang Minker |
Mach. Transl. | 1 |
| 2000 | Towards best practice in the development and evaluation of speech recognition components of a spoken language dialog systemabstractThis article provides a global overview of the main aspects of current practice in the design, implementation and evaluation of speech recognition components for Spoken Language Dialog Systems (SLDSs), and presents the results of the DISC European project related to speech recognition. DISC and its successor DISC-2 are efforts towards the definition of best practice guidelines for SLDS development and evaluation. SLDSs aim at using natural spoken input for performing an information processing task such as automated standards, call routing or travel planning and reservations. The main functionality of an SLDS are speech recognition, natural language understanding, dialog management, database access and interpretation, response generation and speech synthesis. Speech recognition, which transforms the acoustic signal into a string of words, is a key technology in any SLDS. Lori Lamel, Wolfgang Minker, Patrick Paroubek |
Nat. Lang. Eng. | 2 |
| 1999 | Stochastically-based semantic analysis for machine translation
Wolfgang Minker, Marsal Gavaldà, Alex Waibel |
Comput. Speech Lang. | 1 |
| 1999 | Design considerations for knowledge source representations of a stochastically-based natural language understanding component
Wolfgang Minker |
Speech Commun. | 1 |
| 1998 | Evaluation methodologies for interactive speech system
Wolfgang Minker |
LREC | 1 |
| 1998 | Stochastic versus rule-based speech understanding for information retrieval
Wolfgang Minker |
Speech Commun. | 1 |
| 1997 | Stochastically-based natural language understanding across tasks and languagesabstractA stochastically-based method for natural language understanding has been ported from the American ATIS (Air Travel Information Services) to the French MASK (Multimodal-Multimedia Automated Service Kiosk) task. The porting was carried out by designing and annotating a corpus of semantic representations via a semi-automatic iterative labeling. The study shows that domain and language porting is rather flexible, since it is sufficient to train the system on data sets specific to the application and language. A limiting factor of the current implementation is the quality of the semantic representation and the use of query preprocessing strategies which strongly suffer from human influence. The performances of the stochastically-based and a rule-based method are compared on both tasks. 1. INTRODUCTION In this paper, we report on our experience in porting a stochastically-based natural language understanding component across tasks and languages. Stochastically-based methods have been appli... Wolfgang Minker |
EUROSPEECH | 1 |
| 1996 | A stochastic case frame approach for natural language understandingabstractA stochastically based approach for the semantic analysis component of a natural spoken language system for the ATIS task has been developed.The semantic analyzer of the spoken language system already in use at LIMSI makes use of a rule-based case grammar.In this work, the system of rules for the semantic analysis is replaced with a relatively simple, first order Hidden Markov Model.The performance of the two approaches can be compared because they use identical semantic representations despite their rather different methods for meaning extraction.We use an evaluation methodology that assesses performance at different semantic levels, including the database response comparison used in the ARPA ATIS paradigm. Wolfgang Minker, Samir Bennacef, Jean-Luc Gauvain |
ICSLP | 1 |
| 1994 | A spoken language system for information retrieval
Samir Bennacef, Hélène Bonneau-Maynard, Jean-Luc Gauvain, Lori Lamel, Wolfgang Minker |
ICSLP | 5 |