Gerhard Sagerer

dblp:20/6444 · DBLP profile ↗
← Back
96ranked-venue papers
5as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 68 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 40 · 2 first-authorHuman-computer interaction and ubiquitous computing · 20 · 2 first-authorSystems, architecture and hardware · 14Applied, interdisciplinary, general and emerging computing · 13 · 2 first-authorDatabases, data management, data science and information retrieval · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
11 papers
Human-robot interaction · 84% Haptics and multimodal interaction · 14% Ubiquitous computing and smart environments · 1%
Artificial intelligence
7 papers
Face, body and person analysis · 31% Video understanding and tracking · 18% Speech recognition and synthesis · 14%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 97% Medical and health informatics · 3%

Topics — the 29 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Haptics and multimodal interaction
multimodal interaction
0.232008
Human-Oriented Interaction With an Anthropomorphic Robot · IEEE Trans. Robotics 2007
Human-like Person Tracking with an Anthropomorphic Robot · ICRA 2006
"Try something else!" - When users change their discursive behavior in human-robot interaction · ICRA 2008
Human-robot interaction
social robot
0.122010
The Bielefeld anthropomorphic robot head "Flobi" · ICRA 2010
Who am I talking with? A face memory for social robots · ICRA 2008
Human-robot interaction › anthropomorphism
anthropomorphic robot head
0.112010
The Bielefeld anthropomorphic robot head "Flobi" · ICRA 2010
Bioinformatics and computational biology › protein structure prediction
protein-protein docking
0.132005
Database driven test case generation for protein?Cprotein docking · Bioinform. 2005
Estimation and filtering of potential protein-protein docking positions · Bioinform. 1998
Protein Docking: Combining Symbolic Descriptions of Molecular Surfaces and Grid-Based Scoring Functions · ISMB 1995
Computer vision › Face, body and person analysis
face recognition
0.112008
Who am I talking with? A face memory for social robots · ICRA 2008
Human-robot interaction › cognitive human-robot interaction
theory of mind
0.112008
Theory of mind (ToM) on robots: a functional neuroimaging study · HRI 2008
Human-robot interaction › adaptive robot behavior
user adaptation
0.112008
"Try something else!" - When users change their discursive behavior in human-robot interaction · ICRA 2008
Human-robot interaction › robot perception
audio-visual perception
0.112006
Human-like Person Tracking with an Anthropomorphic Robot · ICRA 2006
Human-robot interaction › robot perception
person tracking
0.112006
Human-like Person Tracking with an Anthropomorphic Robot · ICRA 2006
Human-robot interaction › social robot
companion robots
0.112005
A Flexible Infrastructure for the Development of a Robot Companion with Extensible HRI-Capabilities · ICRA 2005
Computer vision › Video understanding and tracking
action recognition
0.012004
A Cognitive Vision System for Action Recognition in Office Environments · CVPR (2) 2004
Robotics › Robot manipulation
service robot
0.012010
The Bielefeld anthropomorphic robot head "Flobi" · ICRA 2010
Computer vision › Image recognition and object detection
object recognition
0.012001
Incorporating Process Knowledge into Object Recognition for Assemblies · ICCV 2001
Human-robot interaction › service robot
domestic robots
0.012009
Systemic interaction analysis (SInA) in HRI · HRI 2009
Human-robot interaction › anthropomorphism
robot appearance
0.012008
Theory of mind (ToM) on robots: a functional neuroimaging study · HRI 2008
Haptics and multimodal interaction › multimodal interaction
speech and gesture interaction
0.012008
"Try something else!" - When users change their discursive behavior in human-robot interaction · ICRA 2008
Human-robot interaction › robot communication
dialogue system
0.012007
Human-Oriented Interaction With an Anthropomorphic Robot · IEEE Trans. Robotics 2007
Bioinformatics and computational biology
protein structure prediction
0.011998
Estimation and filtering of potential protein-protein docking positions · Bioinform. 1998
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.021994
A Speech Understanding and Dialog System with a Homogeneous Linguistic Knowledge Base · IEEE Trans. Pattern Anal. Mach. Intell. 1994
ERNEST: A Semantic Network System for Pattern Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.011994
A Speech Understanding and Dialog System with a Homogeneous Linguistic Knowledge Base · IEEE Trans. Pattern Anal. Mach. Intell. 1994
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
linguistic knowledge integration
0.011994
A Speech Understanding and Dialog System with a Homogeneous Linguistic Knowledge Base · IEEE Trans. Pattern Anal. Mach. Intell. 1994
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems
0.011994
A Speech Understanding and Dialog System with a Homogeneous Linguistic Knowledge Base · IEEE Trans. Pattern Anal. Mach. Intell. 1994
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
semantic networks
0.021990
ERNEST: A Semantic Network System for Pattern Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 1990
A Knowledge Based System for Analysis of Gated Blood Pool Studies · IEEE Trans. Pattern Anal. Mach. Intell. 1985
Knowledge, reasoning and agents › Knowledge representation and reasoning
process knowledge
0.012001
Incorporating Process Knowledge into Object Recognition for Assemblies · ICCV 2001
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge-based systems
0.011985
A Knowledge Based System for Analysis of Gated Blood Pool Studies · IEEE Trans. Pattern Anal. Mach. Intell. 1985
Medical and health informatics › medical imaging
medical image analysis
0.011985
A Knowledge Based System for Analysis of Gated Blood Pool Studies · IEEE Trans. Pattern Anal. Mach. Intell. 1985
Data mining
clustering
0.011982
An Experimental Study of Some Algorithms for Unsupervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 1982
Data mining › clustering
unsupervised learning
0.011982
An Experimental Study of Some Algorithms for Unsupervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 1982
Multimedia analysis and retrieval › image analysis
image understanding
0.011990
ERNEST: A Semantic Network System for Pattern Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 1990

Methods — techniques the papers use, named apart from their topics

stereo vision · 0.4gyroscope motion compensation · 0.2active memory architecture · 0.2systemic interaction analysis · 0.1system-level tracing · 0.1user study · 0.1prisoners' dilemma game · 0.1functional neuroimaging · 0.1active appearance models · 0.1active appearance model · 0.1stereo audio · 0.1cognitive vision system · 0.0root mean square deviation · 0.0global rotation sampling · 0.0symbolic surface description · 0.0procedural knowledge representation · 0.0grid-based scoring · 0.0declarative knowledge representation · 0.0
YearPublicationVenuePosition
2013 Revisioning HRI given exponential technological growth
Peter H. Kahn Jr., Gerhard Sagerer, Andrea Thomaz, Takayuki Kanda 0001
HRI2
2010 Panel 1: grand technical and social challenges in human-robot interaction
abstract
Robots are becoming part of people's everyday social lives - and will increasingly become so. In future years, robots may become caretaking assistants for the elderly, or academic tutors for our children, or medical assistants, day care assistants, or psychological counselors. Robots may become our co-workers in factories and offices, or maids in our homes. They may become our friends. As we move to create our future with robots, hard problems in HRI exist, both technically and socially. The Fifth Annual Conference on HRI seeks to take up grand technical and social challenges in the field - and speak to their integration. This panel brings together 4 leading experts in the field of HRI to speak on this topic.
Nathan G. Freier, Minoru Asada, Pam Hinds, Gerhard Sagerer, J. Gregory Trafton
HRI4
2010 Multiple sequence alignment based bootstrapping for improved incremental word learning
abstract
We investigate incremental word learning with few training examples in a Hidden Markov Model (HMM) framework suitable for an interactive learning scenario with little prior knowledge. When using only a few training examples the initialization of the models is a crucial step. In the bootstrapping approach proposed, an unsupervised initialization of the parameters is performed, followed by the retraining and construction of a new HMM using multiple sequence alignment (MSA). Finally we analyze discriminative training techniques to increase the separability of the classes using minimum classification error (MCE). Recognition results are reported on isolated digits taken from the TIDIGITS database.
Irene Ayllón Clemente, Martin Heckmann, Gerhard Sagerer, Frank Joublin
ICASSP3
2010 The Bielefeld anthropomorphic robot head "Flobi"
abstract
A robot's head is important both for directional sensors and, in human-directed robotics, as the single most visible interaction interface. However, designing a robot's head faces contradicting requirements when integrating powerful sensing with social expression. Furher, reactions of the general public show that current head designs often cause negative user reactions and distract from the functional capabilities. Therefore, this contribution presents a novel anthropomorphic robot head called "Flobi", which combines state-of-the-art sensing functionality with an exterior that elicits a sympathetic emotional response. It can display primary and secondary emotions in a human-like way, to enable intuitive human-robot-interaction. To facilitate further research on facial appearance, the exterior is fully modular and replaceable. While current state-of-the-art still requires trade-offs when integrating sensing and social expression, Flobi has been designed to enable service robotic applications, with high-resolution, wide-angle stereo vision, gyroscope motion compensation and stereo audio. For ease of integration, the head is self-contained, including 18 actuators, sensors and control boards, all in a human-head sized package.
Ingo Lütkebohle, Frank Hegel, Simon Schulz, Matthias Hackel, Britta Wrede, Sven Wachsmuth, Gerhard Sagerer
ICRA7
2009 Understanding Social Robots
abstract
Research on social robots is mainly comprised of research into algorithmic problems in order to expand a robot's capabilities to improve communication with human beings. Also, a large body of research concentrates on the appearance, i.e. aesthetic form of social robots. However, only little reference to their definition is made. In this paper we argue that form, function, and context have to be taken systematically into account in order to develop a model to help us understand social robots. Therefore, we address the questions: What is a social robot, what are the interdisciplinary research aspects of social robotics, and how are these different aspects interlinked? In order to present a comprehensive and concise overview of the various aspects we present a framework for a definition towards social robots.
Frank Hegel, Claudia Muhl, Britta Wrede, Martina Hielscher-Fastabend, Gerhard Sagerer
ACHI5
2009 Systemic interaction analysis (SInA) in HRI
abstract
Recent developments in robotics enable advanced human-robot interaction. Especially interactions of novice users with robots are often unpredictable and, therefore, demand for novel methods for the analysis of the interaction in systemic ways. We propose Systemic Interaction Analysis (SInA) as a method to jointly analyze system level and interaction level in an integrated manner using one tool. The approach allows us to trace back patterns that deviate from prototypical interaction sequences to the distinct system components of our autonomous robot. In this paper, we exemplarily apply the method to the analysis of the follow behavior of our domestic robot BIRON. The analysis is the basis to achieve our goal of improving human-robot interaction iteratively.
Manja Lohse, Marc Hanheide, Katharina J. Rohlfing, Gerhard Sagerer
HRI4
2009 Mediated attention with multimodal augmented reality
abstract
We present an Augmented Reality (AR) system to support collaborative tasks in a shared real-world interaction space by facilitating joint attention. The users are assisted by information about their interaction partner's field of view both visually and acoustically. In our study, the audiovisual improvements are compared with an AR system without these support mechanisms in terms of the participants' reaction times and error rates. The participants performed a simple object-choice task we call the "gaze game" to ensure controlled experimental conditions. Additionally, we asked the subjects to fill in a questionnaire to gain subjective feedback from them. We were able to show an improvement for both dependent variables as well as positive feedback for the visual augmentation in the questionnaire.
Angelika Dierker, Christian Mertes, Thomas Hermann 0001, Marc Hanheide, Gerhard Sagerer
ICMI5
2009 A dynamic attention system that reorients to unexpected motion in real-world traffic environments
abstract
Abstract — In this paper we propose a system architecture that extends the current state-of-the-art in computational visual attention by incorporating the biological concept of ventral attention. According to recent findings regarding the neurobiological foundations of attention, there exist two separate but interacting attention systems in the human brain: the dorsal attention system and the ventral attention system. As opposed to the well-known computational concepts of bottomup and top-down saliency, which both correspond to the dorsal attention system, the ventral attention system is sensitive to behavior-relevant stimuli that are unexpected (i.e. not top-down salient), independent of their perceptual saliency (bottom-up saliency). This results in a dynamic interplay between topdown saliency, bottom-up saliency and ventral attention in the proposed system architecture, enabling the system to redirect its focus of attention to important stimuli while being absorbed in a task, even if their perceptual saliency is low. Our technical system instance implementing the proposed architecture integrates several state-of-the-art methods in a coherent system and concentrates on unexpected motion as a first technical account of ventral attention. In our experiments, we demonstrate that the ventral attention enables our system to detect and reorient to important situations in real-world traffic environments that are relevant for the behavior of driving. I.
Martin Heracles, Ursula Körner, Gerhard Sagerer, Jannik Fritsch, Christian Goerick
IROS4
2009 Feedback interpretation based on facial expressions in human-robot interaction
abstract
In everyday conversation besides speech people also communicate by means of nonverbal cues. Facial expressions are one important cue, as they can provide useful information about the conversation, for instance, whether the interlocutor seems to understand or appears to be puzzled. Similarly, in human-robot interaction facial expressions also give feedback about the interaction situation. We present a Wizard of Oz user study in an object-teaching scenario where subjects showed several objects to a robot and taught the objects' names. Afterward, the robot should term the objects correctly. In a first evaluation, we let other people watch short video sequences of this study. They decided by looking at the face of the human whether the answer of the robot was correct (unproblematic situation) or incorrect (problematic situation). We conducted the experiments under specific conditions by varying the amount of temporal and visual context information and compare the results with related experiments described in the literature.
Christian Lang 0002, Marc Hanheide, Manja Lohse, Heiko Wersing, Gerhard Sagerer
RO-MAN5
2008 Fusion of perceptual processes for real-time object tracking
Kai Jüngling, Michael Arens, Marc Hanheide, Gerhard Sagerer
FUSION4
2008 Theory of mind (ToM) on robots: a functional neuroimaging study
abstract
Theory of Mind (ToM) is not only a key capability for cognitive development but also for successful social interaction. In order for a robot to interact successfully with a human both interaction partners need to have an adequate representation of the other's actions. In this paper we address the question of how a robot's actions are perceived and represented in a human subject interacting with the robot and how this perception is influenced by the appearance of the robot. We present the preliminary results of an fMRI-study in which participants had to play a version of the classical Prisoners' Dilemma Game (PDG) against four opponents: a human partner (HP), an anthropomorphic robot (AR), a functional robot (FR), and a computer (CP). The PDG scenario enables to implicitly measure mentalizing or Theory of Mind (ToM) abilities, a technique commonly applied in functional imaging. As the responses of each game partner were randomized unknowingly to the participants, the attribution of intention or will to an opponent (i.e. HP, AR, FR or CP) was based purely on differences in the perception of shape and embodiment.
Frank Hegel, Soeren Krach, Tilo Kircher, Britta Wrede, Gerhard Sagerer
HRI5
2008 Who am I talking with? A face memory for social robots
abstract
In order to provide personalized services and to develop human-like interaction capabilities robots need to recognize their human partner. Face recognition has been studied in the past decade exhaustively in the context of security systems and with significant progress on huge datasets. However, these capabilities are not in focus when it comes to social interaction situations. Humans are able to remember people seen for a short moment in time and apply this knowledge directly in their engagement in conversation. In order to equip a robot with capabilities to recall human interlocutors and to provide user- aware services, we adopt human-human interaction schemes to propose a face memory on the basis of active appearance models integrated with the active memory architecture. This paper presents the concept of the interactive face memory, the applied recognition algorithms, and their embedding into the robot's system architecture. Performance measures are discussed for general face databases as well as scenario-specific datasets.
Marc Hanheide, Sebastian Wrede 0001, Christian Lang 0002, Gerhard Sagerer
ICRA4
2008 "Try something else!" - When users change their discursive behavior in human-robot interaction
abstract
This paper investigates the influence of feedback provided by an autonomous robot (BIRON) on users' discursive behavior. A user study is described during which users show objects to the robot. The results of the experiment indicate, that the robot's verbal feedback utterances cause the humans to adapt their own way of speaking. The changes in users' verbal behavior are due to their beliefs about the robots knowledge and abilities. In this paper they are identified and grouped. Moreover, the data implies variations in user behavior regarding gestures. Unlike speech, the robot was not able to give feedback with gestures. Due to the lack of feedback, users did not seem to have a consistent mental representation of the robot's abilities to recognize gestures. As a result, changes between different gestures are interpreted to be unconscious variations accompanying speech.
Manja Lohse, Katharina J. Rohlfing, Britta Wrede, Gerhard Sagerer
ICRA4
2008 Automatic Initialization for Facial Analysis in Interactive Robotics
Ahmad Rabie, Christian Lang 0002, Marc Hanheide, Modesto Castrillón-Santana, Gerhard Sagerer
ICVS5
2008 Interacting with a mobile robot: Evaluating gestural object references
abstract
Creating robots able to interact and cooperate with humans in household environments and everyday life is an emerging topic. Our goal is to facilitate a human-like and intuitive interaction with such robots. Besides verbal interaction, gestures are a fundamental aspect in human-human interaction. One typical usage of interactive gestures is referencing of objects. This paper describes a novel integrated vision system combining different algorithms for pose tracking, gesture detection, and object attention in order to enable a mobile robot to resolve gesture-based object references. Results from the evaluation of the individual algorithms as well as the overall system are presented. A total of 20 minutes of video data collected from four subjects performing almost 500 gestures are evaluated to demonstrate the current performance of the approach as well as the overall success rate of gestural object references. This demonstrates that our integrated vision system can serve as the gestural front end that enables an interactive mobile robot to engage in multimodal human-robot interaction.
Joachim Schmidt 0001, Nils Hofemann, Axel Haasch, Jannik Fritsch, Gerhard Sagerer
IROS5
2008 Active memory-based interaction strategies for learning-enabling behaviors
abstract
Despite increasing efforts in the field of social robotics and interactive systems integrated and fully autonomous robots which are capable of learning from interaction with inexperienced and non-expert users are still a rarity. However, in order to tackle the challenge of learning by interaction robots need to be equipped with a set of basic behaviors and abilities which have to be coupled and combined in a flexible manner. This paper presents how a recently proposed information-driven integration concept termed ldquoactive memoryrdquo is adopted to realize learning-enabling behaviors for a domestic robot. These behaviors enable it to (i) learn about its environment, (ii) interact with several humans simultaneously, and (iii) couple learning and interaction tightly. The basic interaction strategies on the basis of information exchange through the active memory are presented. A brief discussion of results obtained from live user trials with inexperienced users in a home tour scenario underpin the relevance and appropriateness of the described concepts.
Marc Hanheide, Gerhard Sagerer
RO-MAN2
2008 Understanding social robots: A user study on anthropomorphism
abstract
Anthropomorphism is one of the keys to understand the expectations people have about social robots. In this paper we address the question of how a robotpsilas actions are perceived and represented in a human subject interacting with the robot and how this perception is influenced only by the appearance of the robot. We present results of an interaction-study in which participants had to play a version of the classical Prisonerspsila Dilemma Game (PDG) against four opponents: a human partner (HP), an anthropomorphic robot (AR), a functional robot (FR), and a computer (CP). As the responses of each game partner were randomized unknowingly to the participants, the attribution of intention or will to an opponent (i.e. HP, AR, FR or CP) was based purely on differences in the perception of shape and embodiment. We hypothesize that the degree of human-likeness of the game partner will modulate what the people attribute to the opponents - the more human like the robot looks the more people attribute human-like qualities to the robot.
Frank Hegel, Soeren Krach, Tilo Kircher, Britta Wrede, Gerhard Sagerer
RO-MAN5
2008 Evaluating extrovert and introvert behaviour of a domestic robot - a video study
abstract
Human-robot interaction (HRI) research is here presented into social robots that have to be able to interact with inexperienced users. In the design of these robots many research findings of human-human interaction and human-computer interaction are adopted but the direct applicability of these theories is limited because a robot is different from both humans and computers. Therefore, new methods have to be developed in HRI in order to build robots that are suitable for inexperienced users. In this paper we present a video study we conducted employing our robot BIRON (Bielefeld robot companion) which is designed for use in domestic environments. Subjects watched the system during the interaction with a human and rated two different robot behaviours (extrovert and introvert). The behaviours differed regarding verbal output and person following of the robot. Aiming to improve human-robot interaction, participantspsila ratings of the behaviours were evaluated and compared.
Manja Lohse, Marc Hanheide, Britta Wrede, Michael L. Walters, Kheng Lee Koay, Dag Sverre Syrdal, Anders Green, Helge Hüttenrauch, Kerstin Dautenhahn, Gerhard Sagerer, Kerstin Severinson Eklundh
RO-MAN10
2008 Memory and learning for social robots
abstract
While most research in social robotics embraces the challenge of designing and studying the interaction between robots and humans itself, this talk will discuss the utility of social interaction in order to facilitate for more flexible robotics. What can a robot gain with respect to learning and adaptation from being able to sociably interact? What are basic learning-enabling behaviors? And how do inexperienced human tutor robots a sociable way? In order to answer these question we consider the challenge of learning by interaction as a systemic one, comprising appropriate perception, system design, and feedback. Basic abilities of robots will be outlined which resemble concepts of developmental learning in infants, apply linguistic models of interaction management, and take tutoring as a joint task of a human and a robot. However, in order to tackle the challenge of learning by interaction the robot has to couple and coordinate these behaviors in a very flexible and adaptive manner. The active memory as an architectural concept in particular suitable for learning-enabled robots will be briefly discussed as a foundation for coordination and integration of such interactive robotic systems. The talk will build a bridge from the construction of integrated robotic systems to their evaluation, analysis, and way back. It will outline why we intend to enable our robots to learn by interacting and how this paradigm impacts the design of systems and interaction behaviors.
Gerhard Sagerer
RO-MAN1
2008 The visual active memory perspective on integrated recognition systems
Christian Bauckhage, Sven Wachsmuth, Marc Hanheide, Sebastian Wrede 0001, Gerhard Sagerer, Gunther Heidemann, Helge J. Ritter
Image Vis. Comput.5
2008 Estimating Object Proper Motion Using Optical Flow, Kinematics, and Depth Information
abstract
For the interaction of a mobile robot with a dynamic environment, the estimation of object motion is desired while the robot is walking and/or turning its head. In this paper, we describe a system which manages this task by combining depth from a stereo camera and computation of the camera movement from robot kinematics in order to stabilize the camera images. Moving objects are detected by applying optical flow to the stabilized images followed by a filtering method, which incorporates both prior knowledge about the accuracy of the measurement and the uncertainties of the measurement process itself. The efficiency of this system is demonstrated in a dynamic real-world scenario with a walking humanoid robot.
Jens Schmüdderich, Volker Willert, Julian Eggert, Sven Rebhan, Christian Goerick, Gerhard Sagerer, Edgar Körner
IEEE Trans. Syst. Man Cybern. Part B6
2007 View-adaptive manipulative action recognition for robot companions
abstract
This paper puts forward an approach for a mobile robot to recognize the human's manipulative actions from different single camera views. While most of the related work in action recognition assume a fixed static camera view that is the same for training and testing, such kind of constraints do not apply for mobile robot companions. We propose a recognition scheme that is able to generalize an action model, that has been learned from a very few data items observed from a single camera view, to variant view points and different settings. We tackle the problem of compensating the view dependence of 2D motion models on three different levels. Firstly, we pre-segment the trajectories based on an object vicinity that depends on the camera tilt and object detections. Secondly, an interactive feature vector is designed that represents the relative movements between the human hand and the objects. Thirdly, we propose an adaptive HMM-based matching process that is based on a particle filter and includes a dynamically adjusted scaling parameter that models the systematic error of the view dependency. Finally, we use a two-layered approach for task recognition which decouples the task knowledge from the view dependent primitive recognition. The results of experiments in an office environment show the applicability of this approach.
Zhe Li 0009, Sven Wachsmuth, Jannik Fritsch, Gerhard Sagerer
IROS4
2007 Human-Oriented Interaction With an Anthropomorphic Robot
abstract
A very important aspect in developing robots capable of human-robot interaction (HRI) is the research in natural, human-like communication, and subsequently, the development of a research platform with multiple HRI capabilities for evaluation. Besides a flexible dialog system and speech understanding, an anthropomorphic appearance has the potential to support intuitive usage and understanding of a robot, e.g., human-like facial expressions and deictic gestures can as well be produced and also understood by the robot. As a consequence of our effort in creating an anthropomorphic appearance and to come close to a human- human interaction model for a robot, we decided to use human-like sensors, i.e., two cameras and two microphones only, in analogy to human perceptual capabilities too. Despite the challenges resulting from these limits with respect to perception, a robust attention system for tracking and interacting with multiple persons simultaneously in real time is presented. The tracking approach is sufficiently generic to work on robots with varying hardware, as long as stereo audio data and images of a video camera are available. To easily implement different interaction capabilities like deictic gestures, natural adaptive dialogs, and emotion awareness on the robot, we apply a modular integration approach utilizing XML-based data exchange. The paper focuses on our efforts to bring together different interaction concepts and perception capabilities integrated on a humanoid robot to achieve comprehending human-oriented interaction.
Thorsten Spexard, Marc Hanheide, Gerhard Sagerer
IEEE Trans. Robotics3
2006 Human-like Person Tracking with an Anthropomorphic Robot
abstract
A very important aspect in developing robots capable of human-robot interaction (HRI) is natural, human-like communication. Besides a flexible dialog system and speech understanding an anthropomorphic appearance has many advantages for intuitive usage and understanding of a robot. As a consequence of our effort in creating an anthropomorphic appearance and to come as close as possible to a human-human interaction, we decided to use human-like sensors, i.e., two cameras and two microphones only, not using a laser range finder or omnidirectional camera for tracking persons. Despite the challenge of a limited field of perception, a robust attention system for tracking and interacting with multiple persons simultaneously in real-time was created. Our approach is sufficiently generic to work on robots with varying hardware, as long as stereo audio data and images of a video camera are available. Since the architecture is designed modular with a XML based data exchange we are able to extend the robot's abilities easily
Thorsten Spexard, Axel Haasch, Jannik Fritsch, Gerhard Sagerer
ICRA4
2006 Integration and Coordination in a Cognitive Vision System
abstract
In this paper, we present a case study that exemplifies general ideas of system integration and coordination. The application field of assistant technology provides an ideal test bed for complex computer vision systems including real-time components, human-computer interaction, dynamic 3-d environments, and information retrieval aspects. In our scenario the user is wearing an augmented reality device that supports her/him in everyday tasks by presenting information that is triggered by perceptual and contextual cues. The system integrates a wide variety of visual functions like localization, object tracking and recognition, action recognition, interactive object learning, etc. We show how different kinds of system behavior are realized using the Active Memory Infrastructure that provides the technical basis for distributed computation and a data- and eventdriven integration approach.
Sebastian Wrede 0001, Marc Hanheide, Sven Wachsmuth, Gerhard Sagerer
ICVS4
2006 Towards a multimodal topic tracking system for a mobile robot
abstract
Topics in situated and task oriented communication depend heavily on the given, often changing environment, making the detection of predetermined topics in many cases useless. Detection of non-predefined topics can enhance Human-Robot-Interaction (HRI) in a variety of ways, though. In this paper we propose a way to dynamically determine topics during Human-Robot-Communication using well established techniques such as Latent Semantic Analysis (LSA). The procedure is based on multimodal cues, supporting the view that topics are not simply a property of spoken or written language, but of multimodal situated communication. An online version of the topic detection system has been developed and is currently being tested on our mobile robot BIRON. To demonstrate the feasibility of our approach, we present the results of an evaluation of our system on the BITT corpus. Index Terms: topic tracking, multimodal dialogue, human robot interaction.
Jan Frederik Maas, Britta Wrede, Gerhard Sagerer
INTERSPEECH3
2006 BIRON, where are you? Enabling a robot to learn new places in a real home environment by integrating spoken dialog and visual localization
abstract
An ambitious goal in modern robotic science is to build mobile robots that are able to interact as companions in real world environments. Especially for caretaking of elderly people a system robustly working at private homes is essential, requiring a very natural and human oriented way of communication. Since home environments are usually very individual a first task for a newly acquired robot is to get familiar with its new environment. This paper gives a short overview on how we integrated a vision based localization using the advantages of a very modular architecture and extending a spoken dialog system for online labeling and interaction about different locations. We present results from the integrated system working in a real, fully furnished home environment where it was able to learn the names of different rooms. This system enables us to perform real user studies in future without the need to fall back to Wizard-of-Oz experiments. Ongoing work aims at enabling the robot to take initiative by asking for unknown locations. A future extension is the ability to generalize over features of known rooms to make predictions when encountering unknown rooms
Thorsten Spexard, Shuyin Li, Britta Wrede, Jannik Fritsch, Gerhard Sagerer, Olaf Booij, Zoran Zivkovic, Bas Terwijn, Ben J. A. Kröse
IROS5
2006 Robust Speech Understanding for Multi-Modal Human-Robot Communication
abstract
In order to model complex human robot interaction researchers not only have to consider different tasks but also to handle the complex interplay of different modules of one single robot system. In our context we constructed a robot assistant integrated in a home or office environment. We allow for a fairly natural communication style, which means that the users communicate using speech but are also allowed to use gestures and moreover to use contextual scene knowledge. Against this background, this paper presents a robust speech understanding component for situated human-robot communication. It serves as interface between speech recognition and dialog management. To increase robustness of speech processing it rates the speech recognition output by means of semantic coherence. Even if the recognized word-stream is not grammatically correct the speech understanding component provides semantic interpretations in context of multi-modal input for dialog management. For the understanding process, we designed special semantic concepts grounded to the domain of situated communication. They also provide additional information about the dialog act. A processing mechanism uses these concept units to generate the most likely semantic interpretation of the utterances
Sonja Hüwel, Britta Wrede, Gerhard Sagerer
RO-MAN3
2006 A dialog system for comparative user studies on robot verbal behavior
abstract
In domestic social robot systems the dialog system is often the main user interface. The verbal behavior of such a robot, therefore, plays crucial role in human-robot interaction. Comparative user studies on various verbal behaviors of a robot can effectively contribute to human-robot interaction research. In this paper we present a dialog system that can be easily configured to demonstrate different verbal, initiative-taking behaviors for a robot and, thus, can be used as a platform for such comparative user studies. The pilot study we conducted does not only provide strong evidence for this suitability, but also reveals benefits of comparative studies on a real robot in general
Shuyin Li, Britta Wrede, Gerhard Sagerer
RO-MAN3
2006 BIRON, what's the topic? A Multi-Modal Topic Tracker for improved Human-Robot Interaction
abstract
Creating robots with extendable social skills and interaction capabilities that suffice their operation in the real world with naive users is a very challenging task. In this paper we present a new approach using topic tracking on multi-modal dialogue to provide a mobile robot with a higher level situation awareness in human-robot interaction. The robot is no longer operating in laboratory surroundings, but in its own real world flat. We describe how our topic tracking approach is implemented in this integrated system, operating on verbal speech input. Different modalities like data from video cameras and laser scans are used as additional cues to a semantic understanding and grouping of user utterances into different topics. Both the amount of topics and the according topic names are created dynamically. Evaluating an offline speech corpus demonstrates the suitability of our approach. It is now possible to ask "BIRON, what's the topic?", making the interaction more social
Jan Frederik Maas, Thorsten Spexard, Jannik Fritsch, Britta Wrede, Gerhard Sagerer
RO-MAN5
2005 Combining environmental cues & head gestures to interact with wearable devices
abstract
As wearable sensors and computing hardware are becoming a reality, new and unorthodox approaches to seamless human-computer interaction can be explored. This paper presents the prototype of a wearable, head-mounted device for advanced human-machine interaction that integrates speech recognition and computer vision with head gesture analysis based on inertial sensor data. We will focus on the innovative idea of integrating visual and inertial data processing for interaction. Fusing head gestures with results from visual analysis of the environment provides rich vocabularies for human-machine communication because it renders the environment into an interface: if objects or items in the surroundings are being associated with system activities, head gestures can trigger commands if the corresponding object is being looked at. We will explain the algorithmic approaches applied in our prototype and present experiments that highlight its potential for assistive technology. Apart from pointing out a new direction for seamless interaction in general, our approach provides a new and easy to use interface for disabled and paralyzed users in particular.
Marc Hanheide, Christian Bauckhage, Gerhard Sagerer
ICMI3
2005 Human-style interaction with a robot for cooperative learning of scene objects
abstract
In research on human-robot interaction the interest is currently shifting from uni-modal dialog systems to multi-modal interaction schemes. We present a system for human-style interaction with a robot that is integrated on our mobile robot BIRON. To model the dialog we adopt an extended grounding concept with a mechanism to handle multi-modal in- and output where object references are resolved by the interaction with an object attention system (OAS). The OAS integrates multiple input from, e.g., the object and gesture recognition systems and provides the information for a common representation. This representation can be accessed by both modules and combines symbolic verbal attributes with sensor-based features. We argue that such a representation is necessary to achieve a robust and efficient information processing.
Shuyin Li, Axel Haasch, Britta Wrede, Jannik Fritsch, Gerhard Sagerer
ICMI5
2005 A Flexible Infrastructure for the Development of a Robot Companion with Extensible HRI-Capabilities
abstract
The development of robot companions with nat ural human-robot interaction (HRI) capabilities is a challeng ing task as it requires incorporating various functionalities. Consequently, a flexible infrastructure for controlling module operation and data exchange between modules is proposed, taking into account insights from software system integration. This is achieved by combining a three-layer control architec ture containing a flexible control component with a powerful communication framework. The use of XML throughout the whole infrastructure facilitates ongoing evolutionary develop ment of the robot companion's capabilities.
Jannik Fritsch, Marcus Kleinehagenbrock, Axel Haasch, Sebastian Wrede 0001, Gerhard Sagerer
ICRA5
2005 A multi-modal object attention system for a mobile robot
abstract
Robot companions are intended for operation in private homes with naive users. For this purpose, they need to be endowed with natural interaction capabilities. Additionally, such robots will need to be taught unknown objects that are present in private homes. We present a multi-modal object attention system that is able to identify objects referenced by the user with gestures and verbal instructions. The proposed system can detect known and unknown objects and stores newly acquired object information in a scene model for later retrieval. This way, the growing knowledge base of the robot companion improves the interaction quality as the robot can more easily focus its attention on objects it has been taught previously.
Axel Haasch, Nils Hofemann, Jannik Fritsch, Gerhard Sagerer
IROS4
2005 Humanoid robot platform suitable for studying embodied interaction
abstract
This paper presents the humanoid robot BARTHOC who has been developed to study human-robot interaction (HRI). The main focus of BARTHOC's design was to realize the expression and behavior of the robot to be as human-like as possible. This allows to apply the platform to manifold research and demonstration areas. With his human-like look and mimic possibilities, he differs from other platforms like ASIMO or QRIO, and enables experiments even close to Mori's 'uncanny valley'. The paper describes details of the mechanical and electrical design of BARTHOC together with its PC control interface. Through its humanoid appearance, it can imitate human behavior with its soft- and hardware. Currently, several components for HRI on a mobile robot platform are being ported to BARTHOC. Starting with these components, the robot's human-like appearance enables us to study embodied interaction and to explore theories of human intelligence.
Matthias Hackel, Stefan Schwope, Jannik Fritsch, Britta Wrede, Gerhard Sagerer
IROS5
2005 Interactive object learning for robot companions using mosaic images
abstract
Natural human-robot interaction (HRI) is a key feature of mobile robot companions collaborating with humans. To achieve natural HRI, multiple communication modalities like vision, speech, and gestures have to be utilized. Besides, capabilities to emulate cognitive processes, e.g., object learning and object recognition, are essential. In this work we present a new approach to interactive object learning enabling multi-view object representation. To overcome a robot's limitation of having only one view point, we make use of an iconic memory consisting of previously acquired images. As the relevant scene area is unknown during construction of the iconic memory, a representation in the form of mosaic images is applied. The relevant image patches describing an object referenced by the user are selected through an object attention mechanism. The resulting multi-view object representations improve the flexibility of our interactive approach for object learning.
Birgit Möller 0001, Stefan Posch, Axel Haasch, Jannik Fritsch, Gerhard Sagerer
IROS5
2005 Database driven test case generation for protein?Cprotein docking
abstract
UNLABELLED: We present a method for automatic test case generation for protein-protein docking. A consensus-type approach is proposed processing the whole PDB and classifying protein structures into complexes and unbound proteins by combining information from three different approaches (current PDB-at-a-glance classification, search of complexes by sequence identical unbound structures and chain naming). Out of this classification test cases are generated automatically. All calculations were run on the database. The information stored is available via a web interface. The user can choose several criteria for generating his own subset out of our test cases, e.g. for testing docking algorithms. AVAILABILITY: http://bibiserv.techfak.uni-bielefeld.de/agt-sdp/ CONTACT: [email protected].
Frank Zöllner 0001, Steffen Neumann, Franz Kummert, Gerhard Sagerer
Bioinform.4
2005 Toward automatic video-based whiteboard reading
Markus Wienecke, Gernot A. Fink, Gerhard Sagerer
Int. J. Document Anal. Recognit.3
2005 Modelling expertise for structure elucidation in organic chemistry using Bayesian networks
Michaela Hohenner, Sven Wachsmuth, Gerhard Sagerer
Knowl. Based Syst.3
2004 A Cognitive Vision System for Action Recognition in Office Environments
Christian Bauckhage, Marc Hanheide, Sebastian Wrede 0001, Gerhard Sagerer
CVPR (2)4
2004 Supporting advanced interaction capabilities on a mobile robot with a flexible control system
abstract
Building a mobile service robot for home and office environments that incorporates skilled interaction capabilities is a challenging task. The control system has to consider various demands: first, it has to manage unstructured and dynamic environments. Second, as humans are around, aspects of safety are of particular importance. This implies that the system has to be highly reactive. Third, the robot also has to be capable of carrying out dialogs to be taught or instructed. Altogether, this requires a highly integrated control framework. In this paper we present an agent-based architecture for our mobile robot BIRON in order to realize sophisticated human-robot interaction. The architecture is built in a modular fashion and controlled by a central execution supervisor using an event queue to handle asynchronous events. This execution supervisor contains an augmented finite state machine which is specified in XML and thus is highly generic. Similarly, the communication between all modules is based on XML. The overall system, therefore, is easily maintainable and extensible, as the architecture's design allows us to add new modules to the system without requiring major modifications on existing components.
Marcus Kleinehagenbrock, Jannik Fritsch, Gerhard Sagerer
IROS3
2004 From Image Features To Symbols And Vice Versa - Using Graphs To Loop Data- And Model-Driven Processing In Visual Assembly Recognition
abstract
Graphs and graph matching are powerful mechanisms for knowledge representation, pattern recognition and machine learning. Especially in computer vision their application is manifold. Graphs can characterize relations among image features like points or regions but they may also represent symbolic object knowledge. Hence, graph matching can accomplish recognition tasks on different levels of abstraction. In this contribution, we demonstrate that graphs may also bridge the gap between different levels of knowledge representation. We present a system for visual assembly monitoring that integrates bottom-up and top-down strategies for recognition and automatically generates and learns graph models to recognize assembled objects. Data-driven processing is subdived into three stages: first, elementary objects are recognized from low-level image features. Then, clusters of elementary objects are analyzed syntactically; if an assembly structure is found, it is translated into a graph that uniquely models the assembly. Finally, symbolic models like this are stored in a database so that individual assemblies can be recognized by means of graph matching. At the same time, these graphs enable top-down knowledge propagation: they are transformed into graphs which represent relations between image features and thus describe the visual appearance of the recently found assembly. Therefore, due to model-driven knowledge propagation assemblies may subsequently be recognized from graph matching on a lower computational level and tedious bottom-up processing becomes superfluous.
Christian Bauckhage, Elke Braun, Gerhard Sagerer
Int. J. Pattern Recognit. Artif. Intell.3
2003 A Structural Framework for Assembly Modeling and Recognition
Christian Bauckhage, Franz Kummert, Gerhard Sagerer
CAIP3
2003 Towards Automatic Video-based Whiteboard Reading
abstract
As whiteboards have become a popular tool in meeting rooms, there has been a growing interest in making use of the whiteboard as a user interface for human computer interaction. Therefore, systems based on electronic whiteboards have been developed in order to serve as meeting assistants for e.g. collaborative working. However, as special pens and erasers are required, the natural interaction is restricted. In order to render this communication method more natural it was proposed to retain ordinary whiteboard and pens and to visually observe the writing process using a video camera by Stafford-Fraser and Robinson (1996). In this paper a prototype system for automatic video-based whiteboard reading is presented. The system is designed for recognizing unconstrained handwritten text and is further characterized by an incremental processing strategy in order to facilitate recognizing portions of text as soon as they have been written on the board. We present the methods employed for extracting text regions, pre-processing, feature extraction, and statistical modeling and recognition. Evaluation results on a writer independent unconstrained handwriting recognition task demonstrate the feasibility of the proposed approach.
Markus Wienecke, Gernot A. Fink, Gerhard Sagerer
ICDAR3
2003 Providing the basis for human-robot-interaction: a multi-modal attention system for a mobile robot
abstract
In order to enable the widespread use of robots in home and office environments, systems with natural interaction capabilities have to be developed. A prerequisite for natural interaction is the robot's ability to automatically recognize when and how long a person's attention is directed towards it for communication. As in open environments several persons can be present simultaneously, the detection of the communication partner is of particular importance. In this paper we present an attention system for a mobile robot which enables the robot to shift its attention to the person of interest and to maintain attention during interaction. Our approach is based on a method for multi-modal person tracking which uses a pan-tilt camera for face recognition, two microphones for sound source localization, and a laser range finder for leg detection. Shifting of attention is realized by turning the camera into the direction of the person which is currently speaking. From the orientation of the head it is decided whether the speaker addresses the robot. The performance of the proposed approach is demonstrated with an evaluation. In addition, qualitative results from the performance of the robot at the exhibition part of the ICVS'03 are provided.
Sebastian Lang 0002, Marcus Kleinehagenbrock, Sascha Hohenner, Jannik Fritsch, Gernot A. Fink, Gerhard Sagerer
ICMI6
2003 Computer vision systems
Bernt Schiele, Gerhard Sagerer
Mach. Vis. Appl.2
2002 INDI - intelligent database navigation by interactive and intuitive content-based image retrieval
abstract
We present a content-based image retrieval system, INDI (techniques for Intelligent Navigation in Digital Image databases), that combines the use of low-level pattern recognition techniques, machine learning and an intuitive human-computer interface in order to support intelligent and user-friendly semantic navigation in large image databases. To keep independence from specific image domains and to encompass different search tasks, the system is highly modular and contains a hierarchical mechanism for the adaptive reweighting of similarity measures implemented by dynamically reloadable modules at different semantic levels.
Tanja Kämpfe, Thomas Käster, Michael Pfeiffer 0003, Helge J. Ritter, Gerhard Sagerer
ICIP (3)5
2002 Evaluating Integrated Speech- and Image Understanding
abstract
The capability to coordinate and interrelate speech and vision is a virtual prerequisite for adaptive, cooperative, and flexible interaction among people. It is therefore fair to assume that human-machine interaction, too, would benefit from intelligent interfaces for integrated speech and image processing. We first sketch an interactive system that integrates automatic speech processing with image understanding. Then, we concentrate on performance assessment which we believe is an emerging key issue in multimodal interaction. We explain the benefit of time scale analysis and usability studies and evaluate our system accordingly.
Christian Bauckhage, Jannik Fritsch, Katharina J. Rohlfing, Sven Wachsmuth, Gerhard Sagerer
ICMI5
2002 Multi-modal human-machine communication for instructing robot grasping tasks
abstract
A major challenge for the realization of intelligent robots is to supply them with cognitive abilities in order to allow ordinary users to program them easily and intuitively. One approach to such programming is teaching work tasks by interactive demonstration. To make this effective and convenient for the user, the machine must be capable of establishing a common focus of attention and be able to use and integrate spoken instructions, visual perception, and non-verbal clues like gestural commands. We report progress in building a hybrid architecture that combines statistical methods, neural networks, and finite state machines into an integrated system for instructing grasping tasks by man-machine interaction. The system combines the GRAVIS-robot for visual attention and gestural instruction with an intelligent interface for speech recognition and linguistic interpretation, and a modality fusion module to allow multi-modal task-oriented man-machine communication with respect to dextrous robot manipulation of objects.
Patrick C. McGuire, Jannik Fritsch, Jochen J. Steil, Frank Röthling, Gernot A. Fink, Sven Wachsmuth, Gerhard Sagerer, Helge J. Ritter
IROS7
2002 Combining acoustic and articulatory feature information for robust speech recognition
Katrin Kirchhoff, Gernot A. Fink, Gerhard Sagerer
Speech Commun.3
2001 Incorporating Process Knowledge into Object Recognition for Assemblies
Elke Braun, Jannik Fritsch, Gerhard Sagerer
ICCV3
2001 Video-Based On-line Handwriting Recognition
abstract
The use of handwriting provides a natural way of interacting with small portable computers. However, in order to capture handwritten text. online, special input devices are necessary. Therefore, M.E. Munich & P. Perona (1996) proposed to use visual input for pen-based computers. Writing can then be performed on ordinary paper, and pen trajectories are automatically extracted from image sequences recorded during the writing process. On the basis of this work, we developed a complete video-based online handwriting recognition system. We will present the techniques applied for pen tracking, pre-processing, feature extraction, and statistical modeling and recognition. Evaluation results on a writer-independent unconstrained handwriting recognition task demonstrate that the inherent limitations of the video-based approach can be compensated using robust modeling combined with adaptation techniques.
Gernot A. Fink, Markus Wienecke, Gerhard Sagerer
ICDAR3
2001 An investigation of modelling aspects for ratedependent speech recognition
abstract
For the modelling of speech rate variation in speech recognition many approaches have been suggested. However, the training of speech-rate dependent models has by far received most of the attention. In order to investigate problematic aspects related with the classification of the speech data which represents one of the major problems of these approaches extensive experiments were carried out on a German corpus of read speech. The results indicate that while the kind of the model-driven speech-rate measure is only of minor importance a data-driven classification of the speech data significantly improves the performance of rate-dependent models. Further results suggest a detailed modelling of speech rate based on more general models. This means that it might be possible to model speech rate adaptation by means of a transformation based on a continuous measure.
Britta Wrede, Gernot A. Fink, Gerhard Sagerer
INTERSPEECH3
2000 Conversational speech recognition using acoustic and articulatory input
abstract
The combination of multiple speech recognizers based on different signal representations is increasingly attracting interest in the speech community. In previous work we presented a hybrid speech recognition system based on the combination of acoustic and articulatory information which achieved significant word error rate reductions under highly noisy conditions on a small-vocabulary numbers recognition task. In this study we extend this approach to large-vocabulary conversational speech recognition using the Gaussian mixture acoustic modeling paradigm. We demonstrate that the articulatory input representation we propose contains information which is complementary to that provided by standard MFCC features, and that their combination can significantly reduce the word error rate on conversational speech. Various combination strategies (feature-level, state-level and word-level combination) are compared and evaluated.
Katrin Kirchhoff, Gernot A. Fink, Gerhard Sagerer
ICASSP3
2000 Detecting Assembly Actions by Scene Observation
abstract
We present a fast and reliable method to analyse an image sequence where two human hands perform assembly actions. Our classifier to detect the initial skin coloured regions is based on a polynomial classifier of sixth degree. For each region a judgement is calculated using two cues. The first cue is motion information obtained from a difference image. The second cue is obtained from confidence mapping performed on the output of the polynomial classifier. All resulting regions with a high judgement indicating movement are tracked using Kalman filters. Based on the trajectories of the Kalman filters action hypotheses are generated. The hypotheses are verified through a detailed analysis using a Fourier transformation on the derivatives of the trajectories.
Jannik Fritsch, Frank Lömker, Markus Wienecke, Gerhard Sagerer
ICIP4
2000 Towards an Integrated Framework for Contour-Based Grouping and Object Recognition Using Markov Random Fields
abstract
We present an integrated approach for contour-based grouping and object recognition. Domain knowledge and domain independent grouping laws are combined in a multi-layered Markov random field (MRF) framework. It provides a basis for propagating top-down knowledge between different processing cues or input modalities. Additionally, the domain dependent MRF-layer can be used in order to evaluate the grouping process with regard to relevant contours for object recognition. Initial results show the approach to be adequate for complex scenes and partially occluded objects.
Daniel Schlüter, Sven Wachsmuth, Gerhard Sagerer, Stefan Posch
ICIP3
2000 Integration of Regions and Contours for Object Recognition
abstract
We present an integrated approach combining region and contour-based techniques to enhance both segmentation and recognition processes. This cue integration operates on the level of contour-based groups and complete regions, which are matched to reflect a common cause in the image (and thus in the scene). Additionally, we realize a top-down scheme controlling the segmentation process on the basis of unresolvable image areas and the integration of processes on different time scales using a region memory. Results for real images are given showing distinct improvements of recognition results.
Daniel Schlüter, Franz Kummert, Gerhard Sagerer, Stefan Posch
ICPR3
2000 A hybrid speech recognizer combining HMMs and polynomial classification
abstract
In this paper, we present a hybrid speech recognizer combining Hidden Markov Models (HMMs) and a polynomial classifier. In our approach the emission probabilities are not modeled as a mixture of Gaussians but are calculated by the polynomial classifier. However, we do not apply the classifier directly to the feature vector but we make use of the density values of L Gaussians clustering the feature space. That means we model the emission probability as a polynomial of Gaussian distributions of n-th degree. As most of these density values are approximately zero for a single feature vector the calculation of a polynomial can be done very efficiently. The usefulness of this hybrid approach was successfully tested on a large conversational speech recognition task. 1.
Franz Kummert, Gernot A. Fink, Gerhard Sagerer
INTERSPEECH3
2000 Influence of duration on static and dynamic properties of German vowels in spontaneous speech
abstract
Changes in speech rate severely affect the performance of continuous speech recognition systems. In order to better understand the underlying effects of speech rate changes an analysis was carried out on the influence of duration on the spectral properties of vowels in a large German corpus of spontaneous speech. The results show a strong centralisation effect of the vowel formant frequencies due to shorter duration while the formant movements are only slightly affected. The data suggest that the movement velocity is not changed in vowels with a limited duration. As the means of the on- and offset frequencies also remain stable only the middle part of the vowels are affected by the centralisation effect. These results are discussed in the light of the modelling of varying speech rate in automatic speech recognition systems. 1.
Britta Wrede, Gernot A. Fink, Gerhard Sagerer
INTERSPEECH3
2000 Bayesian reasoning on qualitative descriptions from images and speech
Gudrun Socher, Gerhard Sagerer, Pietro Perona
Image Vis. Comput.2
1999 Multilevel Integration of Vision and Speech Understanding Using Bayesian Networks
Sven Wachsmuth, Hans Brandt-Pook, Gudrun Socher, Franz Kummert, Gerhard Sagerer
ICVS5
1999 A multi-directional multiple-path recognition scheme for complex objects applied to the domain of a wooden toy kit
Elke Braun, Gunther Heidemann, Helge J. Ritter, Gerhard Sagerer
Pattern Recognit. Lett.4
1998 Hybrid object recognition in image sequences
abstract
We present a hybrid approach attaching probabilistic formalisms, as artificial neural networks or hidden Markov models, to concepts of a semantic network for a robust and efficient detection of objects. Additionally, an efficient processing strategy for image sequences is outlined which propagates the structural results of the semantic network as an expectation for the next image. This method allows one to produce linked results over time supporting the recognition of events and actions.
Franz Kummert, Gernot A. Fink, Gerhard Sagerer, Elke Braun
ICPR3
1998 A HMM-based recognition system for perceptive relevant pitch movements of spontaneous German speech
abstract
This paper presents an HMM-based recognition system for perceptive relevant pitch movements of spontaneous German speech. The pitch movements are defined according to the perceptively and phonetically motivated IPO-approach to intonation. For recognition we use a hybrid approach combining polynomial classification with Hidden Markov Modelling. The recognition is based only on the speech signal, its fundamental frequency and eleven derived features. We evaluate the system on a speaker independant recognition task. 1 Introduction In current speech recognition systems, usually, no prosodic information is used. However, it is a wellknown assumption that prosody can contribute useful information to enhance speech recognition and understanding processes [6]. While for speech recognition, there is no doubt that recognition processes should be resulting in a sequence of word hypotheses for prosody recognition the chosen units depend on several competing linguistic theories. Although their ad...
Christel Brindöpke, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
ICSLP4
1998 Estimation and filtering of potential protein-protein docking positions
abstract
MOTIVATION: Software systems predicting automatically whether and how two proteins may interact are highly desirable, both for understanding biological processes and for the rational design of new proteins. As a part of a future complete solution to this problem, a bundle of programs is presented designed (i) to estimate initial docking positions for a given pair of docking candidates, (ii) to adjust them, and (iii) to filter them, thus preparing more detailed computations of free energies. RESULTS: The system is evaluated on a test set of 51 co-crystallized complexes aiming at redocking the subunits. It works completely automatically and the evaluation is performed using one single set of parameters for all complexes in the test set. The number of solutions is fixed to 50 positions with a median CPU time of 26 min. For 30 complexes, these contain a near-correct solution with root mean square deviation ( RMSD ) </=5.0 A, which is ranked first in five cases. For all complexes, the best solution is scored on rank 16 as the worst case, and has a median RMSD of 4.3 A. Alternatively to this initial estimation of docking positions, a global sampling of rotations was tested. Whereas this yields top-ranked solutions with RMSD </=3.0 A for all 51 complexes, the median CPU time increases to 11 h. This shows that this blind sampling is not feasible for most applications. AVAILABILITY: The system and its components are available on request from the authors. CONTACT: [email protected] or [email protected]
Friedrich Ackermann, Grit Herrmann, Stefan Posch, Gerhard Sagerer
Bioinform.4
1997 Using Markov Random Fields for Contour-Based Grouping
abstract
To overcome fragmentation of an initial contour-based segmentation and to organize contour segments into image primitives on a higher level of abstraction, regularities of the image data are exploited using ideas from the Gestalt psychology. First, groups are hypothesized within a hierarchy based on local evidence only, where the criteria are derived from a hand labelled training set. These hypotheses are subsequently judged in a global context using a Markov random field to derive a global interpretation. Examples of results for real data are given.
A. Mabbmann, Stefan Posch, Gerhard Sagerer, Daniel Schlüter
ICIP (2)3
1997 An environment for the labelling and testing of melodic aspects of speech
abstract
In this paper, we present an environment for labelling and testing of melodic aspects of spoken language. The environment has three modes of application: First, the environment provides labelling facilities for a model-based melodic description for German. Second, it supports a language independent pre-theoretical description of speech melody allowing the development of new melodic categories. Third, our test bed can be used to generate speech samples with controlled melodic parameters for further use in perception experiments. The melodic description facilities (model-based, pre-theoretical) are supported by visual and audible feedback allowing a step-bystep refinement of the melodic description in question. 1 Introduction Melodic aspects of speech are related with several linguistic and extralinguistic phenomenon. Linguistic aspects are for example the realization of accents or the marking of boundaries. Extralinguistic aspects concern e.g. speaker related qualities like emotions. T...
Christel Brindöpke, Arno Pahde, Franz Kummert, Gerhard Sagerer
EUROSPEECH4
1996 A Hybrid Object Recognition Architecture
Gunther Heidemann, Franz Kummert, Helge J. Ritter, Gerhard Sagerer
ICANN4
1996 Talking about 3D scenes: integration of image and speech understanding in a hybrid distributed system
abstract
We present a hybrid system that integrates speech and image understanding. Given spoken references, it is able to identify objects of a 3D scene perceived via a stereo camera. Central to our approach is the extraction of qualitative object features and spatial scene properties from acoustic and visual data. The interaction of the understanding processes is performed using a procedural semantic network that interfaces with signal recognition and reconstruction modules, thus integrating semantic, neural and Bayesian networks and Hidden Markov Models.
Gudrun Socher, Gerhard Sagerer, Franz Kummert, Thomas Fuhr 0002
ICIP (2)2
1996 A robust dialogue system for making an appointment
abstract
A complete dialogue system within the task domain of making an appointment is presented.It is based on a semantic network representation of linguistic knowledge and a word recognition system that communicates with the interpretation component bidirectionally.System robustness is achieved using a special metascore that evaluates the advance of the linguistic interpretation.
Hans Brandt-Pook, Gernot A. Fink, Bernd Hildebrandt, Franz Kummert, Gerhard Sagerer
ICSLP5
1996 Evaluation of spoken language understanding and dialogue systems
abstract
A spoken language understanding and dialogue system in the domain of appointment schedule is presented.The system is capable of understanding complex times, e.g. it correctly combines discontinuous constituents and resolves ambiguities.A distributed representation of surface structure models and an incremental semantic analysis is used to manage the complexity.An elaborate evaluation of the system based on measurements of accuracy was carried out.Our approach combines pattern recognition with linguistic aspects forming a system of measurement consisting of word accuracy, constituent accuracy, and concept accuracy.
Bernd Hildebrandt, Heike Rautenstrauch, Gerhard Sagerer
ICSLP3
1996 Incremental generation of word graphs
Gerhard Sagerer, Heike Rautenstrauch, Gernot A. Fink, Bernd Hildebrandt, A. Jusek, Franz Kummert
ICSLP1
1995 Segmentation of molecular surfaces based on their convex hull
abstract
Docking of two or more proteins to protein complexes is based on the complementarity of shape and chemical attributes of the involved molecular surfaces. A geometry-based segmentation technique is proposed, generating characteristically shaped regions of molecular surfaces, used to predict possible docking sites of proteins. Thus, an unrestricted search in 6D configuration space is avoided. The segmentation algorithm is based on the convex hull of the molecular surface. Its implementation has been shown to produce stable segmentations of the 3D surfaces, well approximating the true contact sides of the involved proteins. Making as few as possible assumptions on the underlying surfaces this technique might be useful to other applications as well.
R. Meier, Friedrich Ackermann, Grit Herrmann, Stefan Posch, Gerhard Sagerer
ICIP (3)5
1995 Statistical classification and segmentation of biomolecular surfaces
abstract
A preprocessing algorithm is presented, integrating the computation of biomolecular surfaces, the calculation of geometrical and chemical features on each surface point, and a complex segmentation into characteristic surface regions. Calculated regions cover the whole computed protein surface and are a feature of shape, roughness and hydrophobicity on the appropriate surface area. Such regions can be paired in complementary parts, used to predict possible docking sites of proteins. The segmentation algorithm of docking sites proceeds in two steps: first, surface points are classified with a vector quantizer. Second, connected areas of surface points are computed by a complex region growing technique. The detected regions were compared to the real docking sites calculated from complexes with known 3D structures with good results of approximation.
Christoph Schillo, Grit Herrmann, Friedrich Ackermann, Stefan Posch, Gerhard Sagerer
ICIP (3)5
1995 Detection of unknown words and its evaluation
abstract
Especially in recognition of spontaneous speech it is necessary to be able to cope with the occurance of unknown words. We present an approach to unknown word detection integrated with speech recognition. Though several schemes exist for the evaluation of detection algorithms all suffer from some deficiencies. Therefore, we also present a new evaluation measure called detection accuracy which is similar to the widely accepted word accuracy. We applied this measure to the evaluation of our approach on a large spontaneous speech recognition task from the German Verbmobil project.
A. Jusek, Gernot A. Fink, Franz Kummert, Heike Rautenstrauch, Gerhard Sagerer
EUROSPEECH5
1995 Generation of language models using the results of image analysis
abstract
We present a new approach towards using contextual information to enhance speech recognition and understanding. Dynamically inferred knowledge about the context is used in addition to the static linguistic and domain specific knowledge. Based on the results of image analysis of a given scene language models for constituents of possible utterances concerning that scene are generated.
Uta Naeve, Gudrun Socher, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
EUROSPEECH5
1995 Protein Docking: Combining Symbolic Descriptions of Molecular Surfaces and Grid-Based Scoring Functions
Friedrich Ackermann, Grit Herrmann, Franz Kummert, Stefan Posch, Gerhard Sagerer, Dietmar Schomburg
ISMB5
1994 A close high-level interaction scheme for recognition and interpretation of speech
abstract
The vast majority of speech understanding systems suffers from a bottleneck between the recognition and the interpretation components. Normally, only a relatively small set of word hypotheses is passed from the recognizer and no flow of information in the opposite direction is even possible. We propose an interaction scheme that tries to overcome many of the disadvantages of traditional systems. It makes use of the possibility to process abstract constituents in our word recognizer and pass them back as complex hypotheses. Predictions that define the complex analysis goal of the recognition can be derived dynamically during the interpretation of an utterance. A left-to-right processing in both recognition and interpretation makes an incremental analysis possible.
Gernot A. Fink, Franz Kummert, Gerhard Sagerer
ICSLP3
1994 Understanding of time constituents in spoken language dialogues
abstract
The analysis and interpretation of time constituents is a rather complex enterprise, since diverse time constituents are distributable in variable positions within an utterance. The first step in order to manage this complexity and variability is the syntactic analysis at phrase structure level. Within a utterance each time constituent is analyzed independently and tested for its syntactic coherence. A semantic interpretation of the time constituent has to follow. The second step consists of the analysis and interpretation at sentence structure level. The time interpretations need to be tested for consistency and merged into a single representation. Here it is usually possible to resolve ambiguities. As a last step, the interpretations of time constituents have to be merged at dialogue level. Although the system asks the user for verification of the time computed, users do seldom just reply ’ or ’. Mostly they add new information about time, and sometimes users correct the system’s interpretation without explicit negation. Thus, the merging of time constituents at dialogue level becomes a rather complex issue
Bernd Hildebrandt, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
ICSLP4
1994 A Speech Understanding and Dialog System with a Homogeneous Linguistic Knowledge Base
abstract
This article presents the speech understanding and dialog system EVAR. All levels of linguistic knowledge are used both to control the analysis process and for the interpretation of an utterance. All kinds of knowledge are integrated in a homogeneous knowledge base. The control algorithm used for the analysis is defined within the representation scheme and does not depend on the application. One of the aims of EVAR is to develop a system structure where linguistic and nonlinguistic expectations could be used not only for the interpretation but also as predictions for the recognition process.>
Marion Mast, Franz Kummert, Ute Ehrlich, Gernot A. Fink, Thomas Kuhn 0002, Heinrich Niemann, Gerhard Sagerer
IEEE Trans. Pattern Anal. Mach. Intell.7
1993 Robust interpretation of speech
abstract
In all fields of pattern recognition there may arise situations where severe distortions of the sensor data or errors of processing prevent a successful analysis. This is especially true in speech recognition. Therefore, a speech understanding system can not rely on being able to interpret the whole or even a given fixed percentage of the input data. Rather a dynamic criterion has to be applied to decide when the analysis has produced the best results obtainable from the input data. We propose an appropriate criterion and show how the requirements for robust interpretation of speech can be met in a dialog system. Keywords: Speech Understanding, Robustness, Search 1. INTRODUCTION In speech understanding systems a main problem is how to deal with poor segmentation results. How can it be decided that no or at least only a partial interpretation of the input data is possible? Let us consider the following two dialogs where the user's utterance (U), the word recognition result (R), and the...
Gernot A. Fink, Franz Kummert, Gerhard Sagerer, Bernd Seestaedt
EUROSPEECH3
1993 Speech recognition using semantic hidden Markov networks
abstract
Semantic Hidden Markov Networks (SHMNs) were first introduced in [2] as a new technique of interfacing between linguistic analysis and word recognition in speech understanding systems. The main difference between SHMNs and the use of traditional language models is that SHMNs always refer to a linguistic concept and impose the linguistic structure as closely as possible on its acoustic counterpart --- a hierarchically structured HMM. Normally the result of decoding a HMM is merely the sequence of best fitting elementary acoustic concepts, e.g. phonemes or words. Taking into account the structure of the recognition task a structured instance can be computed. This complex acoustic instance can easily be transformed into a linguistic instance by a recursive computation but without any searching. In this paper we present an algorithm for generating linguistic instances from word recognition results based on SHMNs. Additionally, we present recognition results obtained when evaluating a set o...
Gernot A. Fink, Franz Kummert, Gerhard Sagerer, Ernst Günter Schukat-Talamazzini
EUROSPEECH3
1993 Modeling of time constituents for speech understanding
abstract
The analysis and interpretation of time constituents is important for most applications of speech understanding systems. Problems can be caused by the varying distribution of constituents. A basic set of time constituents were found in a corpus of domain specific (train schedule) utterances. A distributed representation of surface structure models and an incremental semantic analysis is used to manage the complexity. The knowledge base of the speech understanding system that provides the framework for the analysis and interpretation of time constituents uses the semantic network language ERNEST. Keywords: Speech Understanding, Time Constituents 1. INTRODUCTION In a speech understanding system there are several domains of analysis and interpretation. Firstly, the system has to recognize single words in a torrent of speech sounds. Secondly, the system combines words to constituents, i.e. it performs a syntactic analysis. It also has to reconstruct the meaning of the utterance in questio...
Bernd Hildebrandt, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
EUROSPEECH4
1993 Control and explanation in a signal understanding environment
Franz Kummert, Heinrich Niemann, Reinhard Prechtel, Gerhard Sagerer
Signal Process.4
1992 A problem-independent control algorithm for image understanding
abstract
Presents a problem-independent control algorithm for image understanding providing both data-driven and model-driven control structures. By an easy combination of these structures any mixed strategy can be achieved. The basis is a framework for the representation of declarative and procedural knowledge using a semantic network.>
Franz Kummert, Gerhard Sagerer, Heinrich Niemann
ICPR (1)2
1992 Semantic hidden Markov networks
abstract
Although much effort has been put into speech understanding systems there still exists a rather wide gap between acoustic recognition and linguistic interpretation. We propose a formalism for an extremely close interaction of acoustic recognition and higher level analysis. Instead of a strict horizontal interface at the level of hypothesized word sequences or lattices, a vertical interface to the acoustic component is used that can be accessed from linguistic concepts of any degree of abstraction. As the linguistic knowledge is represented in the formalism of Semantic Networks and acoustic recognition is based on Hidden Markov Models the close interaction between the two components was termed Semantic Hidden Markov Networks. 1 INTRODUCTION Because of the high degree of uncertainty in the recognition of spoken language it is very important to exploit any possible predictions and restrictions to guide acoustic analysis as well as linguistic interpretation. Within traditional systems the...
Gernot A. Fink, Franz Kummert, Gerhard Sagerer, Ernst Günter Schukat-Talamazzini, Heinrich Niemann
ICSLP3
1990 ERNEST: A Semantic Network System for Pattern Understanding
abstract
The authors give a detailed account of a system environment for treating general problems of image and speech understanding. A framework for representing declarative and procedural knowledge based on a suitable definition of a semantic network is provided. The syntax and semantics of the network are clearly defined. The pragmatics of the network in its use for pattern understanding are defined by several rules which are problem-independent. This allows one to formulate problem-independent control algorithms. Complete software environments are available to handle the described structures. The general applicability of the network system is demonstrated by short descriptions of three applications from different task domains.>
Heinrich Niemann, Gerhard Sagerer, Stefan Schröder, Franz Kummert
IEEE Trans. Pattern Anal. Mach. Intell.2
1988 An expert system architecture and its application to the evaluation of scintigraphic image sequences
abstract
A general framework for expert systems in the area of image understanding is described. It represents declarative knowledge in an associative network, allows the attachment of arbitrary procedural knowledge, and is used by a problem-independent control algorithm. This framework is used to implement a complete system for the diagnostic evaluation of scintigraphic images of the heart. The input to the system is an image sequence which is first preprocessed and segmented to obtain the left ventricle and its sectors. The contours of the ventricle and the sectors are the basis for knowledge-based analysis of the images. Analysis proceeds from the contour segments via objects and their motions to diagnostic descriptions of the image sequence. Experiments show that the system successfully completes analysis, and in particular does not make wrong diagnostic suggestions.>
Gerhard Sagerer, Heinrich Niemann
CBMS1
1988 On the accuracy of optical flow computation using global optimization
abstract
The accuracy of optical flow computation is investigated using gradient methods and global optimization. It considered the influence of systematic and statistical errors on the constraint equation, and of systematic and statistical errors on the iterative solution of the optimization of the continuity constraint. Approaches to error reduction are suggested. Experimental results are given using synthetic images with exactly known properties demonstrating the efficiency of error reduction techniques.>
Heinrich Niemann, J. Arnold, Gerhard Sagerer
ICPR3
1988 A flexible control strategy with multilevel judgements for a knowledge based speech understanding system
abstract
A flexible control strategy for a speech understanding and dialog system is presented. The system is organized around a homogeneous knowledge base. The used knowledge representation language is based on semantic networks. It does not only cover declarative but also procedural knowledge. The inference processes can be described by five problem-independent rules which form the procedural semantics of the network language. The control strategy itself is the well-known A*-algorithm, but for scoring purposes complex judgement vectors are used. It is shown that these vectors are admissible for the control algorithm and that they reflect the quality, the certainty, and the priority of both bottom-up and top-down hypotheses. So far, the system has been tested in a rough prototype version, and an efficient realization is being built.>
Gerhard Sagerer, Ute Ehrlich, Franz Kummert, Heinrich Niemann, Ernst Günter Schukat-Talamazzini
ICPR1
1988 A Knowledge Based speech Understanding System
abstract
This article describes the work in the development of the speech understanding and dialog system EVAR. The relevant knowledge bases containing the raw linguistic knowledge and the preprocessors converting this to the specialized form needed by the processing algorithms are treated. Processing so far covers the level of the speech signal up to the level of pragmatic analysis. Some topics of ongoing and future work are mentioned briefly.
Heinrich Niemann, Astrid Brietzmann, Ute Ehrlich, Stefan Posch, Peter Regel-Brietzmann, Gerhard Sagerer, Richard Salzbrunn, Ernst Günter Schukat-Talamazzini
Int. J. Pattern Recognit. Artif. Intell.6
1988 Control Strategies in a Hierarchical Knowledge Structure
abstract
Two control strategies are presented as working on a hierarchical knowledge structure based on a semantic network. The control algorithms cover strict top-down control and a bidirectional control which is a mixture of top-down (model driven) and bottom-up (data driven) analysis. The knowledge used by the algorithm is represented in a semantic network. Besides the network some other knowledge sources may be generated automatically to direct the analysis and limit the search space. The approach was used successfully in image and speech understanding.
Heinrich Niemann, Gerhard Sagerer, Wolfgang Eichhorn
Int. J. Pattern Recognit. Artif. Intell.2
1988 Automatic interpretation of medical image sequences
Gerhard Sagerer
Pattern Recognit. Lett.1
1986 Representation of a continuous speech understanding and dialog system in a homogeneous semantic net achitecture
abstract
Continuous speech understanding systems require the integration of different knowledge sources on different levels of abstraction. In this paper a representation scheme based on semantic networks is presented which allow a homogeneous architechture and a global processing strategy for a complete system.
Heinrich Niemann, Astrid Brietzmann, Ute Ehrlich, Gerhard Sagerer
ICASSP4
1985 A Knowledge Based System for Analysis of Gated Blood Pool Studies
abstract
A system for obtaining a complete diagnostic description of an image sequence taken in nuclear medicine from the human heart has been developed, implemented, and tested. The knowledge about these images is represented in a semantic net, conclusions are drawn by a production rule approach, and scoring of alternative diagnoses is based on fuzzy membership functions. On the low level, image pixels are smoothed and organ contours are extracted; these are the input for the high level processing. Tests with several image sequences gave correct descriptions as compared to the diagnosis of a physician.
Heinrich Niemann, Horst Bunke, Ingrid Hofmann, Gerhard Sagerer, Friedrich Wolf, Herbert Feistel
IEEE Trans. Pattern Anal. Mach. Intell.4
1982 An Experimental Study of Some Algorithms for Unsupervised Learning
abstract
Three well-known algorithms for unsupervised learning using a decision-directed approach are the random labeling of patterns according to the estimated a posteriori probabilities, the classification according to the estimated a posteriori probabilities, and the iterative solution of the maximum likelihood equations. The convergence properties of these algorithms are studied by using a sample of about 10 000 handwritten numerals. It turns out that the iterative solution of the maximum likelihood equations has the best properties among the three approaches. However, even this one fails to yield satisfactory results if the number of unknown parameters becomes large, as is usually the case in realistic problems of pattern recognition.
Heinrich Niemann, Gerhard Sagerer
IEEE Trans. Pattern Anal. Mach. Intell.2