Sven Wachsmuth

dblp:81/6660 · DBLP profile ↗
← Back
56ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0001-5371-7214ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 17 · 3 since 2021Systems, architecture and hardware · 13Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5
YearPublicationVenuePosition
2026 Particle-Driven Robot Placement: A ROS 2 Plugin-Based Framework
abstract
The present work introduces a ROS 2 plugin-based framework for robot placement in dynamic 2D environments, aimed at Human–Robot Interaction (HRI) scenarios such as approaching and person transfer, while remaining applicable to broader service robot tasks. Candidate poses where the robot could be placed are represented as particles and evaluated against a multi-constraint specification composed of critical (hard) and non-critical (soft/scoring) constraints. Constraints are declared in XML using a concise syntax that supports logical (∧, ∨, ¬) and algebraic (∘, +) operators. The particles distribution adapts online to environmental changes via weighting and resampling, yielding feasible goal poses that respect specified constraints. The implementation is written in C++, integrates natively with ROS 2 through a plugin architecture, and exposes an action server interface for seamless use within higher-level planners.
Jesus Enrique Aleman Gallegos, Sven Wachsmuth
HRI2
2025 AURORA: A Platform for Advanced User-Driven Robotics Online Research and Assessment
abstract
AURORAis a software platform, that facilitates scalable deployment of robotic simulations over the web for the Human-Robot Interaction (HRI) community. As robotics is becoming increasingly important in various disciplines, there is a growing need for accessible and scalable research methods. Traditional experiments often require expensive hardware and in-person participation, limiting accessibility and participant diversity. Our platform allows researchers from different fields to easily provide HRI experiences by deploying online studies with robotic simulations paired with customizable surveys, allowing end users worldwide to interact with these simulations. Our platform is entirely open source and can be hosted locally, providing flexibility and control of the research environment. Since AURORA is implemented with Docker, it is platform-independent. By offering a user-friendly interface that can be deployed and used without extensive technical expertise, our platform reduces costs, increases participant diversity, and improves the reproducibility of research in the HRI community.
Phillip Richter, Markus Rothgänger, Arthur Maximilian Noller, Heiko Wersing, Sven Wachsmuth, Anna-Lisa Vollmer
HRI5
2024 A Hybrid Collaboration Design for a Large Scale Virtual Reality Training Environment to Fulfil the Belongingness Needs of Maslow's Theory
Yusra Tehreem, Thies Pfeiffer, Sven Wachsmuth
EuroXR3
2021 Recipe Enrichment: Knowledge Required for a Cooking Assistant
abstract
Neumann N, Wachsmuth S. Recipe Enrichment: Knowledge Required for a Cooking Assistant. In: Proceedings of the 13th International Conference on Agents and Artificial Intelligence. SCITEPRESS - Science and Technology Publications; 2021: 822-829.
Nils Neumann, Sven Wachsmuth
ICAART (2)2
2020 A Three-Site Reproduction of the Joint Simon Effect with the NAO Robot
abstract
The generalizability of empirical research depends on the reproduction of findings across settings and populations. Consequently, generalizations demand resources beyond that which is typically available to any one laboratory. With collective interest in the joint Simon effect (JSE) -- a phenomenon that suggests people work more effectively with humanlike (as opposed to mechanomorphic) robots -- we pursued a multi-institutional research cooperation between robotics researchers, social scientists, and software engineers. To evaluate the robustness of the JSE in dyadic human-robot interactions, we constructed an experimental infrastructure for exact, lab-independent reproduction of robot behavior. Deployment of our infrastructure across three institutions with distinct research orientations (well-resourced versus resource-constrained) provides initial demonstration of the success of our approach and the degree to which it can alleviate technical barriers to HRI reproducibility. Moreover, with the three deployments situated in culturally distinct contexts (Germany, the U.S. Midwest, and the Mexico-U.S. Border), observation of a JSE at each site provides evidence its generalizability across settings and populations.
Megan K. Strait, Florian Lier, Jasmin Bernotat, Sven Wachsmuth, Friederike Eyssel, Robert L. Goldstone, Selma Sabanovic
HRI4
2020 Weakly-Supervised Learning for Multimodal Human Activity Recognition in Human-Robot Collaboration Scenarios
abstract
The ability to synchronize expectations among human-robot teams and understand discrepancies between expectations and reality is essential for human-robot collaboration scenarios. To ensure this, human activities and intentions must be interpreted quickly and reliably by the robot using various modalities. In this paper we propose a multimodal recognition system designed to detect physical interactions as well as nonverbal gestures. Existing approaches feature high post-transfer recognition rates which, however, can only be achieved based on well-prepared and large datasets. Unfortunately, the acquisition and preparation of domain-specific samples especially in industrial context is time consuming and expensive. To reduce this effort we introduce a weakly-supervised classification approach. Therefore, we learn a latent representation of the human activities with a variational autoencoder network. Additional modalities and unlabeled samples are incorporated by a scalable product-of-expert sampling approach. The applicability in industrial context is evaluated by two domain-specific collaborative robot datasets. Our results demonstrate, that we can keep the number of labeled samples constant while increasing the network performance by providing additional unprocessed information.
Clemens Pohlt, Thomas Schlegl, Sven Wachsmuth
IROS3
2019 See and Be Seen - Rapid and Likeable High-Definition Camera-Eye for Anthropomorphic Robots
abstract
While many social robots already include carefully designed robotic faces, functional robot eyes that meet human expectations are still an open challenge. As a consequence, many robots either have cameras separated from their robot eyes or active camera heads missing an anthropomorphic face. In this paper, we propose a new robot eye that is integrated in the anthropomorphic robot head Floka and fulfills a similar technical specification as a human eye including zero backlash, an increased range of motion, high velocities and accelerations, an integrated high resolution camera, and fast actuated eyelids. The robot eye is built using state-of-the-art off-the-shelf components and the CAD model of our prototype is available free of charge on request for non-commercial applications. We evaluate the technical properties of the robot eye and show that it meets and partially outperforms human eye movements and saccades.
Simon Schulz, Sebastian Meyer zu Borgsen, Sven Wachsmuth
ICRA3
2019 An Evaluation of Robot-to-Human Handover Configurations for Commercial Robots
abstract
The handover of objects is a fundamental task in robot-human interaction. The nature of the handover can be varied by adjusting various parameters in order to achieve a human-like interaction. In this paper the final handover configuration is examined and analyzed. For this purpose, we conducted an experiment in which subjects taught a robot proper and improper end poses for different objects in different situations. We analyze the learned poses to determine what factors proper poses depend on. In addition to the poses, the configurations of the individual joints of the commercial robot are evaluated. Our results demonstrate that proper poses can be grouped into three clusters. These differ mainly in the rotation of the forearm. The height of a proper handover is also determined based on external factors.
Robin Rasch, Sven Wachsmuth, Matthias König 0001
IROS2
2019 Human Work Activity Recognition for Working Cells in Industrial Production Contexts
abstract
Collaboration between robots and humans requires communicative skills on both sides. The robot has to understand the conscious and unconscious activities of human workers. Many state-of-the-art activity recognition algorithms with high performance rates on existing benchmark datasets are available for this task. This paper re-evaluates appropriate architectures in light of human work activity recognition for working cells in industrial production contexts. The specific constraints of such a domain is elaborated and used as prior knowledge. We utilize state-of-the-art algorithms as spatiotemporal feature encoders and search for appropriate classification and fusion strategies. Furthermore, we combine keypoint-based with appearance-based approaches to a multi-stream recognition system. Due to data protection rules and the high effort of data annotation within industrial domains only small datasets are available that reflect production aspects. Therefore, we use transfer learning approaches to reduce the dependency on data volume and variance in the target domain. The resulting recognition system achieves high performance for both singular person action and human-object interaction.
Clemens Pohlt, Thomas Schlegl, Sven Wachsmuth
SMC3
2018 ToBI - Team of Bielefeld Enhancing the Robot Capabilities of the Social Standard Platform Pepper
Florian Lier, Johannes Kummert, Patrick Renner, Sven Wachsmuth
RoboCup4
2018 Effects on User Experience During Human-Robot Collaboration in Industrial Scenarios
abstract
In smart manufacturing environments robots collaborate with human operators as peers. They even share the same working space and time. An intuitive interaction with different input modalities is decisive to reduce workload and training periods for collaboration. We introduce our interaction system that is able to recognize gestures, actions and objects in a typical smart working scenario. As key aspect, this article considers an empirical investigation of input modalities (touch, gesture), individual differences (performance, recognition rate, previous knowledge) and boundary conditions (level of automation) on user experience. Therefore, answers from 31 participants within two experiments are collected. We show that the arrangement of the human-robot collaboration (input modalities, boundary conditions) has a significant effect on user experience in real-world environments. This effect and the individual differences between participants can be measured utilizing recognition rates and standardized usability questionnaires.
Clemens Pohlt, Franz Haubner, Jonas Lang, Sandra Rochholz, Thomas Schlegl, Sven Wachsmuth
SMC6
2017 Impact of Spontaneous Human Inputs during Gesture based Interaction on a Real-World Manufacturing Scenario
abstract
Seamless human-robot collaboration depends on high non-verbal behaviour recognition rates. To realize that in real-world manufacturing scenarios with an ecological valid setup, a lot of effort has to be invested. In this paper, we evaluate the impact of spontaneous inputs on the robustness of human-robot collaboration during gesture-based interaction. A high share of these spontaneous inputs lead to a reduced capability to predict behaviour and subsequently to a loss of robustness. We observe body and hand behaviour during interactive manufacturing of a collaborative task within two experiments. First, we analyse the occurrence frequency, reason and manner of human inputs in specific situations during a human-human experiment. We show the high impact of spontaneous inputs, especially in situations that differ from the typical working procedure. Second, we concentrate on implicit inputs during a real-world Wizard of Oz experiment using our human-robot working cell. We show that hand positions can be used to anticipate user needs in a semi-structured environment by applying knowledge about the semi-structured human behaviour which is distributed over working space and time in a typical manner.
Clemens Pohlt, Sebastian Hell, Thomas Schlegl, Sven Wachsmuth
HAI4
2016 Are you talking to me?: Improving the Robustness of Dialogue Systems in a Multi Party HRI Scenario by Incorporating Gaze Direction and Lip Movement of Attendees
abstract
In this paper, we present our humanoid robot "Meka", participating in a multi party human robot dialogue scenario. Active arbitration of the robot's attention based on multi-modal stimuli is utilised to observe persons which are outside of the robots field of view. We investigate the impact of this attention management and addressee recognition on the robot's capability to distinguish utterances directed at it from communication between humans. Based on the results of a user study, we show that mutual gaze at the end of an utterance, as a means of yielding a turn, is a substantial cue for addressee recognition. Verification of a speaker through the detection of lip movements can be used to further increase precision. Furthermore, we show that even a rather simplistic fusion of gaze and lip movement cues allows a considerable enhancement in addressee estimation, and can be altered to adapt to the requirements of a particular scenario.
Viktor Richter, Birte Richter, Florian Lier, Sebastian Meyer zu Borgsen, David Schlangen, Franz Kummert, Sven Wachsmuth, Britta Wrede
HAI7
2016 Humotion: A Human Inspired Gaze Control Framework for Anthropomorphic Robot Heads
abstract
In recent years, an attempt is being made to control robots more intuitive and intelligible by exploiting and integrating anthropomorphic features to boost social human-robot interaction. The design and construction of anthropomorphic robots for this kind of interaction is not the only challenging issue -- smooth and expectation-matching motion control is still an unsolved topic. In this work we present a highly configurable, portable, and open control framework that facilitates anthropomorphic motion generation for humanoid robot heads by enhancing state-of-the-art neck-eye coordination with human-like eyelid saccades and animation. On top of that, the presented framework supports dynamic neck offset angles that allow animation overlays and changes in alignment to the robots communication partner whileretaining visual focus on a given target. In order to demonstrate the universal applicability of the proposed ideas we used this framework to control the Flobi and the iCub robot head, both in simulation and on the physical robot. In order to foster further comparative studies of different robot heads, we will release all software, based on this contribution, under an open-source license.
Simon Schulz, Florian Lier, Andreas Kipp 0002, Sven Wachsmuth
HAI4
2016 1st international workshop on embodied interaction with smart environments (workshop summary)
abstract
The first workshop on embodied interaction with smart environments aims to bring together the very active community of multi-modal interaction research and the rapidly evolving field of smart home technologies. Besides addressing the software architecture of such very complex systems, it puts an emphasis on questions regarding an intuitive interaction with the environment. Thereby, especially the role of agency leads to interesting challenges in the light of user interactions. We therefore encourage a lively discussion on the design and concepts of social robots and virtual avatars as well as innovative ambient devices and their implementation into smart environments.
Patrick Holthaus, Thomas Hermann 0001, Sebastian Wrede 0001, Sven Wachsmuth, Britta Wrede
ICMI4
2016 Towards automated system and experiment reproduction in robotics
abstract
Even though research on autonomous robots and human-robot interaction accomplished great progress in recent years, and reusable soft- and hardware components are available, many of the reported findings are only hardly reproducible by fellow scientists. Usually, reproducibility is impeded because required information, such as the specification of software versions and their configuration, required data sets, and experiment protocols are not mentioned or referenced in most publications. In order to address these issues, we recently introduced an integrated tool chain and its underlying development process to facilitate reproducibility in robotics. In this contribution we instantiate the complete tool chain in a unique user study in order to assess its applicability and usability. To this end, we chose three different robotic systems from independent institutions and modeled them in our tool chain, including three exemplary experiments. Subsequently, we asked twelve researchers to reproduce one of the formerly unknown systems and the associated experiment. We show that all twelve scientists were able to replicate a formerly unknown robotics experiment using our tool chain.
Florian Lier, Marc Hanheide, Lorenzo Natale, Simon Schulz, Jonathan Weisz, Sven Wachsmuth, Sebastian Wrede 0001
IROS6
2016 How to Address Smart Homes with a Social Robot? A Multi-modal Corpus of User Interactions with an Intelligent Environment
Patrick Holthaus, Christian Leichsenring, Jasmin Bernotat, Viktor Richter, Marian Pohling, Birte Richter, Norman Köster, Sebastian Meyer zu Borgsen, René Zorn, Birte Schiffhauer, Kai Frederic Engelmann, Florian Lier, Simon Schulz, Philipp Cimiano, Friederike Eyssel, Thomas Hermann 0001, Franz Kummert, David Schlangen, Sven Wachsmuth, Petra Wagner, Britta Wrede, Sebastian Wrede 0001
LREC19
2016 ToBI - Team of Bielefeld: Enhancing Robot Behaviors and the Role of Multi-robotics in RoboCup@Home
Sebastian Meyer zu Borgsen, Timo Korthals, Florian Lier, Sven Wachsmuth
RoboCup4
2015 RoboCup@Home - Benchmarking Domestic Service Robots
abstract
The RoboCup@Home league has been founded in 2006with the idea to drive research in AI and related fieldstowards autonomous and interactive robots that copewith real life tasks in supporting humans in everday life.The yearly competition format establishes benchmarkingas a continuous process with yearly changes insteadof a single challenge. We discuss the current state andfuture perspectives of this endeavor.
Sven Wachsmuth, Dirk Holz, Maja Rudinac, Javier Ruiz-del-Solar
AAAI1
2014 The receptionist robot
abstract
In this demonstration, a humanoid robot interacts with an interlocutor through speech and gestures in order to give directions on a map. The interaction is specifically designed to provide an enhanced user experience by being aware of non-verbal social signals. Therefore, we take spatial communicative cues into account and to react to them accordingly.
Patrick Holthaus, Sven Wachsmuth
HRI2
2014 Towards automated execution and evaluation of simulated prototype HRI experiments
abstract
Autonomous robots are highly relevant targets for interaction studies, but can exhibit behavioral variability that confounds experimental validity. Currently, testing on real systems is the only means to prevent this, but remains very labour-intensive and often happens too late. To improve this situation, we are working towards early testing by means of partial simulation, with automated assessment, and based upon continuous software integration to prevent regressions. We will introduce the concept and describe a proof-of-concept that demonstrates fast feedback and coherent experiment results across repeated trials.
Florian Lier, Ingo Lütkebohle, Sven Wachsmuth
HRI3
2014 Reality check!: a physical robot versus its simulation
abstract
Simulated environments usually provide the most frequent test environment for robotic systems, often due to their cost and availability advantages. The crucial question is: how precisely must a simulation match the real world in order to produce realistic results? In this demonstration the humanoid robot head Flobi, as well as its simulation, will react to diverse visual stimuli in order to indicate visual attention which is an important cue in social interaction. We directly compare the behavior of the virtual robot, triggered by simulated stimuli, to the physical robot, triggered by real world stimuli. Therefore, we are able to assess and demonstrate the current limitations and advantages of our simulation compared to real world interactions.
Florian Lier, Simon Schulz, Sven Wachsmuth
HRI3
2014 On RoboCup@Home - Past, Present and Future of a Scientific Competition for Service Robots
Dirk Holz, Javier Ruiz-del-Solar, Komei Sugiura, Sven Wachsmuth
RoboCup4
2013 A Computational Model for Reference Object Selection in Spatial Relations
Katrin Johannsen, Agnes Swadzba, Leon Ziegler, Sven Wachsmuth, Jan Peter de Ruiter
COSIT4
2013 Facial communicative signal interpretation in human-robot interaction by discriminative video subsequence selection
abstract
Facial communicative signals (FCSs) such as head gestures, eye gaze, and facial expressions can provide useful feedback in conversations between people and also in human-robot interaction. This paper presents a pattern recognition approach for the interpretation of FCSs in terms of valence, based on the selection of discriminative subsequences in video data. These subsequences capture important temporal dynamics and are used as prototypical reference subsequences in a classification procedure based on dynamic time warping and feature extraction with active appearance models. Using this valence classification, the robot can discriminate positive from negative interaction situations and react accordingly. The approach is evaluated on a database containing videos of people interacting with a robot by teaching the names of several objects to it. The verbal answer of the robot is expected to elicit the display of spontaneous FCSs by the human tutor, which were classified in this work. The achieved classification accuracies are comparable to the average human recognition performance and outperformed our previous results on this task.
Christian Lang 0002, Sven Wachsmuth, Marc Hanheide, Heiko Wersing
ICRA2
2013 Robot reality - A motion capture system that makes robots become human and vice versa
abstract
A key issue for interactive robots is smooth and natural motion. When robots are tele-operated, e.g. for tele-presence or during “Wizard of Oz” style studies, this requires an intuitive and responsive control interface. Furthermore, this interface should not interfere with perceiving the robot's sensors.
Simon Schulz, Florian Lier, Ingo Lütkebohle, Sven Wachsmuth
ICRA4
2013 Integrating Multiple Viewpoints for Articulated Scene Model Aquisition
Leon Ziegler, Agnes Swadzba, Sven Wachsmuth
ICVS3
2013 Direct on-line imitation of human faces with hierarchical ART networks
abstract
This work-in-progress paper presents an on-line system for robotic heads capable of mimicking humans. The marker-less method solely depends on the interactant's face as an input and does not use a set of basic emotions and is thus capable of displaying a large variety of facial expressions. A preliminary evaluation assigns solid performance with potential for improvement.
Patrick Holthaus, Sven Wachsmuth
RO-MAN2
2013 ART-based fusion of multi-modal perception for robots
Elmar Berghöfer, Denis Schulze, Christian Rauch 0004, Marko Tscherepanow, Tim Köhler, Sven Wachsmuth
Neurocomputing6
2012 Integrating PAMOCAT in the research cycle: linking motion capturing and conversation analysis
abstract
In order to understand and model the non-verbal communicative conduct of humans, it seems fruitful to combine qualitative (Conversation Analysis [6] [10] [11]) and quantitative analytical (motion capturing) methods. Tools for data visualization and annotation are important as they constitute a central interface between different research approaches and methodologies. With this aim we have developed the pre-annotation tool "PAMOCAT - Pre Annotation Motion Capture Analysis Tool" that detects different phenomena. These phenomena are in the category single person and person overlapping phenomena. Included are functions for the analysis of head focused objects, hand activities, single DOF- degree of freedom activity, posture detection and intrusions into the co-participant's space. These detected phenomena related to the frames will be displayed in an overview. The phenomena can be chosen to search for a specific constellation between these different phenomena. A sophisticated user interface easily allows the annotating person to find correlations between different joints and phenomena, to analyze the corresponding 3D pose in a reconstructed virtual environment, and to export combined qualitative and quantitative annotations to standard annotation tools. Using this technique we are able to examine complex setups with three participants engaged in conversation. In this paper we propose how PAMOCAT can be integrated in the research cycle by showing a concrete PAMOCAT-based micro-analysis of a multimodal phenomenon, which deals with kinetic procedures to claim the floor.
Bernhard Andreas Brüning, Christian Schnier, Karola Pitsch, Sven Wachsmuth
ICMI4
2012 User Behavior Recognition for an Automatic Prompting System - A Structured Approach based on Task Analysis
Christian Peters, Thomas Hermann 0001, Sven Wachsmuth
ICPRAM (2)3
2012 An affordable, 3D-printable camera eye with two active degrees of freedom for an anthropomorphic robot
abstract
Movable cameras, or robotic eyes, have been of longstanding interest in robotics, but it remains challenging to achieve the various desirable characteristics (such as fast, small, robust, good-looking, etc.) at the same time. We consider robotic eyes targeted at interactive humanoid robots, and present a novel construction that provides human-level accelerations and velocities in a robust way, while its appearance avoids the uncanny valley effect by hiding all mechanics.
Simon Schulz, Ingo Lütkebohle, Sven Wachsmuth
IROS3
2012 PAMOCAT: Automatic retrieval of specified postures
Bernhard Andreas Brüning, Christian Schnier, Karola Pitsch, Sven Wachsmuth
LREC4
2011 Automatic Enhancement of Correspondence Detection in an Object Tracking System
Denis Schulze, Sven Wachsmuth, Katharina J. Rohlfing
ESANN2
2011 View-Tuned Approximate Partial Matching Kernel from Hierarchical Growing Neural Gases
Marco Kortkamp, Sven Wachsmuth
ICANN (2)2
2010 Indoor Scene Classification Using Combined 3D and Gist Features
Agnes Swadzba, Sven Wachsmuth
ACCV (2)2
2010 Continuous Visual Codebooks with a Limited Branching Tree Growing Neural Gas
Marco Kortkamp, Sven Wachsmuth
ICANN (3)2
2010 The Bielefeld anthropomorphic robot head "Flobi"
abstract
A robot's head is important both for directional sensors and, in human-directed robotics, as the single most visible interaction interface. However, designing a robot's head faces contradicting requirements when integrating powerful sensing with social expression. Furher, reactions of the general public show that current head designs often cause negative user reactions and distract from the functional capabilities. Therefore, this contribution presents a novel anthropomorphic robot head called "Flobi", which combines state-of-the-art sensing functionality with an exterior that elicits a sympathetic emotional response. It can display primary and secondary emotions in a human-like way, to enable intuitive human-robot-interaction. To facilitate further research on facial appearance, the exterior is fully modular and replaceable. While current state-of-the-art still requires trade-offs when integrating sensing and social expression, Flobi has been designed to enable service robotic applications, with high-resolution, wide-angle stereo vision, gyroscope motion compensation and stereo audio. For ease of integration, the head is self-contained, including 18 actuators, sensors and control boards, all in a human-head sized package.
Ingo Lütkebohle, Frank Hegel, Simon Schulz, Matthias Hackel, Britta Wrede, Sven Wachsmuth, Gerhard Sagerer
ICRA6
2010 Dynamic 3D scene analysis for acquiring articulated scene models
abstract
In this paper we present a new system for a mobile robot to generate an articulated scene model by analyzing complex dynamic 3D scenes. The system extracts essential knowledge about the foreground, like moving persons, and the background, which consists of all visible static scene parts. In contrast to other 3D reconstruction approaches, we suggest to additionally distinguish between static parts, like walls, and movable objects like chairs or doors. The discrimination supports the reconstruction process and additionally, delivers important information about interaction objects. Here, the movable object detection is realized object independent by analyzing changes in the scenery. Furthermore, in the proposed system the background scene is feedbacked to the tracking part yielding a much better tracking and detection result which improves again the 3D reconstruction. We show in our experiments that we are able to provide a sound background model and to extract simultaneously persons and object regions representing chairs, doors, and even smaller movable objects.
Agnes Swadzba, Niklas Beuter, Sven Wachsmuth, Franz Kummert
ICRA3
2010 Using Language to Learn Structured Appearance Models for Image Annotation
abstract
Given an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to simultaneously learn the names and appearances of the objects. Only a small fraction of local features within any given image are associated with a particular caption word, and captions may contain irrelevant words not associated with any image object. We propose a novel algorithm that uses the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to learn meaningful feature configurations (representing named objects). We also introduce a graph-based appearance model that captures some of the structure of an object by encoding the spatial relationships among the local visual features. In an iterative procedure, we use language (the words) to drive a perceptual grouping process that assembles an appearance model for a named object. Results of applying our method to three data sets in a variety of conditions demonstrate that, from complex, cluttered, real-world scenes with noisy captions, we can learn both the names and appearances of objects, resulting in a set of models invariant to translation, scale, orientation, occlusion, and minor changes in viewpoint or articulation. These named models, in turn, are used to automatically annotate new, uncaptioned images, thereby facilitating keyword-based image retrieval.
Michael Jamieson, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson, Sven Wachsmuth
IEEE Trans. Pattern Anal. Mach. Intell.5
2009 The curious robot - Structuring interactive robot learning
abstract
If robots are to succeed in novel tasks, they must be able to learn from humans. To improve such human-robot interaction, a system is presented that provides dialog structure and engages the human in an exploratory teaching scenario. Thereby, we specifically target untrained users, who are supported by mixed-initiative interaction using verbal and non-verbal modalities. We present the principles of dialog structuring based on an object learning and manipulation scenario. System development is following an interactive evaluation approach and we will present both an extensible, event-based interaction architecture to realize mixed-initiative and evaluation results based on a video-study of the system. We show that users benefit from the provided dialog structure to result in predictable and successful human-robot interaction.
Ingo Lütkebohle, Julia Peltason, Lars Schillingmann, Britta Wrede, Sven Wachsmuth, Christof Elbrechter, Robert Haschke
ICRA5
2009 Laser-based navigation enhanced with 3D time-of-flight data
abstract
Navigation and obstacle avoidance in robotics using planar laser scans has matured over the last decades. They basically enable robots to penetrate highly dynamic and populated spaces, such as people's home, and move around smoothly. However, in an unconstrained environment the two-dimensional perceptual space of a fixed mounted laser is not sufficient to ensure safe navigation. In this paper, we present an approach that pools a fast and reliable motion generation approach with modern 3D capturing techniques using a time-of-flight camera. Instead of attempting to implement full 3D motion control, which is computationally more expensive and simply not needed for the targeted scenario of a domestic robot, we introduce a ldquovirtual laserrdquo. For the originally solely laser-based motion generation the technique of fusing real laser measurements and 3D point clouds into a continuous data stream is 100% compatible and transparent. The paper covers the general concept, the necessary extrinsic calibration of two very different types of sensors, and exemplarily illustrates the benefit which is to avoid obstacles not being perceivable in the original laser scan.
Agnes Swadzba, Roland Philippsen, Orhan Engin, Marc Hanheide, Sven Wachsmuth
ICRA6
2009 A Computational Model for the Alignment of Hierarchical Scene Representations in Human-Robot Interaction
Agnes Swadzba, Sven Wachsmuth, Constanze Vorwerg, Gert Rickheit
IJCAI2
2008 Reducing noise and redundancy in registered range data for planar surface extraction
abstract
This paper presents a new method for detecting and merging redundant points in registered range data. Given a global representation from sequences of 3D points, the points are projected onto a virtual image plane computed from the intrinsic parameters of the sensor. Candidates for redundancy are collected per pixel which then are clustered locally via region growing and replaced by the clusterpsilas mean value. As data is provided in a certain manner defined by camera characteristics, this processing step preserves the structural information of the data. For evaluation, our approach is compared to two other algorithms. Applied to two different sequences, it is shown that the presented method gives smooth results within planar regions of the point clouds by successfully reducing noise and redundancy and thus improves registered range data.
Agnes Swadzba, Anna-Lisa Vollmer, Marc Hanheide, Sven Wachsmuth
ICPR4
2008 The visual active memory perspective on integrated recognition systems
Christian Bauckhage, Sven Wachsmuth, Marc Hanheide, Sebastian Wrede 0001, Gerhard Sagerer, Gunther Heidemann, Helge J. Ritter
Image Vis. Comput.2
2007 Learning Structured Appearance Models from Captioned Images of Cluttered Scenes
abstract
Given an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to learn both the names and appearances of the objects. Only a small number of local features within any given image are associated with a particular caption word. We describe a connected graph appearance model where vertices represent local features and edges encode spatial relationships. We use the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to guide the search for meaningful feature configurations. We demonstrate improved results on a dataset to which an unstructured object model was previously applied. We also apply the new method to a more challenging collection of captioned images from the Web, detecting and annotating objects within highly cluttered realistic scenes.
Michael Jamieson, Afsaneh Fazly, Sven J. Dickinson, Suzanne Stevenson, Sven Wachsmuth
ICCV5
2007 View-adaptive manipulative action recognition for robot companions
abstract
This paper puts forward an approach for a mobile robot to recognize the human's manipulative actions from different single camera views. While most of the related work in action recognition assume a fixed static camera view that is the same for training and testing, such kind of constraints do not apply for mobile robot companions. We propose a recognition scheme that is able to generalize an action model, that has been learned from a very few data items observed from a single camera view, to variant view points and different settings. We tackle the problem of compensating the view dependence of 2D motion models on three different levels. Firstly, we pre-segment the trajectories based on an object vicinity that depends on the camera tilt and object detections. Secondly, an interactive feature vector is designed that represents the relative movements between the human hand and the objects. Thirdly, we propose an adaptive HMM-based matching process that is based on a particle filter and includes a dynamically adjusted scaling parameter that models the systematic error of the view dependency. Finally, we use a two-layered approach for task recognition which decouples the task knowledge from the view dependent primitive recognition. The results of experiments in an office environment show the applicability of this approach.
Zhe Li 0009, Sven Wachsmuth, Jannik Fritsch, Gerhard Sagerer
IROS2
2007 Classes of Applications for Social Robots: A User Study
abstract
The paper introduces an online user study on applications for social robots with 127 participants. The potential users proposed 570 application scenarios based on the appearance and functionality of four robots presented (AIBO, BARTHOC, BIRON, iCat). The items were grouped into 13 categories which are interpreted and discussed by means of four dimensions: public vs. private use, intensity of interaction, complexity of interaction model, and functional vs. human-like appearance. The interpretation lead to three classes of applications for social robots according to the degree of social interaction: (1) Specialized Applications where the robot has to perform clearly defined tasks which are delegated by a user, (2) Public Applications which are directed to the communication with many users, and (3) Individual Applications with the need of a highly elaborated social model to maintain a variety of situations with few people.
Frank Hegel, Manja Lohse, Agnes Swadzba, Sven Wachsmuth, Katharina J. Rohlfing, Britta Wrede
RO-MAN4
2007 Coordinating interactive vision behaviors for cognitive assistance
Sven Wachsmuth, Sebastian Wrede 0001, Marc Hanheide
Comput. Vis. Image Underst.1
2006 Using Language to Drive the Perceptual Grouping of Local Image Features
abstract
We address the problem of learning both the semantics (names) and the visual features (SIFT collections) of objects appearing in a training set of unstructured, captioned images of cluttered scenes. Prior work in applying machine translation models to learn the associations between image features and caption nouns has assumed a one-toone correspondence between features and nouns. However, each training image may contain thousands of SIFT features belonging to multiple objects. Our challenge is two-fold: 1) grouping the SIFT features into meaningful collections, and 2) learning the object names associated with those collections. Since better collections tend to have stronger associations with object names, we offer an integrated solution that uses the caption words to drive the feature grouping process. The result is a more general model acquisition framework that does not assume words correspond to individual features and does not require training images with isolated objects or unambiguous labels. The model that is learned performs well at labeling cluttered scenes in a set of test images.
Michael Jamieson, Sven J. Dickinson, Suzanne Stevenson, Sven Wachsmuth
CVPR (2)4
2006 Integration and Coordination in a Cognitive Vision System
abstract
In this paper, we present a case study that exemplifies general ideas of system integration and coordination. The application field of assistant technology provides an ideal test bed for complex computer vision systems including real-time components, human-computer interaction, dynamic 3-d environments, and information retrieval aspects. In our scenario the user is wearing an augmented reality device that supports her/him in everyday tasks by presenting information that is triggered by perceptual and contextual cues. The system integrates a wide variety of visual functions like localization, object tracking and recognition, action recognition, interactive object learning, etc. We show how different kinds of system behavior are realized using the Active Memory Infrastructure that provides the technical basis for distributed computation and a data- and eventdriven integration approach.
Sebastian Wrede 0001, Marc Hanheide, Sven Wachsmuth, Gerhard Sagerer
ICVS3
2005 Modelling expertise for structure elucidation in organic chemistry using Bayesian networks
Michaela Hohenner, Sven Wachsmuth, Gerhard Sagerer
Knowl. Based Syst.2
2002 Evaluating Integrated Speech- and Image Understanding
abstract
The capability to coordinate and interrelate speech and vision is a virtual prerequisite for adaptive, cooperative, and flexible interaction among people. It is therefore fair to assume that human-machine interaction, too, would benefit from intelligent interfaces for integrated speech and image processing. We first sketch an interactive system that integrates automatic speech processing with image understanding. Then, we concentrate on performance assessment which we believe is an emerging key issue in multimodal interaction. We explain the benefit of time scale analysis and usability studies and evaluate our system accordingly.
Christian Bauckhage, Jannik Fritsch, Katharina J. Rohlfing, Sven Wachsmuth, Gerhard Sagerer
ICMI4
2002 Multi-modal human-machine communication for instructing robot grasping tasks
abstract
A major challenge for the realization of intelligent robots is to supply them with cognitive abilities in order to allow ordinary users to program them easily and intuitively. One approach to such programming is teaching work tasks by interactive demonstration. To make this effective and convenient for the user, the machine must be capable of establishing a common focus of attention and be able to use and integrate spoken instructions, visual perception, and non-verbal clues like gestural commands. We report progress in building a hybrid architecture that combines statistical methods, neural networks, and finite state machines into an integrated system for instructing grasping tasks by man-machine interaction. The system combines the GRAVIS-robot for visual attention and gestural instruction with an intelligent interface for speech recognition and linguistic interpretation, and a modality fusion module to allow multi-modal task-oriented man-machine communication with respect to dextrous robot manipulation of objects.
Patrick C. McGuire, Jannik Fritsch, Jochen J. Steil, Frank Röthling, Gernot A. Fink, Sven Wachsmuth, Gerhard Sagerer, Helge J. Ritter
IROS6
2000 Towards an Integrated Framework for Contour-Based Grouping and Object Recognition Using Markov Random Fields
abstract
We present an integrated approach for contour-based grouping and object recognition. Domain knowledge and domain independent grouping laws are combined in a multi-layered Markov random field (MRF) framework. It provides a basis for propagating top-down knowledge between different processing cues or input modalities. Additionally, the domain dependent MRF-layer can be used in order to evaluate the grouping process with regard to relevant contours for object recognition. Initial results show the approach to be adequate for complex scenes and partially occluded objects.
Daniel Schlüter, Sven Wachsmuth, Gerhard Sagerer, Stefan Posch
ICIP2
1999 Multilevel Integration of Vision and Speech Understanding Using Bayesian Networks
Sven Wachsmuth, Hans Brandt-Pook, Gudrun Socher, Franz Kummert, Gerhard Sagerer
ICVS1