VLDB 2026 Research / reviewers in the wild / expert
Sven Behnke
dblp:16/6112
· DBLP profile ↗
210ranked-venue papers
17as first author
51since 2021 · last 2026
0000-0002-5040-7525ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 195 · 16 first-author · 44 since 2021Systems, architecture and hardware · 73 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from TranslationabstractCross-lingual, cross-task transfer is challenged by task-specific data scarcity, which becomes more severe as language support grows and is further amplified in vision-language models (VLMs). We investigate multilingual generalization in encoder-decoder transformer VLMs to enable zero-shot image captioning in languages encountered only in the translation task. In this setting, the encoder must learn to generate generalizable, task-aware latent vision representations to instruct the decoder via inserted cross-attention layers. To analyze scaling behavior, we train Florence-2 based and Gemma-2 based models (0.4B to 11.2B parameters) on a synthetic dataset using varying compute budgets. While all languages in the dataset have image-aligned translations, only a subset of them include image captions. Notably, we show that captioning can emerge using a language prefix, even when this language only appears in the translation task. We find that indirect learning of unseen task-language pairs adheres to scaling laws that are governed by the multilinguality of the model, model size, and seen training samples. Finally, we demonstrate that the scaling laws extend to downstream tasks, achieving competitive performance through fine-tuning in multimodal machine translation (Multi30K, CoMMuTE), lexical disambiguation (CoMMuTE), and image captioning (Multi30K, XM3600, COCO Karpathy). Julian Spravil, Sebastian Houben, Sven Behnke |
AAAI | 3 |
| 2026 | Dexterous Pre-Grasp Manipulation for Human-Like Functional Categorical Grasping: Deep Reinforcement Learning and Grasp RepresentationsabstractMany objects such as tools and household items can be used only if grasped in a very specific way—grasped functionally. Often, a direct functional grasp is not possible, though. We propose a method for learning a dexterous pre-grasp manipulation policy to achieve human-like functional grasps using deep reinforcement learning. We introduce a dense multi-component reward function that enables learning a single policy, capable of dexterous pre-grasp manipulation of novel instances of several known object categories with an anthropomorphic hand. The policy is learned purely by means of reinforcement learning from scratch, without any expert demonstrations, and implicitly learns to reposition and reorient objects of complex shapes to achieve given functional grasps. In addition, we explore two different ways to represent a desired grasp: explicit and more abstract, constraint-based. We show that our method learns to successfully manipulate and achieve desired grasps on previously unseen instances of known categories using both grasp representations. Learning is done on a single GPU in less than three hours. Dmytro Pavlichenko, Sven Behnke |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Feature-Preserving Mesh Decimation for Normal IntegrationabstractNormal integration reconstructs 3D surfaces from normal maps obtained e.g. by photometric stereo. These normal maps capture surface details down to the pixel level but require large computational resources for integration at high resolutions. In this work, we replace the dense pixel grid with a sparse anisotropic triangle mesh prior to normal integration. We adapt the triangle mesh to the local geometry in the case of complex surface structures and remove oversampling from flat featureless regions. For high-resolution images, the resulting compression reduces normal integration runtimes from hours to minutes while maintaining high surface accuracy. Our main contribution is the derivation of the well-known quadric error measure from mesh decimation for screen space applications and its combination with optimal Delaunay triangulation. Code is available at https://moritzheep.github.io/anisotropic-screen-meshing. Moritz Heep, Sven Behnke, Eduard Zell |
CVPR | 2 |
| 2025 | VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks
Noel José Rodrigues Vicente, Enrique Lehner, Angel Villar-Corrales, Jan Nogga, Sven Behnke |
ICANN (2) | 5 |
| 2025 | SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from PixelsabstractLearning a latent dynamics model provides a task-agnostic representation of an agent's understanding of its environment. Leveraging this knowledge for model-based reinforcement learning (RL) holds the potential to improve sample efficiency over model-free methods by learning from imagined rollouts. Furthermore, because the latent space serves as input to behavior models, the informative representations learned by the world model facilitate efficient learning of desired skills. Most existing methods rely on holistic representations of the environment’s state. In contrast, humans reason about objects and their interactions, predicting how actions will affect specific parts of their surroundings. Inspired by this, we propose *Slot-Attention for Object-centric Latent Dynamics (SOLD)*, a novel model-based RL algorithm that learns object-centric dynamics models in an unsupervised manner from pixel inputs. We demonstrate that the structured latent space not only improves model interpretability but also provides a valuable input space for behavior models to reason over. Our results show that SOLD outperforms DreamerV3 and TD-MPC2 - state-of-the-art model-based RL algorithms - across a range of multi-object manipulation environments that require both relational reasoning and dexterous control. Videos and code are available at https:// slot-latent-dynamics.github.io. Malte Mosbach, Jan Niklas Ewertz, Angel Villar-Corrales, Sven Behnke |
ICML | 4 |
| 2025 | PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and PlanningabstractPredicting future scene representations is a crucial task for enabling robots to understand and interact with the environment. However, most existing methods rely on videos and simulations with precise action annotations, limiting their ability to leverage the large amount of available unlabeled video data. To address this challenge, we propose PlaySlot, an object-centric video prediction model that infers object representations and latent actions from unlabeled video sequences. It then uses these representations to forecast future object states and video frames. PlaySlot allows the generation of multiple possible futures conditioned on latent actions, which can be inferred from video dynamics, provided by a user, or generated by a learned action policy, thus enabling versatile and interpretable world modeling. Our results show that PlaySlot outperforms both stochastic and object-centric baselines for video prediction across different environments. Furthermore, we show that our inferred latent actions can be used to learn robot behaviors sample-efficiently from unlabeled video demonstrations. Videos and code are available on our project website. Angel Villar-Corrales, Sven Behnke |
ICML | 2 |
| 2025 | Prompt-Responsive Object Retrieval with Memory-Augmented Student-Teacher LearningabstractBuilding models responsive to input prompts represents a transformative shift in machine learning. This paradigm holds significant potential for robotics problems, such as targeted manipulation amidst clutter. In this work, we present a novel approach to combine promptable foundation models with reinforcement learning (RL), enabling robots to perform dexterous manipulation tasks in a prompt-responsive manner. Existing methods struggle to link high-level commands with fine-grained dexterous control. We address this gap with a memory-augmented student-teacher learning framework. We use the Segment-Anything 2 (SAM2) model as a perception backbone to infer an object of interest from user prompts. While detections are imperfect, their temporal sequence provides rich information for implicit state estimation by memory-augmented models. Our approach successfully learns prompt-responsive policies, demonstrated in picking objects from cluttered scenes. Videos and code are available at https://memory-student-teacher.github.io Malte Mosbach, Sven Behnke |
ICRA | 2 |
| 2025 | DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic ModelsabstractPerception systems play a crucial role in autonomous driving, incorporating multiple sensors and corresponding computer vision algorithms. 3D LiDAR sensors are widely used to capture sparse point clouds of the vehicle’s surroundings. However, such systems struggle to perceive occluded areas and gaps in the scene due to the sparsity of these point clouds and their lack of semantics. To address these challenges, Semantic Scene Completion (SSC) jointly predicts unobserved geometry and semantics in the scene given raw LiDAR measurements, aiming for a more complete scene representation. Building on promising results of diffusion models in image generation and super-resolution tasks, we propose their extension to SSC by implementing the noising and denoising diffusion processes in the point and semantic spaces individually. To control the generation, we employ semantic LiDAR point clouds as conditional input and design local and global regularization losses to stabilize the denoising process. We evaluate our approach on autonomous driving datasets, and it achieves state-of-the-art performance for SSC, surpassing most existing methods. Helin Cao, Sven Behnke |
IROS | 2 |
| 2025 | OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric AwarenessabstractAutonomous driving perception faces significant challenges due to occlusions and incomplete scene data in the environment. To overcome these issues, the task of semantic occupancy prediction (SOP) is proposed, which aims to jointly infer both the geometry and semantic labels of a scene from images. However, conventional camera-based methods typically treat all categories equally and primarily rely on local features, leading to suboptimal predictions, especially for dynamic foreground objects. To address this, we propose Object-Centric SOP (OC-SOP), a framework that integrates high-level object-centric cues extracted via a detection branch into the semantic occupancy prediction pipeline. This object-centric integration significantly enhances the prediction accuracy for foreground objects and achieves state-of-the-art performance among all categories on SemanticKITTI. Helin Cao, Sven Behnke |
SMC | 2 |
| 2025 | SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous DrivingabstractPerception systems in autonomous driving rely on sensors such as LiDAR and cameras to perceive the 3D environment. However, due to occlusions and data sparsity, these sensors often fail to capture complete information. Semantic Occupancy Prediction (SOP) addresses this challenge by inferring both occupancy and semantics of unobserved regions. Existing transformer-based SOP methods lack explicit modeling of spatial structure in attention computation, resulting in limited geometric awareness and poor performance in sparse or occluded areas. To this end, we propose Spatially-aware Window Attention (SWA), a novel mechanism that incorporates local spatial context into attention. SWA significantly improves scene completion and achieves state-of-the-art results on LiDAR-based SOP benchmarks. We further validate its generality by integrating SWA into a camera-based SOP pipeline, where it also yields consistent gains across modalities. Helin Cao, Rafael Materla, Sven Behnke |
SMC | 3 |
| 2025 | LiLMaps: Learnable Implicit Language MapsabstractOne of the current trends in robotics is to employ large language models (LLMs) to provide non-predefined commands execution and natural human-robot interaction. It is useful to have an environment map together with its language representation, which can be further utilized by LLMs. Such a comprehensive scene representation enables numerous ways of interaction with the map for autonomously operating robots. In this work, we present an approach that enhances incremental implicit mapping through the integration of visual-language features. Specifically, we (i) propose a decoder optimization technique for implicit language maps which can be used when new objects appear on the scene, and (ii) address the problem of inconsistent visual-language predictions between different viewing positions. Our experiments demonstrate the effectiveness of LiLMaps and solid improvements in performance. Evgeny Kruzhkov, Sven Behnke |
WACV | 2 |
| 2024 | MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
Angel Villar-Corrales, Moritz Austermann, Sven Behnke |
BMVC | 3 |
| 2024 | FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-Pose, and Facial Expression FeaturesabstractThe task of face reenactment is to transfer the head motion and facial expressions from a driving video to the appearance of a source image, which may be of a different person (cross-reenactment). Most existing methods are CNN-based and estimate optical flow from the source image to the current driving frame, which is then inpainted and refined to produce the output animation. We propose a transformer-based encoder for computing a set-latent representation of the source image(s). We then predict the output color of a query pixel using a transformer-based decoder, which is conditioned with keypoints and a facial expression vector extracted from the driving frame. Latent representations of the source person are learned in a self-supervised manner that factorize their appearance, head pose, and facial expressions. Thus, they are perfectly suited for cross-reenactment. In contrast to most related work, our method naturally extends to multiple source images and can thus adapt to person-specific facial dynamics. We also propose data augmentation and regularization schemes that are necessary to prevent overfitting and support generalizability of the learned representations. We evaluated our approach in a randomized user study. The results indicate superior performance compared to the state-of-the-art in terms of motion transfer quality and temporal consistency.11Code & Video: https://andrerochow.github.io/fsrt Andre Rochow, Max Schwarz, Sven Behnke |
CVPR | 3 |
| 2024 | HyenaPixel: Global Image Context with ConvolutionsabstractIn computer vision, a larger effective receptive field (ERF) is associated with better performance. While attention natively supports global context, its quadratic complexity limits its applicability to tasks that benefit from high-resolution input. In this work, we extend Hyena, a convolution-based attention replacement, from causal sequences to bidirectional data and two-dimensional image space. We scale Hyena’s convolution kernels beyond the feature map size, up to 191×191, to maximize ERF while maintaining sub-quadratic complexity in the number of pixels. We integrate our two-dimensional Hyena, HyenaPixel, and bidirectional Hyena into the MetaFormer framework. For image categorization, HyenaPixel and bidirectional Hyena achieve a competitive ImageNet-1k top-1 accuracy of 84.9% and 85.2%, respectively, with no additional training data, while outperforming other convolutional and large-kernel networks. Combining HyenaPixel with attention further improves accuracy. We attribute the success of bidirectional Hyena to learning the data-dependent geometric arrangement of pixels without a fixed neighborhood definition. Experimental results on downstream tasks suggest that HyenaPixel with large filters and a fixed neighborhood leads to better localization performance. Julian Spravil, Sebastian Houben, Sven Behnke |
ECAI | 3 |
| 2024 | SLCF-Net: Sequential LiDAR-Camera Fusion for Semantic Scene Completion using a 3D Recurrent U-NetabstractWe introduce SLCF-Net, a novel approach for the Semantic Scene Completion (SSC) task that sequentially fuses LiDAR and camera data. It jointly estimates missing geometry and semantics in a scene from sequences of RGB images and sparse LiDAR measurements. The images are semantically segmented by a pre-trained 2D U-Net and a dense depth prior is estimated from a depth-conditioned pipeline fueled by Depth Anything. To associate the 2D image features with the 3D scene volume, we introduce Gaussian-decay Depth-prior Projection (GDP). This module projects the 2D features into the 3D volume along the line of sight with a Gaussian-decay function, centered around the depth prior. Volumetric semantics is computed by a 3D U-Net. We propagate the hidden 3D U-Net state using the sensor motion and design a novel loss to ensure temporal consistency. We evaluate our approach on the SemanticKITTI dataset and compare it with leading SSC approaches. The SLCF-Net excels in all SSC metrics and shows great temporal consistency. Helin Cao, Sven Behnke |
ICRA | 2 |
| 2024 | Grasp Anything: Combining Teacher-Augmented Policy Gradient Learning with Instance Segmentation to Grasp Arbitrary ObjectsabstractInteractive grasping from clutter, akin to human dexterity, is one of the longest-standing problems in robot learning. Challenges stem from the intricacies of visual perception, the demand for precise motor skills, and the complex interplay between the two. In this work, we present Teacher-Augmented Policy Gradient (TAPG), a novel two-stage learning framework that synergizes reinforcement learning and policy distillation. After training a teacher policy to master the motor control based on object pose information, TAPG facilitates guided, yet adaptive, learning of a sensorimotor policy, based on object segmentation. We zero-shot transfer from simulation to a real robot by using Segment Anything Model for promptable object segmentation. Our trained policies adeptly grasp a wide variety of objects from cluttered scenarios in simulation and the real world based on human-understandable prompts. Furthermore, we show robust zero-shot transfer to novel objects. Videos of our experiments are available at https://maltemosbach.github.io/grasp_anything. Malte Mosbach, Sven Behnke |
ICRA | 2 |
| 2024 | MOTPose: Multi-object 6D Pose Estimation for Dynamic Video Sequences using Attention-based Temporal FusionabstractCluttered bin-picking environments are challenging for pose estimation models. Despite the impressive progress enabled by deep learning, single-view RGB pose estimation models perform poorly in cluttered dynamic environments. Imbuing the rich temporal information contained in the video of scenes has the potential to enhance models’ ability to deal with the adverse effects of occlusion and the dynamic nature of the environments. Moreover, joint object detection and pose estimation models are better suited to leverage the co-dependent nature of the tasks for improving the accuracy of both tasks. To this end, we propose attention-based temporal fusion for multi-object 6D pose estimation that accumulates information across multiple frames of a video sequence. Our MOTPose method takes a sequence of images as input and performs joint object detection and pose estimation for all objects in one forward pass. It learns to aggregate both object embeddings and object parameters over multiple time steps using cross-attention-based fusion modules. We evaluate our method on the physically-realistic cluttered bin-picking dataset SynPick and the YCB-Video dataset and demonstrate improved pose estimation accuracy as well as better object detection accuracy. Arul Selvam Periyasamy, Sven Behnke |
ICRA | 2 |
| 2024 | Spiking CenterNet: A Distillation-boosted Spiking Neural Network for Object DetectionabstractIn the era of AI at the edge, self-driving cars, and climate change, the need for energy-efficient, small, embedded AI is growing. Spiking Neural Networks (SNNs) are a promising approach to address this challenge, with their event-driven information flow and sparse activations. We propose Spiking CenterNet for object detection on event data. It combines an SNN CenterNet adaptation with an efficient M2U-Net-based decoder. Our model significantly outperforms comparable previous work on Prophesee’s challenging GEN1 Automotive Detection Dataset while using less than half the energy. Distilling the knowledge of a non-spiking teacher into our SNN further increases performance. To the best of our knowledge, our work is the first approach that takes advantage of knowledge distillation in the field of spiking object detection. Lennard Bodden, Duc Bach Ha, Franziska Schwaiger, Lars Kreuzberg, Sven Behnke |
IJCNN | 5 |
| 2024 | HortiBot: An Adaptive Multi-Arm System for Robotic Horticulture of Sweet PeppersabstractHorticultural tasks such as pruning and selective harvesting are labor intensive and horticultural staff are hard to find. Automating these tasks is challenging due to the semi-structured greenhouse workspaces, changing environmental conditions such as lighting, dense plant growth with many occlusions, and the need for gentle manipulation of non-rigid plant organs. In this work, we present the three-armed system HortiBot, with two arms for manipulation and a third arm as an articulated head for active perception using stereo cameras. Its perception system detects not only peppers, but also peduncles and stems in real time, and performs online data association to build a world model of pepper plants. Collision-aware online trajectory generation allows all three arms to safely track their respective targets for observation, grasping, and cutting. We integrated perception and manipulation to perform selective harvesting of peppers and evaluated the system in lab experiments. Using active perception coupled with end-effector force torque sensing for compliant manipulation, HortiBot achieves high success rates in our indoor pepper plant mock-up. Christian Lenz, Rohit U. Menon, Michael Schreiber, Melvin Paul Jacob, Sven Behnke, Maren Bennewitz |
IROS | 5 |
| 2024 | RoboCup@Home 2024 OPL Winner NimbRo: Anthropomorphic Service Robots Using Foundation Models for Perception and Planning
Raphael Memmesheimer, Jan Nogga, Bastian Pätzold, Evgeny Kruzhkov, Simon Bultmann, Michael Schreiber, Jonas Bode, Bertan Karacora, Juhui Park, Alena Savinykh, Sven Behnke |
RoboCup | 11 |
| 2024 | Cleaning Robots in Public Spaces: A Survey and Proposal for Benchmarking Based on Stakeholders Interviews
Raphael Memmesheimer, Martina Overbeck, Björn Kral, Lea Steffen, Sven Behnke, Martin Gersch, Arne Roennau |
RoboCup | 5 |
| 2023 | PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces Using Permutohedral LatticesabstractNeural radiance-density field methods have become increasingly popular for the task of novel-view rendering. Their recent extension to hash-based positional encoding ensures fast training and inference with visually pleasing results. However, density-based methods struggle with recovering accurate surface geometry. Hybrid methods alleviate this issue by optimizing the density based on an underlying SDF. However, current SDF methods are overly smooth and miss fine geometric details. In this work, we combine the strengths of these two lines of work in a novel hash-based implicit surface representation. We propose improvements to the two areas by replacing the voxel hash encoding with a permutohedral lattice which optimizes faster, especially for higher dimensions. We additionally propose a regularization scheme which is crucial for recovering high-frequency geometric detail. We evaluate our method on multiple datasets and show that we can recover geometric detail at the level of pores and wrinkles while using only RGB images for supervision. Furthermore, using sphere tracing we can render novel views at 30 fps on an RTX 3090. Code is publicly available at https://radualexandru.github.io/permuto_sdf. Radu Alexandru Rosu, Sven Behnke |
CVPR | 2 |
| 2023 | Object-Centric Video Prediction Via Decoupling of Object Dynamics and InteractionsabstractWe present a framework for object-centric video prediction, i.e., parsing a video sequence into objects, and modeling their dynamics and interactions in order to predict the future object states from which video frames are rendered. To facilitate the learning of meaningful spatio-temporal object representations and forecasting of their states, we propose two novel object-centric video prediction (OCVP) transformer modules, which decouple the processing of temporal dynamics and object interactions. We show how OCVP predictors outperform object-agnostic video prediction models on two different datasets. Furthermore, we observe that OCVP modules learn consistent and interpretable object representations. Animations and code to reproduce our results can be found in our project website1. Angel Villar-Corrales, Ismail Wahdan, Sven Behnke |
ICIP | 3 |
| 2023 | External Camera-Based Mobile Robot Pose Estimation for Collaborative Perception with Smart Edge SensorsabstractWe present an approach for estimating a mobile robot's pose w.r.t. the allocentric coordinates of a network of static cameras using multi-view RGB images. The images are processed online, locally on smart edge sensors by deep neural networks to detect the robot and estimate 2D keypoints defined at distinctive positions of the 3D robot model. Robot keypoint detections are synchronized and fused on a central backend, where the robot's pose is estimated via multi-view minimization of reprojection errors. Through the pose estimation from external cameras, the robot's localization can be initialized in an allocentric map from a completely unknown state (kidnapped robot problem) and robustly tracked over time. We conduct a series of experiments evaluating the accuracy and robustness of the camera-based pose estimation compared to the robot's internal navigation stack, showing that our camera-based method achieves pose errors below 3 cm and 1° and does not drift over time, as the robot is localized allocentrically. With the robot's pose precisely estimated, its observations can be fused into the allocentric scene model. We show a real-world application, where observations from mobile robot and static smart edge sensors are fused to collaboratively build a 3D semantic map of a ~240 m2indoor environment. Simon Bultmann, Raphael Memmesheimer, Sven Behnke |
ICRA | 3 |
| 2023 | Dynamic Hybrid Locomotion and Jumping for Wheeled-Legged QuadrupedsabstractHybrid wheeled-legged quadrupeds have the potential to navigate challenging terrain with agility and speed and over long distances. However, obstacles can impede their progress by requiring the robots to either slow down to step over obstacles or modify their path to circumvent the obstacles. We propose a motion optimization framework for quadruped robots that incorporates non-steerable wheels and dynamic jumps, enabling them to perform hybrid wheeled-legged locomotion while overcoming obstacles without slowing down. Our approach involves a model predictive controller that uses a time-varying rigid body dynamics model of the robot, including legs and wheels, to track dynamic motions such as jumping. We also introduce a method for driving with minimal leg swings to reduce energy consumption by sparing the effort involved in lifting the wheels. Our method was tested successfully on the wheeled Mini Cheetah and the Unitree AlienGo robots. Further videos and results are available at https://www.ais.uni-bonn.de/∼hosseini/iros2023 Mojtaba Hosseini, Diego Rodriguez, Sven Behnke |
IROS | 3 |
| 2023 | Attention-Based VR Facial Animation with Visual Mouth Camera Guidance for Immersive Telepresence AvatarsabstractFacial animation in virtual reality environments is essential for applications that necessitate clear visibility of the user's face and the ability to convey emotional signals. In our scenario, we animate the face of an operator who controls a robotic Avatar system. The use of facial animation is particularly valuable when the perception of interacting with a specific individual, rather than just a robot, is intended. Purely keypoint-driven animation approaches struggle with the complexity of facial movements. We present a hybrid method that uses both keypoints and direct visual guidance from a mouth camera. Our method generalizes to unseen operators and requires only a quick enrolment step with capture of two short videos. Multiple source images are selected with the intention to cover different facial expressions. Given a mouth camera frame from the HMD, we dynamically construct the target keypoints and apply an attention mechanism to determine the importance of each source image. To resolve keypoint ambiguities and animate a broader range of mouth expressions, we propose to inject visual mouth camera information into the latent space. We enable training on large-scale speaking head datasets by simulating the mouth camera input with its perspective differences and facial deformations. Our method outperforms a baseline in quality, capability, and temporal consistency. In addition, we highlight how the facial animation contributed to our victory at the ANA Avatar XPRIZE Finals. Andre Rochow, Max Schwarz, Sven Behnke |
IROS | 3 |
| 2023 | Quadrupedal Footstep Planning Using Learned Motion Models of a Black-Box ControllerabstractLegged robots are increasingly entering new domains and applications, including search and rescue, inspection, and logistics. However, for such a systems to be valuable in real-world scenarios, they must be able to autonomously and robustly navigate irregular terrains. In many cases, robots that are sold on the market do not provide such abilities, being able to perform only blind locomotion. Furthermore, their controller cannot be easily modified by the end-user, requiring a new and time-consuming control synthesis. In this work, we present a fast local motion planning pipeline that extends the capabilities of a black-box walking controller that is only able to track high-level reference velocities. More precisely, we learn a set of motion models for such a controller that maps high-level velocity commands to Center of Mass (CoM) and footstep motions. We then integrate these models with a variant of the$A$* algorithm to plan the CoM trajectory, footstep sequences, and corresponding high-level velocity commands based on visual information, allowing the quadruped to safely traverse irregular terrains at demand. Ilyass Taouil, Giulio Turrisi, Daniel Schleich, Victor Barasuol, Claudio Semini, Sven Behnke |
IROS | 6 |
| 2023 | RoboCup 2023 Humanoid AdultSize Winner NimbRo: NimbRoNet3 Visual Perception and Responsive Gait with Waveform In-Walk Kicks
Dmytro Pavlichenko, Grzegorz Ficht, Angel Villar-Corrales, Luis Denninger, Julia Brocker, Tim Sinen, Michael Schreiber, Sven Behnke |
RoboCup | 8 |
| 2023 | Audio-Based Roughness Sensing and Tactile Feedback for Haptic Perception in TelepresenceabstractHaptic perception is highly important for immersive teleoperation of robots, especially for accomplishing manipulation tasks. We propose a low-cost haptic sensing and rendering system, which is capable of detecting and displaying surface roughness. As the robot fingertip moves across a surface of interest, two microphones capture sound coupled directly through the fingertip and through the air, respectively. A learning-based detector system analyzes the data in real time and gives roughness estimates with both high temporal resolution and low latency. Finally, an audio-based vibrational actuator displays the result to the human operator. We demonstrate the effectiveness of our system through lab experiments and our winning entry in the ANA Avatar XPRIZE competition finals, where briefly trained judges solved a roughness-based selection task even without additional vision feedback. We publish our dataset used for training and evaluation together with our trained models to enable reproducibility of results. Bastian Pätzold, Andre Rochow, Michael Schreiber, Raphael Memmesheimer, Christian Lenz, Max Schwarz, Sven Behnke |
SMC | 7 |
| 2022 | MSPred: Video Prediction at Multiple Spatio-Temporal Scales with Hierarchical Recurrent Networks
Angel Villar-Corrales, Ani Karapetyan, Andreas Boltres, Sven Behnke |
BMVC | 4 |
| 2022 | Neural Strands: Learning Hair Geometry and Appearance from Multi-view Images
Radu Alexandru Rosu, Shunsuke Saito, Chenglei Wu, Sven Behnke, Giljoo Nam |
ECCV (33) | 5 |
| 2022 | Intention-Aware Frequency Domain Transformer Networks for Video Prediction
Hafez Farazi, Sven Behnke |
ICANN (4) | 2 |
| 2022 | Real-Robot Deep Reinforcement Learning: Improving Trajectory Tracking of Flexible-Joint Manipulator with Reference CorrectionabstractFlexible-joint manipulators are governed by complex nonlinear dynamics, defining a challenging control problem. In this work, we propose an approach to learn an outer-loop joint trajectory tracking controller with deep reinforcement learning. The controller represented by a stochastic policy is learned in under two hours directly on the real robot. This is achieved through bounded reference correction actions and use of a model-free off-policy learning method. In addition, an informed policy initialization is proposed, where the agent is pre-trained in a learned simulation. We test our approach on the 7 DOF manipulator of a Baxter robot. We demonstrate that the proposed method is capable of consistent learning across multiple runs when applied directly on the real robot. Our method yields a policy which significantly improves the trajectory tracking accuracy in comparison to the vendor-provided controller, generalizing to an unseen payload. Dmytro Pavlichenko, Sven Behnke |
ICRA | 2 |
| 2022 | Abstract Flow for Temporal Semantic Segmentation on the Permutohedral LatticeabstractSemantic segmentation is a core ability required by autonomous agents, as being able to distinguish which parts of the scene belong to which object class is crucial for navigation and interaction with the environment. Approaches which use only one time-step of data cannot distinguish between moving objects nor can they benefit from temporal integration. In this work, we extend a backbone LatticeNet to process temporal point cloud data. Additionally, we take inspiration from optical flow methods and propose a new module called Abstract Flow which allows the network to match parts of the scene with similar abstract features and gather the information temporally. We obtain state-of-the-art results on the SemanticKITTI dataset that contains LiDAR scans from real urban environments. We share the PyTorch implementation of TemporalLatticeNet at https://github.com/AIS-Bonn/temporal_latticenet. Peer Schütt, Radu Alexandru Rosu, Sven Behnke |
ICRA | 3 |
| 2022 | Predicting Physical Object Properties from VideoabstractWe present a novel approach to estimating physical properties of objects from video. Our approach consists of a physics engine and a correction estimator. Starting from the initial observed state, object behavior is simulated forward in time. Based on the simulated and observed behavior, the correction estimator then determines refined physical parameters for each object. The method can be iterated for increased precision. Our approach is generic, as it allows for the use of an arbitrary-not necessarily differentiable-physics engine and correction estimator. For the latter, we evaluate both gradient-free hyperparameter optimization and a deep convolutional neural network. We demonstrate faster and more robust convergence of the learned method in several simulated 2D scenarios focusing on bin situations. Martin Link, Max Schwarz, Sven Behnke |
IJCNN | 3 |
| 2022 | NeuralMVS: Bridging Multi-View Stereo and Novel View SynthesisabstractMulti-View Stereo (MVS) is a core task in 3D computer vision. With the surge of novel deep learning methods, learned MVS achieves more complete depth maps than classical approaches, but still relies on building a memory intensive dense cost volume. Novel View Synthesis (NVS) is a parallel line of research and has recently seen an increase in popularity with Neural Radiance Field (NeRF) models, which optimize a per scene radiance field. However, NeRF methods do not generalize to novel scenes and are slow to train and test. We propose to bridge the gap between these two methodologies with a novel network that can recover 3D scene geometry as a distance function, together with high-resolution color images. Our method uses only a sparse set of images as input and can generalize well to novel scenes. Additionally, we propose a coarse-to-fine sphere tracing approach in order to significantly increase speed. We show on various datasets that our method reaches comparable accuracy to per-scene optimized methods while being able to generalize and running significantly faster. Radu Alexandru Rosu, Sven Behnke |
IJCNN | 2 |
| 2022 | VR Facial Animation for Immersive Telepresence AvatarsabstractVR Facial Animation is necessary in applications requiring clear view of the face, even though a VR headset is worn. In our case, we aim to animate the face of an operator who is controlling our robotic avatar system. We propose a real-time capable pipeline with very fast adaptation for specific operators. In a quick enrollment step, we capture a sequence of source images from the operator without the VR headset which contain all the important operator-specific appearance information. During inference, we then use the operator keypoint information extracted from a mouth camera and two eye cameras to estimate the target expression and head pose, to which we map the appearance of a source still image. In order to enhance the mouth expression accuracy, we dynamically select an auxiliary expression frame from the captured sequence. This selection is done by learning to transform the current mouth keypoints into the source camera space, where the alignment can be determined accurately. We, furthermore, demonstrate an eye tracking pipeline that can be trained in less than a minute, a time efficient way to train the whole pipeline given a dataset that includes only complete faces, show exemplary results generated by our method, and discuss performance at the ANA Avatar XPRIZE semifinals. Andre Rochow, Max Schwarz, Michael Schreiber, Sven Behnke |
IROS | 4 |
| 2022 | Predictive Angular Potential Field-based Obstacle Avoidance for Dynamic UAV FlightsabstractIn recent years, unmanned aerial vehicles (UAVs) are used for numerous inspection and video capture tasks. Manually controlling UAVs in the vicinity of obstacles is challenging, however, and poses a high risk of collisions. Even for autonomous flight, global navigation planning might be too slow to react to newly perceived obstacles. Disturbances such as wind might lead to deviations from the planned trajectories. In this work, we present a fast predictive obstacle avoidance method that does not depend on higher-level localization or mapping and maintains the dynamic flight capabilities of UAVs. It directly operates on LiDAR range images in real time and adjusts the current flight direction by computing angular potential fields within the range image. The velocity magnitude is subsequently determined based on a trajectory prediction and time-to-contact estimation. Our method is evaluated using Hardware-in-the-Loop simulations. It keeps the UAV at a safe distance to obstacles, while allowing higher flight velocities than previous reactive obstacle avoidance methods that directly operate on sensor data. Daniel Schleich, Sven Behnke |
IROS | 2 |
| 2022 | A Study on the Ambiguity in Human Annotation of German Oral History Interviews for Perceived Emotion Recognition and Sentiment AnalysisabstractFor research in audiovisual interview archives often it is not only of interest what is said but also how. Sentiment analysis and emotion recognition can help capture, categorize and make these different facets searchable. In particular, for oral history archives, such indexing technologies can be of great interest. These technologies can help understand the role of emotions in historical remembering. However, humans often perceive sentiments and emotions ambiguously and subjectively. Moreover, oral history interviews have multi-layered levels of complex, sometimes contradictory, sometimes very subtle facets of emotions. Therefore, the question arises of the chance machines and humans have capturing and assigning these into predefined categories. This paper investigates the ambiguity in human perception of emotions and sentiment in German oral history interviews and the impact on machine learning systems. Our experiments reveal substantial differences in human perception for different emotions. Furthermore, we report from ongoing machine learning experiments with different modalities. We show that the human perceptual ambiguity and other challenges, such as class imbalance and lack of training data, currently limit the opportunities of these technologies for oral history archives. Nonetheless, our work uncovers promising observations and possibilities for further research. Michael Gref, Nike Matthiesen, Sreenivasa Hikkal Venugopala, Shalaka Satheesh, Aswinkumar Vijayananth, Duc Bach Ha, Sven Behnke, Joachim Köhler |
LREC | 7 |
| 2022 | RoboCup 2022 AdultSize Winner NimbRo: Upgraded Perception, Capture Steps Gait and Phase-Based In-Walk Kicks
Dmytro Pavlichenko, Grzegorz Ficht, Arash Amini, Mojtaba Hosseini, Raphael Memmesheimer, Angel Villar-Corrales, Stefan M. Schulz, Marcell Missura, Maren Bennewitz, Sven Behnke |
RoboCup | 10 |
| 2022 | Benchmarking table recognition performance on biomedical literature on neurological disordersabstractMOTIVATION: Table recognition systems are widely used to extract and structure quantitative information from the vast amount of documents that are increasingly available from different open sources. While many systems already perform well on tables with a simple layout, tables in the biomedical domain are often much more complex. Benchmark and training data for such tables are however very limited. RESULTS: To address this issue, we present a novel, highly curated benchmark dataset based on a hand-curated literature corpus on neurological disorders, which can be used to tune and evaluate table extraction applications for this challenging domain. We evaluate several state-of-the-art table extraction systems based on our proposed benchmark and discuss challenges that emerged during the benchmark creation as well as factors that can impact the performance of recognition methods. For the evaluation procedure, we propose a new metric as well as several improvements that result in a better performance evaluation. AVAILABILITY AND IMPLEMENTATION: The resulting benchmark dataset (https://zenodo.org/record/5549977) as well as the source code to our novel evaluation approach can be openly accessed. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tim Adams, Marcin Namysl, Alpha Tom Kodamullil, Sven Behnke, Marc Jacobs 0001 |
Bioinform. | 4 |
| 2021 | Semantic Prediction: Which One Should Come First, Recognition or Prediction?abstractThe ultimate goal of video prediction is not forecasting future pixel-values given some previous frames.Rather, the end goal of video prediction is to discover valuable internal representations from the vast amount of available unlabeled video data in a self-supervised fashion for downstream tasks.One of the primary downstream tasks is interpreting the scene's semantic composition and using it for decision-making.For example, by predicting human movements, an observer can anticipate human activities and collaborate in a shared workspace.There are two main ways to achieve the same outcome, given a pre-trained video prediction and pre-trained semantic extraction model; one can first apply predictions and then extract semantics or first extract semantics and then predict.We investigate these configurations using the Local Frequency Domain Transformer Network (LFDTN) as the video prediction model and U-Net as the semantic extraction model on synthetic and real datasets. Hafez Farazi, Jan Nogga, Sven Behnke |
ESANN | 3 |
| 2021 | Fourier-based Video Prediction through Relational Object MotionabstractThe ability to predict future outcomes conditioned on observed video frames is crucial for intelligent decision-making in autonomous systems.Recently, deep recurrent architectures have been applied to the task of video prediction.However, this often results in blurry predictions and requires tedious training on large datasets.Here, we explore a different approach by (1) using frequency-domain approaches for video prediction and (2) explicitly inferring object-motion relationships in the observed scene.The resulting predictions are consistent with the observed dynamics in a scene and do not suffer from blur. Malte Mosbach, Sven Behnke |
ESANN | 2 |
| 2021 | DeepWalk: Omnidirectional Bipedal Gait by Deep Reinforcement LearningabstractBipedal walking is one of the most difficult but exciting challenges in robotics. The difficulties arise from the complexity of high-dimensional dynamics, sensing and actuation limitations combined with real-time and computational constraints. Deep Reinforcement Learning (DRL) holds the promise to address these issues by fully exploiting the robot dynamics with minimal craftsmanship. In this paper, we propose a novel DRL approach that enables an agent to learn omnidirectional locomotion for humanoid (bipedal) robots. Notably, the locomotion behaviors are accomplished by a single control policy (a single neural network). We achieve this by introducing a new curriculum learning method that gradually increases the task difficulty by scheduling target velocities. In addition, our method does not require reference motions which facilities its application to robots with different kinematics, and reduces the overall complexity. Finally, different strategies for sim-to-real transfer are presented which allow us to transfer the learned policy to a real humanoid robot. Diego Rodriguez, Sven Behnke |
ICRA | 2 |
| 2021 | Search-based Planning of Dynamic MAV Trajectories Using Local Multiresolution State LatticesabstractSearch-based methods that use motion primitives can incorporate the system's dynamics into the planning and thus generate dynamically feasible MAV trajectories that are globally optimal. However, searching high-dimensional state lattices is computationally expensive. Local multiresolution is a commonly used method to accelerate spatial path planning. While paths within the vicinity of the robot are represented at high resolution, the representation gets coarser for more distant parts. In this work, we apply the concept of local multiresolution to high-dimensional state lattices that include velocities and accelerations. Experiments show that our proposed approach significantly reduces planning times. Thus, it increases the applicability to large dynamic environments, where frequent replanning is necessary. Daniel Schleich, Sven Behnke |
ICRA | 2 |
| 2021 | Local Frequency Domain Transformer Networks for Video PredictionabstractVideo prediction is commonly referred to as forecasting future frames of a video sequence provided several past frames thereof. It remains a challenging domain as visual scenes evolve according to complex underlying dynamics, such as the camera's egocentric motion or the distinct motility per individual object viewed. These are mostly hidden from the observer and manifest as often highly non-linear transformations between consecutive video frames. Therefore, video prediction is of interest not only in anticipating visual changes in the real world but has, above all, emerged as an unsupervised learning rule targeting the formation and dynamics of the observed environment. Many of the deep learning-based state-of-the-art models for video prediction utilize some form of recurrent layers like Long Short-Term Memory (LSTMs) or Gated Recurrent Units (GRUs) at the core of their models. Although these models can predict the future frames, they rely entirely on these recurrent structures to simultaneously perform three distinct tasks: extracting transformations, projecting them into the future, and transforming the current frame. In order to completely interpret the formed internal representations, it is crucial to disentangle these tasks. This paper proposes a fully differentiable building block that can perform all of those tasks separately while maintaining interpretability. We derive the relevant theoretical foundations and showcase results on synthetic as well as real data. We demonstrate that our method is readily extended to perform motion segmentation and account for the scene's composition, and learns to produce reliable predictions in an entirely interpretable manner by only observing unlabeled video data. Hafez Farazi, Jan Nogga, Sven Behnke |
IJCNN | 3 |
| 2021 | Mapless Humanoid Navigation Using Learned Latent DynamicsabstractIn this paper, we propose a novel Deep Reinforcement Learning approach to address the mapless navigation problem, in which the locomotion actions of a humanoid robot are taken online based on the knowledge encoded in learned models. Planning happens by generating open-loop trajectories in a learned latent space that captures the dynamics of the environment. Our planner considers visual (RGB images) and non-visual observations (e.g., attitude estimations). This confers the agent upon awareness not only of the scenario, but also of its own state. In addition, we incorporate a termination likelihood predictor model as an auxiliary loss function of the control policy, which enables the agent to anticipate terminal states of success and failure. In this manner, the sample efficiency of the approach for episodic tasks is increased. Our model is evaluated on the NimbRo-OP2X humanoid robot that navigates in scenes avoiding collisions efficiently in simulation and with the real hardware. André Brandenburger, Diego Rodriguez, Sven Behnke |
IROS | 3 |
| 2021 | Real-time Multi-Adaptive-Resolution-Surfel 6D LiDAR Odometry using Continuous-time Trajectory OptimizationabstractSimultaneous Localization and Mapping (SLAM) is an essential capability for autonomous robots, but due to high data rates of 3D LiDARs real-time SLAM is challenging. We propose a real-time method for 6D LiDAR odometry. Our approach combines a continuous-time B-Spline trajectory representation with a Gaussian Mixture Model (GMM) formulation to jointly align local multi-resolution surfel maps. Sparse voxel grids and permutohedral lattices ensure fast access to map surfels, and an adaptive resolution selection scheme effectively speeds up registration. A thorough experimental evaluation shows the performance of our approach on multiple datasets and during real-robot experiments. Jan Quenzel, Sven Behnke |
IROS | 2 |
| 2021 | NimbRo Avatar: Interactive Immersive Telepresence with Force-Feedback TelemanipulationabstractRobotic avatars promise immersive teleoperation with human-like manipulation and communication capabilities. We present such an avatar system, based on the key components of immersive 3D visualization and transparent force-feedback telemanipulation. Our avatar robot features an anthropomorphic bimanual arm configuration with dexterous hands. The remote human operator drives the arms and fingers through an exoskeleton-based operator station, which provides force feedback both at the wrist and for each finger. The robot torso is mounted on a holonomic base, providing locomotion capability in typical indoor scenarios, controlled using a 3D rudder device. Finally, the robot features a 6D movable head with stereo cameras, which stream images to a VR HMD worn by the operator. Movement latency is hidden using spherical rendering. The head also carries a telepresence screen displaying a synthesized image of the operator with facial animation, which enables direct interaction with remote persons. We evaluate our system successfully both in a user study with untrained operators as well as a longer and more complex integrated mission. We discuss lessons learned from the trials and possible improvements. Max Schwarz, Christian Lenz, Andre Rochow, Michael Schreiber, Sven Behnke |
IROS | 5 |
| 2021 | Real-Time Pose Estimation from Images for Multiple Humanoid Robots
Arash Amini, Hafez Farazi, Sven Behnke |
RoboCup | 3 |
| 2021 | 6D Object Pose Estimation Using Keypoints and Part Affinity Fields
Moritz Zappel, Simon Bultmann, Sven Behnke |
RoboCup | 3 |
| 2020 | NAT: Noise-Aware Training for Robust Neural Sequence LabelingabstractSequence labeling systems should perform reliably not only under ideal conditions but also with corrupted inputs-as these systems often process user-generated text or follow an errorprone upstream component.To this end, we formulate the noisy sequence labeling problem, where the input may undergo an unknown noising process and propose two Noise-Aware Training (NAT) objectives that improve robustness of sequence labeling performed on perturbed input: Our data augmentation method trains a neural model using a mixture of clean and noisy samples, whereas our stability training algorithm encourages the model to create a noise-invariant latent representation.We employ a vanilla noise model at training time.For evaluation, we use both the original data and its variants perturbed with real OCR errors and misspellings.Extensive experiments on English and German named entity recognition benchmarks confirmed that NAT consistently improved robustness of popular sequence labeling models, preserving accuracy on the original input.We make our code and data publicly available for the research community. Marcin Namysl, Sven Behnke, Joachim Köhler |
ACL | 2 |
| 2020 | Motion Segmentation using Frequency Domain Transformer Networks
Hafez Farazi, Sven Behnke |
ESANN | 2 |
| 2020 | Object-centered Fourier Motion Estimation and Segment-Transformation Prediction
Moritz Wolter, Angela Yao, Sven Behnke |
ESANN | 3 |
| 2020 | 3D U-Net for Segmentation of Plant Root MRI Images in Super-Resolution
Yi Zhao 0012, Nils Wandel, Magdalena Landl, Andrea Schnepf, Sven Behnke |
ESANN | 5 |
| 2020 | Robust Skeletonization for Plant Root Structure Reconstruction from MRIabstractStructural reconstruction of plant roots from MRI is challenging, because of low resolution and low signal-to-noise ratio of the 3D measurements which may lead to disconnectivities and wrongly connected roots. We propose a two-stage approach for this task. The first stage is based on semantic root vs. soil segmentation and finds lowest-cost paths from any root voxel to the shoot. The second stage takes the largest fully connected component generated in the first stage and uses 3D skeletonization to extract a graph structure. We evaluate our method on 22 MRI scans and compare to human expert reconstructions. Jannis Horn, Yi Zhao 0012, Nils Wandel, Magdalena Landl, Andrea Schnepf, Sven Behnke |
ICPR | 6 |
| 2020 | Fast Whole-Body Motion Control of Humanoid Robots with Inertia ConstraintsabstractWe introduce a new, analytical method for generating whole-body motions for humanoid robots, which approximate the desired Composite Rigid Body (CRB) inertia. Our approach uses a reduced five mass model, where four of the masses are attributed to the limbs and one is used for the trunk. This compact formulation allows for finding an analytical solution that combines the kinematics with mass distribution and inertial properties of a humanoid robot. The positioning of the masses in Cartesian space is then directly used to obtain joint angles with relations based on simple geometry. Motions are achieved through the time evolution of poses generated through the desired foot positioning and CRB inertia properties. As a result, we achieve short computation times in the order of tens of microseconds. This makes the method suited for applications with limited computation resources, or leaving them to be spent on higher-layer tasks such as model predictive control. The approach is evaluated by performing a dynamic kicking motion with an igus®Humanoid Open Platform robot. Grzegorz Ficht, Sven Behnke |
ICRA | 2 |
| 2020 | Beyond Photometric Consistency: Gradient-based Dissimilarity for Improving Visual Odometry and Stereo MatchingabstractPose estimation and map building are central ingredients of autonomous robots and typically rely on the registration of sensor data. In this paper, we investigate a new metric for registering images that builds upon on the idea of the photometric error. Our approach combines a gradient orientation-based metric with a magnitude-dependent scaling term. We integrate both into stereo estimation as well as visual odometry systems and show clear benefits for typical disparity and direct image registration tasks when using our proposed metric. Our experimental evaluation indicate that our metric leads to more robust and more accurate estimates of the scene depth as well as camera trajectory. Thus, the metric improves camera pose estimation and in turn the mapping capabilities of mobile robots. We believe that a series of existing visual odometry and visual SLAM systems can benefit from the findings reported in this paper. Jan Quenzel, Radu Alexandru Rosu, Thomas Läbe, Cyrill Stachniss, Sven Behnke |
ICRA | 5 |
| 2020 | Stillleben: Realistic Scene Synthesis for Deep Learning in RoboticsabstractTraining data is the key ingredient for deep learning approaches, but difficult to obtain for the specialized domains often encountered in robotics. We describe a synthesis pipeline capable of producing training data for cluttered scene perception tasks such as semantic segmentation, object detection, and correspondence or pose estimation. Our approach arranges object meshes in physically realistic, dense scenes using physics simulation. The arranged scenes are rendered using high-quality rasterization with randomized appearance and material parameters. Noise and other transformations introduced by the camera sensors are simulated. Our pipeline can be run online during training of a deep neural network, yielding applications in life-long learning and in iterative render-and-compare approaches. We demonstrate the usability by learning semantic segmentation on the challenging YCB-Video dataset without actually using any training frames, where our method achieves performance comparable to a conventionally trained model. Additionally, we show successful application in a real-world regrasping system. Max Schwarz, Sven Behnke |
ICRA | 2 |
| 2020 | Category-Level 3D Non-Rigid Registration from Single-View RGB ImagesabstractIn this paper, we propose a novel approach to solve the 3D non-rigid registration problem from RGB images using Convolutional Neural Networks (CNNs). Our objective is to find a deformation field (typically used for transferring knowledge between instances, e.g., grasping skills) that warps a given 3D canonical model into a novel instance observed by a single-view RGB image. This is done by training a CNN that infers a deformation field for the visible parts of the canonical model and by employing a learned shape (latent) space for inferring the deformations of the occluded parts. As result of the registration, the observed model is reconstructed. Because our method does not need depth information, it can register objects that are typically hard to perceive with RGB-D sensors, e.g. with transparent or shiny surfaces. Even without depth data, our approach outperforms the Coherent Point Drift (CPD) registration method for the evaluated object categories. Diego Rodriguez, Florian Huber 0003, Sven Behnke |
IROS | 3 |
| 2020 | Multi-Staged Cross-Lingual Acoustic Model Adaption for Robust Speech Recognition in Real-World Applications - A Case Study on German Oral History InterviewsabstractWhile recent automatic speech recognition systems achieve remarkable performance when large amounts of adequate, high quality annotated speech data is used for training, the same systems often only achieve an unsatisfactory result for tasks in domains that greatly deviate from the conditions represented by the training data. For many real-world applications, there is a lack of sufficient data that can be directly used for training robust speech recognition systems. To address this issue, we propose and investigate an approach that performs a robust acoustic model adaption to a target domain in a cross-lingual, multi-staged manner. Our approach enables the exploitation of large-scale training data from other domains in both the same and other languages. We evaluate our approach using the challenging task of German oral history interviews, where we achieve a relative reduction of the word error rate by more than 30% compared to a model trained from scratch only on the target domain, and 6-7% relative compared to a model trained robustly on 1000 hours of same-language out-of-domain training data. Michael Gref, Oliver Walter, Christoph Schmidt 0001, Sven Behnke, Joachim Köhler |
LREC | 4 |
| 2020 | Semi-supervised Semantic Mapping Through Label Propagation with Semantic Texture Meshes
Radu Alexandru Rosu, Jan Quenzel, Sven Behnke |
Int. J. Comput. Vis. | 3 |
| 2019 | Interpretable and Fine-Grained Visual Explanations for Convolutional Neural NetworksabstractTo verify and validate networks, it is essential to gain insight into their decisions, limitations as well as possible shortcomings of training data. In this work, we propose a post-hoc, optimization based visual explanation method, which highlights the evidence in the input image for a specific prediction. Our approach is based on a novel technique to defend against adversarial evidence (i.e. faulty evidence due to artefacts) by filtering gradients during optimization. The defense does not depend on human-tuned parameters. It enables explanations which are both fine-grained and preserve the characteristics of images, such as edges and colors. The explanations are interpretable, suited for visualizing detailed evidence and can be tested as they are valid model inputs. We qualitatively and quantitatively evaluate our approach on a multitude of models and datasets. Jan Mathias Köhler, Tobias Gindele, Leon Hetzel, Thaddäus Wiedemer, Sven Behnke |
CVPR | 6 |
| 2019 | Complex Valued Gated Auto-encoder for Video Frame Prediction
Niloofar Azizi, Nils Wandel, Sven Behnke |
ESANN | 3 |
| 2019 | Frequency Domain Transformer Networks for Video Prediction
Hafez Farazi, Sven Behnke |
ESANN | 2 |
| 2019 | Learning super-resolution 3D segmentation of plant root MRI images from few examples
Ali Oguz Uzman, Jannis Horn, Sven Behnke |
ESANN | 3 |
| 2019 | SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesabstractSemantic scene understanding is important for various applications. In particular, self-driving cars need a fine-grained understanding of the surfaces and objects in their vicinity. Light detection and ranging (LiDAR) provides precise geometric information about the environment and is thus a part of the sensor suites of almost all self-driving cars. Despite the relevance of semantic scene understanding for this application, there is a lack of a large dataset for this task which is based on an automotive LiDAR. In this paper, we introduce a large dataset to propel research on laser-based semantic segmentation. We annotated all sequences of the KITTI Vision Odometry Benchmark and provide dense point-wise annotations for the complete 360-degree field-of-view of the employed automotive LiDAR. We propose three benchmark tasks based on this dataset: (i) semantic segmentation of point clouds using a single scan, (ii) semantic segmentation using multiple past scans, and (iii) semantic scene completion, which requires to anticipate the semantic scene in the future. We provide baseline experiments and show that there is a need for more sophisticated models to efficiently tackle these tasks. Our dataset opens the door for the development of more advanced methods, but also provides plentiful data to investigate new research directions. Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, Juergen Gall |
ICCV | 5 |
| 2019 | Two-Staged Acoustic Modeling Adaption for Robust Speech Recognition by the Example of German Oral History InterviewsabstractIn automatic speech recognition, often little training data is available for specific challenging tasks, but training of state-of-the-art automatic speech recognition systems requires large amounts of annotated speech. To address this issue, we propose a two-staged approach to acoustic modeling that combines noise and reverberation data augmentation with transfer learning to robustly address challenges such as difficult acoustic recording conditions, spontaneous speech, and speech of elderly people. We evaluate our approach using the example of German oral history interviews, where a relative average reduction of the word error rate by 19.3% is achieved. Michael Gref, Christoph Schmidt 0001, Sven Behnke, Joachim Köhler |
ICME | 3 |
| 2019 | Towards Learning Abstract Representations for Locomotion Planning in High-dimensional State SpacesabstractGround robots which are able to navigate a variety of terrains are needed in many domains. One of the key aspects is the capability to adapt to the ground structure, which can be realized through movable body parts coming along with additional degrees of freedom (DoF). However, planning respective locomotion is challenging since suitable representations result in large state spaces. Employing an additional abstract representation-which is coarser, lower-dimensional, and semantically enriched-can support the planning. While a desired robot representation and action set of such an abstract representation can be easily defined, the cost function requires large tuning efforts. We propose a method to represent the cost function as a CNN. Training of the network is done on generated artificial data, while it generalizes well to the abstraction of real world scenes. We further apply our method to the problem of search-based planning of hybrid driving-stepping locomotion. The abstract representation is used as a powerful informed heuristic which accelerates planning by multiple orders of magnitude. Tobias Klamt, Sven Behnke |
ICRA | 2 |
| 2019 | Search-based 3D Planning and Trajectory Optimization for Safe Micro Aerial Vehicle Flight Under Sensor Visibility ConstraintsabstractSafe navigation of Micro Aerial Vehicles (MAVs) requires not only obstacle-free flight paths according to a static environment map, but also the perception of and reaction to previously unknown and dynamic objects. This implies that the onboard sensors cover the current flight direction. Due to the limited payload of MAVs, full sensor coverage of the environment has to be traded off with flight time. Thus, often only a part of the environment is covered. We present a combined allocentric complete planning and trajectory optimization approach taking these sensor visibility constraints into account. The optimized trajectories yield flight paths within the apex angle of a Velodyne Puck Lite 3D laser scanner enabling low-level collision avoidance to perceive obstacles in the flight direction. Furthermore, the optimized trajectories take the flight dynamics into account and contain the velocities and accelerations along the path. We evaluate our approach with a DJI Matrice 600 MAV and in simulation employing hardware-in-the-loop. Matthias Nieuwenhuisen, Sven Behnke |
ICRA | 2 |
| 2019 | Detection and Tracking of Small Objects in Sparse 3D Laser Range DataabstractDetection and tracking of dynamic objects is a key feature for autonomous behavior in a continuously changing environment. With the increasing popularity and capability of micro aerial vehicles (MAVs) efficient algorithms have to be utilized to enable multi object tracking on limited hardware and data provided by lightweight sensors. We present a novel segmentation approach based on a combination of median filters and an efficient pipeline for detection and tracking of small objects within sparse point clouds generated by a Velodyne VLP-16 sensor. We achieve real-time performance on a single core of our MAV hardware by exploiting the inherent structure of the data. Our approach is evaluated on simulated and real scans of in- and outdoor environments, obtaining results comparable to the state of the art. Additionally, we provide an application for filtering the dynamic and mapping the static part of the data, generating further insights into the performance of the pipeline on unlabeled data. Jan Razlaw, Jan Quenzel, Sven Behnke |
ICRA | 3 |
| 2019 | Fast Time-optimal Avoidance of Moving Obstacles for High-Speed MAV FlightabstractIn this work, we propose a method to efficiently compute smooth, time-optimal trajectories for micro aerial vehicles (MAVs) evading a moving obstacle. Our approach first computes an n-dimensional trajectory from the start- to an arbitrary target state including position, velocity and acceleration. It respects input- and state-constraints and is thus dynamically feasible. The trajectory is then efficiently checked for collisions, exploiting the piecewise polynomial formulation. If collisions occur, viastates are inserted into the trajectory to circumvent the obstacle and still maintain time-optimality. These viastates are described by position, velocity, and acceleration. The evaluation shows that the computational demands of the proposed method are minimal such that obstacle avoidance can begin within few milliseconds. Optimality of generated trajectories, combined with the ability for frequent online re-planning from non-hover initial conditions, make the approach well suited for evasion of suddenly perceived obstacles during fast flight. Marius Beul, Sven Behnke |
IROS | 2 |
| 2019 | Directional TSDF: Modeling Surface Orientation for Coherent MeshesabstractReal-time 3D reconstruction from RGB-D sensor data plays an important role in many robotic applications, such as object modeling and mapping. The popular method of fusing depth information into a truncated signed distance function (TSDF) and applying the marching cubes algorithm for mesh extraction has severe issues with thin structures: not only does it lead to loss of accuracy, but it can generate completely wrong surfaces. To address this, we propose the directional TSDF-a novel representation that stores opposite surfaces separate from each other. The marching cubes algorithm is modified accordingly to retrieve a coherent mesh representation. We further increase the accuracy by using surface gradient-based ray casting for fusing new measurements. We show that our method outperforms state-of-the-art TSDF reconstruction algorithms in mesh accuracy. Malte Splietker, Sven Behnke |
IROS | 2 |
| 2019 | A VR System for Immersive Teleoperation and Live Exploration with a Mobile RobotabstractApplications like disaster management and industrial inspection often require experts to enter contaminated places. To circumvent the need for physical presence, it is desirable to generate a fully immersive individual live teleoperation experience. However, standard video-based approaches suffer from a limited degree of immersion and situation awareness due to the restriction to the camera view, which impacts the navigation. In this paper, we present a novel VR-based practical system for immersive robot teleoperation and scene exploration. While being operated through the scene, a robot captures RGB-D data that is streamed to a SLAM-based live multiclient telepresence system. Here, a global 3D model of the already captured scene parts is reconstructed and streamed to the individual remote user clients where the rendering for e.g. head-mounted display devices (HMDs) is performed. We introduce a novel lightweight robot client component which transmits robot-specific data and enables a quick integration into existing robotic systems. This way, in contrast to first- person exploration systems, the operators can explore and navigate in the remote site completely independent of the current position and view of the capturing robot, complementing traditional input devices for teleoperation. We provide a proof-of-concept implementation and demonstrate the capabilities as well as the performance of our system regarding interactive object measurements and bandwidth-efficient data streaming and visualization. Furthermore, we show its benefits over purely video-based teleoperation in a user study revealing a higher degree of situation awareness and a more precise navigation in challenging environments. Patrick Stotko, Stefan Krumpen, Max Schwarz, Christian Lenz, Sven Behnke, Reinhard Klein, Michael Weinmann |
IROS | 5 |
| 2019 | Utilizing Temporal Information in Deep Convolutional Network for Efficient Soccer Ball Detection and Tracking
Anna Kukleva, Mohammad Asif Khan 0001, Hafez Farazi, Sven Behnke |
RoboCup | 4 |
| 2019 | RoboCup 2019 AdultSize Winner NimbRo: Deep Learning Perception, In-Walk Kick, Push Recovery, and Team Play Capabilities
Diego Rodriguez, Hafez Farazi, Grzegorz Ficht, Dmytro Pavlichenko, André Brandenburger, Mojtaba Hosseini, Oleg Kosenko, Michael Schreiber, Marcell Missura, Sven Behnke |
RoboCup | 10 |
| 2018 | Functionally Modular and Interpretable Temporal Filtering for Robust Segmentation
Volker Fischer 0003, Michael Herman, Sven Behnke |
BMVC | 4 |
| 2018 | Hierarchical Recurrent Filtering for Fully Convolutional DenseNets
Volker Fischer 0003, Michael Herman, Sven Behnke |
ESANN | 4 |
| 2018 | Location Dependency in Video Prediction
Niloofar Azizi, Hafez Farazi, Sven Behnke |
ICANN (3) | 3 |
| 2018 | Efficient Continuous-Time SLAM for 3D Lidar-Based Online MappingabstractModern 3D laser-range scanners have a high data rate, making online simultaneous localization and mapping (SLAM) computationally challenging. Recursive state estimation techniques are efficient but commit to a state estimate immediately after a new scan is made, which may lead to misalignments of measurements. We present a 3D SLAM approach that allows for refining alignments during online mapping. Our method is based on efficient local mapping and a hierarchical optimization back-end. Measurements of a 3D laser scanner are aggregated in local multiresolution maps by means of surfel-based registration. The local maps are used in a multi-level graph for allocentric mapping and localization. In order to incorporate corrections when refining the alignment, the individual 3D scans in the local map are modeled as a sub-graph and graph optimization is performed to account for drift and misalignments in the local maps. Furthermore, in each sub-graph, a continuous-time representation of the sensor trajectory allows to correct measurements between scan poses. We evaluate our approach in multiple experiments by showing qualitative results. Furthermore, we quantify the map quality by an entropy-based measure. David Droeschel, Sven Behnke |
ICRA | 2 |
| 2018 | Planning Hybrid Driving-Stepping Locomotion on Multiple Levels of AbstractionabstractNavigating in search and rescue environments is challenging, since a variety of terrains has to be considered. Hybrid driving-stepping locomotion, as provided by our robot Momaro, is a promising approach. Similar to other locomotion methods, it incorporates many degrees of freedom - offering high flexibility but making planning computationally expensive for larger environments. We propose a navigation planning method, which unifies different levels of representation in a single planner. In the vicinity of the robot, it provides plans with a fine resolution and a high robot state dimensionality. With increasing distance from the robot, plans become coarser and the robot state dimensionality decreases. We compensate this loss of information by enriching coarser representations with additional semantics. Experiments show that the proposed planner provides plans for large, challenging scenarios in feasible time. Tobias Klamt, Sven Behnke |
ICRA | 2 |
| 2018 | Transferring Grasping Skills to Novel Instances by Latent Space Non-Rigid RegistrationabstractRobots acting in open environments need to be able to handle novel objects. Based on the observation that objects within a category are often similar in their shapes and usage, we propose an approach for transferring grasping skills from known instances to novel instances of an object category. Correspondences between the instances are established by means of a non-rigid registration method that combines the Coherent Point Drift approach with subspace methods. The known object instances are modeled using a canonical shape and a transformation which deforms it to match the instance shape. The principle axes of variation of these deformations define a low-dimensional latent space. New instances can be generated through interpolation and extrapolation in this shape space. For inferring the shape parameters of an unknown instance, an energy function expressed in terms of the latent variables is minimized. Due to the class-level knowledge of the object, our method is able to complete novel shapes from partial views. Control poses for generating grasping motions are transferred efficiently to novel instances by the estimated non-rigid transformation. Diego Rodriguez, Corbin Cogswell, Seongyong Koo, Sven Behnke |
ICRA | 4 |
| 2018 | Fast Object Learning and Dual-arm Coordination for Cluttered Stowing, Picking, and PackingabstractRobotic picking from cluttered bins is a demanding task, for which Amazon Robotics holds challenges. The 2017 Amazon Robotics Challenge (ARC) required stowing items into a storage system, picking specific items, and packing them into boxes. In this paper, we describe the entry of team NimbRo Picking. Our deep object perception pipeline can be quickly and efficiently adapted to new items using a custom turntable capture system and transfer learning. It produces high-quality item segments, on which grasp poses are found. A planning component coordinates manipulation actions between two robot arms, minimizing execution time. The system has been demonstrated successfully at ARC, where our team reached second places in both the picking task and the final stow-and-pick task. We also evaluate individual components. Max Schwarz, Christian Lenz, Germán Martín García, Seongyong Koo, Arul Selvam Periyasamy, Michael Schreiber, Sven Behnke |
ICRA | 7 |
| 2018 | Fused Angles and the Deficiencies of Euler AnglesabstractJust like the well-established Euler angles representation, fused angles are a convenient parameterisation for rotations in three-dimensional Euclidean space. They were developed in the context of balancing bodies, most specifically walking bipedal robots, but have since found wider application due to their useful properties. A comparative analysis between fused angles and Euler angles is presented in this paper, delineating the specific differences between the two representations that make fused angles more suitable for representing orientations in balance-related scenarios. Aspects of comparison include the locations of the singularities, the associated parameter sensitivities, the level of mutual independence of the parameters, and the axisymmetry of the parameters. Philipp Allgeuer, Sven Behnke |
IROS | 2 |
| 2018 | Supervised Autonomous Locomotion and Manipulation for Disaster Response with a Centaur-Like RobotabstractMobile manipulation tasks are one of the key challenges in the field of search and rescue (SAR) robotics requiring robots with flexible locomotion and manipulation abilities. Since the tasks are mostly unknown in advance, the robot has to adapt to a wide variety of terrains and workspaces during a mission. The centaur-like robot Centauro has a hybrid legged-wheeled base and an anthropomorphic upper body to carry out complex tasks in environments too dangerous for humans. Due to its high number of degrees of freedom, controlling the robot with direct teleoperation approaches is challenging and exhausting. Supervised autonomy approaches are promising to increase quality and speed of control while keeping the flexibility to solve unknown tasks. We developed a set of operator assistance functionalities with different levels of autonomy to control the robot for challenging locomotion and manipulation tasks. The integrated system was evaluated in disaster response scenarios and showed promising performance. Tobias Klamt, Diego Rodriguez, Max Schwarz, Christian Lenz, Dmytro Pavlichenko, David Droeschel, Sven Behnke |
IROS | 7 |
| 2018 | Robust 6D Object Pose Estimation in Cluttered Scenes Using Semantic Segmentation and Pose Regression NetworksabstractObject pose estimation is a crucial prerequisite for robots to perform autonomous manipulation in clutter. Real-world bin-picking settings such as warehouses present additional challenges, e.g., new objects are added constantly. Most of the existing object pose estimation methods assume that 3D models of the objects is available beforehand. We present a pipeline that requires minimal human intervention and circumvents the reliance on the availability of 3D models by a fast data acquisition method and a synthetic data generation procedure. This work builds on previous work on semantic segmentation of cluttered bin-picking scenes to isolate individual objects in clutter. An additional network is trained on synthetic scenes to estimate object poses from a cropped object-centered encoding extracted from the segmentation results. The proposed method is evaluated on a synthetic validation dataset and cluttered realworld scenes. Arul Selvam Periyasamy, Max Schwarz, Sven Behnke |
IROS | 3 |
| 2018 | Keyframe-Based Photometric Online Calibration and Color CorrectionabstractFinding the parameters of a vignetting function for a camera currently involves the acquisition of several images in a given scene under very controlled lighting conditions, a cumbersome and error-prone task where the end result can only be confirmed visually. Many computer vision algorithms assume photoconsistency, constant intensity between scene points in different images, and tend to perform poorly if this assumption is violated. We present a real-time online vignetting and response calibration with additional exposure estimation for global-shutter color cameras. Our method does not require uniformly illuminated surfaces, known texture or specific geometry. The only assumptions are that the camera is moving, the illumination is static and reflections are Lambertian. Our method estimates the camera view poses by sparse visual SLAM and models the vignetting function by a small number of thin plate splines (TPS) together with a sixth-order polynomial to provide a dense estimation of attenuation from sparsely sampled scene points. The camera response function (CRF) is jointly modeled by a TPS and a Gamma curve. We evaluate our approach on synthetic datasets and in real-world scenarios with reference data from a Structure-from-Motion (SfM) system. We show clear visual improvement on textured meshes without the need for extensive meshing algorithms. A useful calibration is obtained from a few keyframes which makes an on-the-fly deployment conceivable. Jan Quenzel, Jannis Horn, Sebastian Houben, Sven Behnke |
IROS | 4 |
| 2018 | NimbRo Robots Winning RoboCup 2018 Humanoid AdultSize Soccer Competitions
Hafez Farazi, Grzegorz Ficht, Philipp Allgeuer, Dmytro Pavlichenko, Diego Rodriguez, André Brandenburger, Mojtaba Hosseini, Sven Behnke |
RoboCup | 8 |
| 2018 | Combining Simulations and Real-Robot Experiments for Bayesian Optimization of Bipedal Gait Stabilization
Diego Rodriguez, André Brandenburger, Sven Behnke |
RoboCup | 3 |
| 2017 | Learning Semantic Prediction using Pretrained Deep Feedforward Networks
Volker Fischer 0003, Michael Herman, Sven Behnke |
ESANN | 4 |
| 2017 | NimbRo picking: Versatile part handling for warehouse automationabstractPart handling in warehouse automation is challenging if a large variety of items must be accommodated and items are stored in unordered piles. To foster research in this domain, Amazon holds picking challenges. We present our system which achieved second and third place in the Amazon Picking Challenge 2016 tasks. The challenge required participants to pick a list of items from a shelf or to stow items into the shelf. Using two deep-learning approaches for object detection and semantic segmentation and one item model registration method, our system localizes the requested item. Manipulation occurs using suction on points determined heuristically or from 6D item model registration. Parametrized motion primitives are chained to generate motions. We present a full-system evaluation during the APC 2016 and component-level evaluations of the perception system on an annotated dataset. Max Schwarz, Anton Milan, Christian Lenz, Aura Munoz, Arul Selvam Periyasamy, Michael Schreiber, Sebastian Schüller, Sven Behnke |
ICRA | 8 |
| 2017 | Online visual robot tracking and identification using deep LSTM networksabstractCollaborative robots working on a common task are necessary for many applications. One of the challenges for achieving collaboration in a team of robots is mutual tracking and identification. We present a novel pipeline for online vision-based detection, tracking and identification of robots with a known and identical appearance. Our method runs in realtime on the limited hardware of the observer robot. Unlike previous works addressing robot tracking and identification, we use a data-driven approach based on recurrent neural networks to learn relations between sequential inputs and outputs. We formulate the data association problem as multiple classification problems. A deep LSTM network was trained on a simulated dataset and fine-tuned on small set of real data. Experiments on two challenging datasets, one synthetic and one real, which include long-term occlusions, show promising results. Hafez Farazi, Sven Behnke |
IROS | 2 |
| 2017 | Anytime hybrid driving-stepping locomotion planningabstractHybrid driving-stepping locomotion is an effective approach for navigating in a variety of environments. Long, sufficiently even distances can be quickly covered by driving while obstacles can be overcome by stepping. Our quadruped robot Momaro, with steerable pairs of wheels located at the end of each of its compliant legs, allows such locomotion. Planning respective paths attracted only little attention so far. We propose a navigation planning method which generates hybrid locomotion paths. The planner chooses driving mode whenever possible and takes into account the detailed robot footprint. If steps are required, the planner includes those. To accelerate planning, steps are planned first as abstract manoeuvres and are expanded afterwards into detailed motion sequences. Our method ensures at all times that the robot stays stable. Experiments show that the proposed planner is capable of providing paths in feasible time, even for challenging terrain. Tobias Klamt, Sven Behnke |
IROS | 2 |
| 2017 | Efficient stochastic multicriteria arm trajectory optimizationabstractPerforming manipulation with robotic arms requires a method for planning trajectories that takes multiple factors into account: collisions, joint limits, orientation constraints, torques, and duration of a trajectory. We present an approach to efficiently optimize arm trajectories with respect to multiple criteria. Our work extends Stochastic Trajectory Optimization for Motion Planning (STOMP). We optimize trajectory duration by including velocity into the optimization. We propose an efficient cost function with normalized components, which allows prioritizing components depending on user-specified requirements. Optimization is done in two stages: first with a partial cost function and in the second stage with full costs. We compare our method to state-of-the art methods. In addition, we perform experiments on real robots: centaur-like robot Momaro and an industrial manipulator. Dmytro Pavlichenko, Sven Behnke |
IROS | 2 |
| 2017 | Online depth calibration for RGB-D cameras using visual SLAMabstractModern consumer RGB-D cameras are affordable and provide dense depth estimates at high frame rates. Hence, they are popular for building dense environment representations. Yet, the sensors often do not provide accurate depth estimates since the factory calibration exhibits a static deformation. We present a novel approach to online depth calibration that uses a visual SLAM system as reference for the measured depth. A sparse map is generated and the visual information is used to correct the static deformation of the measured depth while missing data is extrapolated using a small number of thin plate splines (TPS). The corrected depth can then be used to improve the accuracy of the sparse RGB-D map and the 3D environment reconstruction. As more data becomes available, the depth calibration is updated on the fly. Our method does not rely on a planar geometry like walls or a one-to-one-pixel correspondence between color and depth camera. Our approach is evaluated in real-world scenarios and against ground truth data. Comparison against two popular self-calibration methods is performed. Furthermore, we show clear visual improvement on aggregated point clouds with our method. Jan Quenzel, Radu Alexandru Rosu, Sebastian Houben, Sven Behnke |
IROS | 4 |
| 2017 | Grown-Up NimbRo Robots Winning RoboCup 2017 Humanoid AdultSize Soccer Competitions
Grzegorz Ficht, Dmytro Pavlichenko, Philipp Allgeuer, Hafez Farazi, Diego Rodriguez, André Brandenburger, Johannes Kürsch, Michael Schreiber, Sven Behnke |
RoboCup | 9 |
| 2017 | Advanced Soccer Skills and Team Play of RoboCup 2017 TeenSize Winner NimbRo
Diego Rodriguez, Hafez Farazi, Philipp Allgeuer, Dmytro Pavlichenko, Grzegorz Ficht, André Brandenburger, Johannes Kürsch, Sven Behnke |
RoboCup | 8 |
| 2017 | Object class segmentation of RGB-D video using recurrent convolutional neural networks
Mircea Serban Pavel, Hannes Schulz, Sven Behnke |
Neural Networks | 3 |
| 2016 | Multispectral Pedestrian Detection using Deep Fusion Convolutional Neural Networks
Volker Fischer 0003, Michael Herman, Sven Behnke |
ESANN | 4 |
| 2016 | Semantic segmentation priors for object discoveryabstractReliable object discovery in realistic indoor scenes is a necessity for many computer vision and service robot applications. In these scenes, semantic segmentation methods have made huge advances in recent years. Such methods can provide useful prior information for object discovery by removing false positives and by delineating object boundaries. We propose a novel method that combines bottom-up object discovery and semantic priors for producing generic object candidates in RGB-D images. We use a deep learning method for semantic segmentation to classify colour and depth superpixels into meaningful categories. Separately for each category, we use saliency to estimate the location and scale of objects, and superpixels to find their precise boundaries. Finally, object candidates of all categories are combined and ranked. We evaluate our approach on the NYU Depth V2 dataset and show that we outperform other state-of-the-art object discovery methods in terms of recall. Germán Martín García, Farzad Husain, Hannes Schulz, Simone Frintrop, Carme Torras, Sven Behnke |
ICPR | 6 |
| 2016 | Focused online visual-motor coordination for a dual-arm robot manipulatorabstractCoordination between visual sensors and robot manipulators is necessary for successful manipulation. This paper proposes a novel visual-motor coordination method that performs online parameter estimation of an RGB-D camera mounted in a robot head without any external markers. Through self-observation of a dual-arm robot manipulator, the method updates parameters to reduce the discrepancy between observed point cloud data and 3D mesh models of the current robot con guration. With the estimated parameters at each time step, visual data is adjusted to the focused workspace of the 14DOF dual-arm robot manipulator. The online and realtime algorithm was developed by using a GPU-based particle filtering method. Experimental results show that our method outperforms state-of-the-art offline registration methods in terms of accuracy and computation time. We also analyzed the dependence of the results on prior parameters to demonstrate the online capability of our method. Seongyong Koo, Sven Behnke |
ICRA | 2 |
| 2016 | Hybrid driving-stepping locomotion with the wheeled-legged robot MomaroabstractLocomotion in uneven terrain is important for a wide range of robotic applications, including Search&Rescue operations. Our mobile manipulation robot Momaro features a unique locomotion design consisting of four legs ending in pairs of steerable wheels, allowing the robot to omnidirectionally drive on sufficiently even terrain, step over obstacles, and also to overcome height differences by climbing. We demonstrate the feasibility and usefulness of this design on the example of the DARPA Robotics Challenge, where our team NimbRo Rescue solved seven out of eight tasks in only 34 minutes. We also introduce a method for semi-autonomous execution of weight-shifting and stepping actions based on a 2D heightmap generated from 3D laser data. Max Schwarz, Tobias Rodehutskors, Michael Schreiber, Sven Behnke |
ICRA | 4 |
| 2016 | Efficient multi-camera visual-inertial SLAM for micro aerial vehiclesabstractVisual SLAM is an area of vivid research and bears countless applications for moving robots. In particular, micro aerial vehicles benefit from visual sensors due to their low weight. Their motion is, however, often faster and more complex than that of ground-based robots which is why systems with multiple cameras are currently evaluated and deployed. This, in turn, drives the computational demand for visual SLAM algorithms. Sebastian Houben, Jan Quenzel, Nicola Krombach, Sven Behnke |
IROS | 4 |
| 2016 | Local multiresolution trajectory optimization for micro aerial vehicles employing continuous curvature transitionsabstractComplex indoor and outdoor missions for autonomous micro aerial vehicles (MAV) require fast generation of collision-free paths in 3D space. Often not all obstacles in an environment are known prior to the mission execution. Consequently, the ability for replanning during a flight is key for success. Our approach locally optimizes trajectories of grid-based path planning. It preserves obstacle-freeness of the path and ensures smoothness with continuous curvature transition segments. Fast optimization and frequent reoptimization is made possible by means of local multiresolution time discretization. With our extensions, high dimensional flight trajectories incorporating velocities and accelerations can be planned with a time discretization of 100 Hz within the prediction horizon of the underlying controller. Matthias Nieuwenhuisen, Sven Behnke |
IROS | 2 |
| 2016 | First International HARTING Open Source Prize Winner: The igus Humanoid Open Platform
Philipp Allgeuer, Grzegorz Ficht, Hafez Farazi, Michael Schreiber, Sven Behnke |
RoboCup | 5 |
| 2016 | MRSLaserMap: Local Multiresolution Grids for Efficient 3D Laser Mapping and Localization
David Droeschel, Sven Behnke |
RoboCup | 2 |
| 2016 | RoboCup 2016 Humanoid TeenSize Winner NimbRo: Robust Visual Perception and Soccer Behaviors
Hafez Farazi, Philipp Allgeuer, Grzegorz Ficht, André Brandenburger, Dmytro Pavlichenko, Michael Schreiber, Sven Behnke |
RoboCup | 7 |
| 2016 | Real-Time Visual Tracking and Identification for a Team of Homogeneous Humanoid Robots
Hafez Farazi, Sven Behnke |
RoboCup | 2 |
| 2015 | Depth and height aware semantic RGB-D perception with convolutional neural networks
Hannes Schulz, Nico Höft, Sven Behnke |
ESANN | 3 |
| 2015 | A skill-based system for object perception and manipulation for automating kitting tasksabstractThe automation of kitting tasks-collecting a set of parts for one particular car into a kit-has a huge impact in the automotive industry. It considerably increases the automation levels of tasks typically conducted by human workers. Collecting the parts involves picking up objects from pallets and bins as well as placing them in the respective compartments of the kitting box. In this paper, we present a complete system for automated kitting with a mobile manipulator thereby focusing on depalletizing tasks and placing. In order to allow for low cycle times, we present particularly efficient solutions to object perception as well as motion planning and execution. For easy portability to different platforms, all components are integrated into a skill-based framework that is tightly coupled with a task planning component. We present results of experiments at both a research laboratory environment and at the industrial site of PSA Peugeot Citroën serving as a proof of concept for the overall system design and implementation. Dirk Holz, Angeliki Topalidou-Kyniazopoulou, Francesco Rovida, Mikkel Rath Pedersen, Volker Krüger, Sven Behnke |
ETFA | 6 |
| 2015 | RGB-D object recognition and pose estimation based on pre-trained convolutional neural network featuresabstractObject recognition and pose estimation from RGB-D images are important tasks for manipulation robots which can be learned from examples. Creating and annotating datasets for learning is expensive, however. We address this problem with transfer learning from deep convolutional neural networks (CNN) that are pre-trained for image categorization and provide a rich, semantically meaningful feature set. We incorporate depth information, which the CNN was not trained with, by rendering objects from a canonical perspective and colorizing the depth channel according to distance from the object center. We evaluate our approach on the Washington RGB-D Objects dataset, where we find that the generated feature set naturally separates classes and instances well and retains pose manifolds. We outperform state-of-the-art on a number of subtasks and show that our approach can yield superior results when only little training data is available. Max Schwarz, Hannes Schulz, Sven Behnke |
ICRA | 3 |
| 2015 | Recurrent convolutional neural networks for object-class segmentation of RGB-D videoabstractObject-class segmentation is a computer vision task which requires labeling each pixel of an image with the class of the object it belongs to. Deep convolutional neural networks (DNN) are able to learn and exploit local spatial correlations required for this task. They are, however, restricted by their small, fixed-sized filters, which limits their ability to learn long-range dependencies. Recurrent Neural Networks (RNN), on the other hand, do not suffer from this restriction. Their iterative interpretation allows them to model long-range dependencies by propagating activity. This property might be especially useful when labeling video sequences, where both spatial and temporal long-range dependencies occur. In this work, we propose novel RNN architectures for object-class segmentation. We investigate three ways to consider past and future context in the prediction process by comparing networks that process the frames one by one with networks that have access to the whole sequence. We evaluate our models on the challenging NYU Depth v2 dataset for object-class segmentation and obtain competitive results. Mircea Serban Pavel, Hannes Schulz, Sven Behnke |
IJCNN | 3 |
| 2015 | Fused angles: A representation of body orientation for balanceabstractThe parameterisation of rotations in three dimensional Euclidean space is an area of applied mathematics that has long been studied, dating back to the original works of Euler in the 18thcentury. As such, many ways of parameterising a rotation have been developed over the years. Motivated by the task of representing the orientation of a balancing body, the fused angles parameterisation is developed and introduced in this paper. This novel representation is carefully defined both mathematically and geometrically, and thoroughly investigated in terms of the properties it possesses, and how it relates to other existing representations. A second intermediate representation, tilt angles, is also introduced as a natural consequence thereof. Philipp Allgeuer, Sven Behnke |
IROS | 2 |
| 2015 | Real-time object detection, localization and verification for fast robotic depalletizingabstractDepalletizing is a challenging task for manipulation robots. Key to successful application are not only robustness of the approach, but also achievable cycle times in order to keep up with the rest of the process. In this paper, we propose a system for depalletizing and a complete pipeline for detecting and localizing objects as well as verifying that the found object does not deviate from the known object model, e.g., if it is not the object to pick. In order to achieve high robustness (e.g., with respect to different lighting conditions) and generality with respect to the objects to pick, our approach is based on multi-resolution surfel models. All components (both software and hardware) allow operation at high frame rates and, thus, allow for low cycle times. In experiments, we demonstrate depalletizing of automotive and other prefabricated parts with both high reliability (w.r.t. success rates) and efficiency (w.r.t. low cycle times). Dirk Holz, Angeliki Topalidou-Kyniazopoulou, Jörg Stückler, Sven Behnke |
IROS | 4 |
| 2015 | Gradient-driven online learning of bipedal push recoveryabstractBipedal walking is a complex and dynamic whole-body motion with balance constraints. Due to the inherently unstable inverted pendulum-like dynamics of walking, the design of robust walking controllers proved to be particularly challenging. While a controller could potentially be learned with a robot in the loop, the destructive nature of losing balance and the impracticality of a high number of repetitions render most existing learning methods unsuitable for an online learning setting with real hardware. We propose a model-driven learning method that enables a humanoid robot to quickly learn how to maintain its balance. We bootstrap the learning process with a central pattern generator for stepping motions that abstracts from the complexity of the walking motion and simplifies the problem setting to the learning of a small number of leg swing amplitude parameters. A simple physical model that represents the dominant dynamics of bipedal walking estimates an approximate gradient and suggests how to modify the swing amplitude to restore balance. In experiments with a real robot, we show that only a few failed steps are sufficient for our biped to learn strong push recovery skills in the sagittal direction. Marcell Missura, Sven Behnke |
IROS | 2 |
| 2015 | Efficient Dense Rigid-Body Motion Segmentation and Estimation in RGB-D Video
Jörg Stückler, Sven Behnke |
Int. J. Comput. Vis. | 2 |
| 2015 | Corrigendum to "Multi-resolution surfel maps for efficient dense 3D modeling and tracking" [J. Visual Communication and Image Representation 25 (1) (2014) 137-147]
Jörg Stückler, Sven Behnke |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Two-layer contractive encodings for learning stable nonlinear features
Hannes Schulz, Kyunghyun Cho, Tapani Raiko, Sven Behnke |
Neural Networks | 4 |
| 2014 | Mobile teleoperation interfaces with adjustable autonomy for personal service robotsabstractPersonal service robots require a comprehensive set of perception, control, and planning skills to perform everyday tasks autonomously. While achieving full autonomy is an ongoing research topic, first real-world applications of personal robots may come into reach, if state-of-the-art autonomous capabilities are combined with the intelligence of the users in a complementary way. We report on handheld user interfaces for personal robots that allow for teleoperating the robot on three levels of autonomy: body, skill, and task control. On the higher levels, autonomous behavior of the robot relieves the user from significant workload. If autonomous execution fails, or autonomous functionality is not provided by the robot system, the user can select a lower level of autonomy to solve a task. The benefits of providing adjustable autonomy in teleoperation have been successfully demonstrated at [email protected] competitions. Max Schwarz, Jörg Stückler, Sven Behnke |
HRI | 3 |
| 2014 | Structured Prediction for Object Detection in Deep Neural Networks
Hannes Schulz, Sven Behnke |
ICANN | 2 |
| 2014 | Local multi-resolution representation for 6D motion estimation and mapping with a continuously rotating 3D laser scannerabstractMicro aerial vehicles (MAV) pose a challenge in designing sensory systems and algorithms due to their size and weight constraints and limited computing power. We present an efficient 3D multi-resolution map that we use to aggregate measurements from a lightweight continuously rotating laser scanner. We estimate the robot's motion by means of visual odometry and scan registration, aligning consecutive 3D scans with an incrementally built map. By using local multi-resolution, we gain computational efficiency by having a high resolution in the near vicinity of the robot and a lower resolution with increasing distance from the robot, which correlates with the sensor's characteristics in relative distance accuracy and measurement density. Compared to uniform grids, local multi-resolution leads to the use of fewer grid cells without loosing information and consequently results in lower computational costs. We efficiently and accurately register new 3D scans with the map in order to estimate the motion of the MAV and update the map in-flight. In experiments, we demonstrate superior accuracy and efficiency of our registration approach compared to state-of-the-art methods such as GICP. Our approach builds an accurate 3D obstacle map and estimates the vehicle's trajectory in real-time. David Droeschel, Jörg Stückler, Sven Behnke |
ICRA | 3 |
| 2014 | Bayesian exploration and interactive demonstration in continuous state MAXQ-learningabstractDeploying robots for service tasks requires learning algorithms that scale to the combinatorial complexity of our daily environment. Inspired by the way humans decompose complex tasks, hierarchical methods for robot learning have attracted significant interest. In this paper, we apply the MAXQ method for hierarchical reinforcement learning to continuous state spaces. By using Gaussian Process Regression for MAXQ value function decomposition, we obtain probabilistic estimates of primitive and completion values for every subtask within the MAXQ hierarchy. From these, we recursively compute probabilistic estimates of state-action values. Based on the expected deviation of these estimates, we devise a Bayesian exploration strategy that balances optimization of expected values and risk from exploring unknown actions. To further reduce risk and to accelerate learning, we complement MAXQ with learning from demonstrations in an interactive way. In every situation and subtask, the system may ask for a demonstration if there is not enough knowledge available to determine a safe action for exploration. We demonstrate the ability of the proposed system to efficiently learn solutions to complex tasks on a box stacking scenario. Kathrin Gräve, Sven Behnke |
ICRA | 2 |
| 2014 | Learning depth-sensitive conditional random fields for semantic segmentation of RGB-D imagesabstractWe present a structured learning approach to semantic annotation of RGB-D images. Our method learns to reason about spatial relations of objects and fuses low-level class predictions to a consistent interpretation of a scene. Our model incorporates color, depth and 3D scene features, on which an energy function is learned to directly optimize object class prediction using the loss-based maximum-margin principle of structural support vector machines. We evaluate our approach on the NYU V2 dataset of indoor scenes, a challenging dataset covering a wide variety of scene layouts and object classes. We hard-code much less information about the scene layout into our model then previous approaches, and instead learn object relations directly from the data. We find that our conditional random field approach improves upon previous work, setting a new state-of-the-art for the dataset. Andreas C. Müller 0001, Sven Behnke |
ICRA | 2 |
| 2014 | Efficient deformable registration of multi-resolution surfel maps for object manipulation skill transferabstractEndowing mobile manipulation robots with skills to use objects and tools often involves the programming or training on specific object instances. To apply this knowledge to novel instances from the same class of objects, a robot requires generalization capabilities for control as well as perception. In this paper, we propose an efficient approach to deformable registration of RGB-D images that enables robots to transfer skills between object instances. Our method provides a dense deformation field between the current image and an object model which allows for estimating local rigid transformations on the object's surface. Since we define grasp and motion strategies as poses and trajectories with respect to the object models, these strategies can be transferred to novel instances through local transformations derived from the deformation field. In experiments, we demonstrate the accuracy and runtime efficiency of our registration method. We also report on the use of our skill transfer approach in a public demonstration. Jörg Stückler, Sven Behnke |
ICRA | 2 |
| 2014 | RoboCup Humanoid League Rule Developments 2002-2014 and Future Perspectives
Jacky Baltes, Soroush Sadeghnejad, Daniel Seifert, Sven Behnke |
RoboCup | 4 |
| 2014 | Balanced Walking with Capture Steps
Marcell Missura, Sven Behnke |
RoboCup | 2 |
| 2014 | Cosero, Find My Keys! Object Localization and Retrieval Using Bluetooth Low Energy Tags
David Schwarz, Max Schwarz, Jörg Stückler, Sven Behnke |
RoboCup | 4 |
| 2014 | PyStruct: learning structured prediction in python
Andreas C. Müller 0001, Sven Behnke |
J. Mach. Learn. Res. | 2 |
| 2014 | Multi-resolution surfel maps for efficient dense 3D modeling and tracking
Jörg Stückler, Sven Behnke |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Efficient Dense 3D Rigid-Body Motion Segmentation in RGB-D VideoabstractMotion is a fundamental segmentation cue in video.Many current approaches segment 3D motion in monocular or stereo image sequences, mostly relying on sparse interest points or being dense but computationally demanding.We propose an efficient expectation-maximization (EM) framework for dense 3D segmentation of moving rigid parts in RGB-D video.Our approach segments two images into pixel regions that undergo coherent 3D rigid-body motion.Our formulation treats background and foreground objects equally and poses no further assumptions on the motion of the camera or the objects than rigidness.While our EM-formulation is not restricted to a specific image representation, we supplement it with efficient image representation and registration for rapid segmentation of RGB-D video.In experiments we demonstrate that our approach recovers segmentation and 3D motion at good precision. Jörg Stückler, Sven Behnke |
BMVC | 2 |
| 2013 | Combining contour and shape primitives for object detection and pose estimation of prefabricated partsabstractMan-made objects such as mechanical construction parts can typically be described as a composition of shape primitives like cylinders, planes, cones and spheres. We propose a robust method for the detection and pose estimation of such objects in 3D point clouds. Our main contribution is to enhance a probabilistic graph-matching approach that detects objects using 3D shape primitives with distinct 2D primitives such as circular contours. With this extension, our method copes with difficult occlusion situations and can be applied for object manipulation in complex scenarios such as grasping from a pile or bin-picking. We demonstrate the performance of our approach in a comparison with a state-of-the-art feature-based method for objects of generic shape and a primitive-based approach using only 3D shapes and no contours. Alexander Berner, Jun Li 0042, Dirk Holz, Jörg Stückler, Sven Behnke, Reinhard Klein |
ICIP | 5 |
| 2013 | Two-Layer Contractive Encodings with Shortcuts for Semi-supervised Learning
Hannes Schulz, Kyunghyun Cho, Tapani Raiko, Sven Behnke |
ICONIP (1) | 4 |
| 2013 | Mobile bin picking with an anthropomorphic service robotabstractGrasping individual objects from an unordered pile in a box has been investigated in static scenarios so far. In this paper, we demonstrate bin picking with an anthropomorphic mobile robot. To this end, we extend global navigation techniques by precise local alignment with a transport box. Objects are detected in range images using a shape primitive-based approach. Our approach learns object models from single scans and employs active perception to cope with severe occlusions. Grasps and arm motions are planned in an efficient local multiresolution height map. All components are integrated and evaluated in a bin picking and part delivery task. Matthias Nieuwenhuisen, David Droeschel, Dirk Holz, Jörg Stückler, Alexander Berner, Jun Li 0042, Reinhard Klein, Sven Behnke |
ICRA | 8 |
| 2013 | Hierarchical Object Discovery and Dense Modelling From Motion Cues in RGB-D Video
Jörg Stückler, Sven Behnke |
IJCAI | 2 |
| 2013 | Learning sequential tasks interactively from demonstrations and own experienceabstractDeploying robots to our day-to-day life requires them to have the ability to learn from their environment in order to acquire new task knowledge and to flexibly adapt existing skills to various situations. For typical real-world tasks, it is not sufficient to endow robots with a set of primitive actions. Rather, they need to learn how to sequence these in order to achieve a desired effect on their environment. In this paper, we propose an intuitive learning method for a robot to acquire sequences of motions by combining learning from human demonstrations and reinforcement learning. In every situation, our approach treats both ways of learning as alternative control flows to optimally exploit their strengths without inheriting their shortcomings. Using a Gaussian Process approximation of the state-action sequence value function, our approach generalizes values observed from demonstrated and autonomously generated action sequences to unknown inputs. This approximation is based on a kernel we designed to account for different representations of tasks and action sequences as well as inputs of variable length. From the expected deviation of value estimates, we devise a greedy exploration policy following a Bayesian optimization criterion that quickly converges learning to promising action sequences while protecting the robot from sequences with unpredictable outcome. We demonstrate the ability of our approach to efficiently learn appropriate action sequences in various situations on a manipulation task involving stacked boxes. Kathrin Gräve, Sven Behnke |
IROS | 2 |
| 2013 | Learning to Improve Capture Steps for Disturbance Rejection in Humanoid Soccer
Marcell Missura, Cedrick Münstermann, Philipp Allgeuer, Max Schwarz, Julio Pastrana, Sebastian Schüller, Michael Schreiber, Sven Behnke |
RoboCup | 8 |
| 2013 | Compliant Robot Behavior Using Servo Actuator Models Identified by Iterative Learning Control
Max Schwarz, Sven Behnke |
RoboCup | 2 |
| 2013 | Humanoid TeenSize Open Platform NimbRo-OP
Max Schwarz, Julio Pastrana, Philipp Allgeuer, Michael Schreiber, Sebastian Schüller, Marcell Missura, Sven Behnke |
RoboCup | 7 |
| 2013 | Increasing Flexibility of Mobile Manipulation and Intuitive Human-Robot Interaction in RoboCup@Home
Jörg Stückler, David Droeschel, Kathrin Gräve, Dirk Holz, Michael Schreiber, Angeliki Topalidou-Kyniazopoulou, Max Schwarz, Sven Behnke |
RoboCup | 8 |
| 2012 | Model Learning and Real-Time Tracking Using Multi-Resolution Surfel MapsabstractFor interaction with its environment, a robot is required to learn models of objects and to perceive these models in the livestreams from its sensors. In this paper, we propose a novel approach to model learning and real-time tracking. We extract multi-resolution 3D shape and texture representations from RGB-D images at high frame-rates. An efficient variant of the iterative closest points algorithm allows for registering maps in real-time on a CPU. Our approach learns full-view models of objects in a probabilistic optimization framework in which we find the best alignment between multiple views. Finally, we track the pose of the camera with respect to the learned model by registering the current sensor view to the model. We evaluate our approach on RGB-D benchmarks and demonstrate its accuracy, efficiency, and robustness in model learning and tracking. We also report on the successful public demonstration of our approach in a mobile manipulation task. Jörg Stückler, Sven Behnke |
AAAI | 2 |
| 2012 | Learning Object-Class Segmentation with Convolutional Neural Networks
Hannes Schulz, Sven Behnke |
ESANN | 2 |
| 2012 | Learning Two-Layer Contractive Encodings
Hannes Schulz, Sven Behnke |
ICANN (1) | 2 |
| 2012 | Incremental action recognition and generalizing motion generation based on goal-directed featuresabstractThe ability to recognize human actions is a fundamental problem in many areas of robotics research concerned with human-robot interaction or learning from human demonstration. In this paper, we present a new integrated approach to identifying and recognizing actions in human movement sequences and their reproduction in unknown situations. We propose a set of task-space features to construct probabilistic models of action classes. Based on this representation, we suggest a combined segmentation and classification algorithm which processes data non-greedily using an incremental lookahead to reliably locate transitions between actions. In a programming by demonstration scenario, our action models afford the generalization and reproduction of learned movements to previously unseen situations. To evaluate the performance of our approach, we consider typical manipulation tasks in a table top setting. In a sequence of human demonstrations, our approach successfully extracts and recognizes actions from different classes and subsequently generalizes them to unknown situations. Kathrin Gräve, Sven Behnke |
IROS | 2 |
| 2012 | Semantic mapping using object-class segmentation of RGB-D imagesabstractFor task planning and execution in unstructured environments, a robot needs the ability to recognize and localize relevant objects. When this information is made persistent in a semantic map, it can be used, e. g., to communicate with humans. In this paper, we propose a novel approach to learning such maps. Our approach registers measurements of RGB-D cameras by means of simultaneous localization and mapping. We employ random decision forests to segment object classes in images and exploit dense depth measurements to obtain scale-invariance. Our object recognition method integrates shape and texture seamlessly. The probabilistic segmentation from multiple views is filtered in a voxel-based 3D map using a Bayesian framework. We report on the quality of our object-class segmentation method and demonstrate the benefits in accuracy when fusing multiple views in a semantic map. Jörg Stückler, Nenad Biresev, Sven Behnke |
IROS | 3 |
| 2012 | Adjustable autonomy for mobile teleoperation of personal service robotsabstractAbstract — Controlling personal service robots is an impor-tant, complex task which must be accomplished by the users themselves. In this paper, we propose a novel user interface for personal service robots that allows the user to teleoperate the robot using handheld computers. The user can adjust the autonomy of the robot between three levels: body control, skill control, and task control. We design several user interfaces for teleoperation on these levels. On the higher levels, au-tonomous behavior of the robot relieves the user from significant workload. If the autonomous execution fails, or autonomous functionality is not provided by the robot system, the user can select a lower level of autonomy, e.g., direct body control, to solve a task. In a qualitative user study we evaluate usability aspects of our teleoperation interface with our domestic service robots Cosero and Dynamaid. We demonstrate the benefits of providing adjustable manual and autonomous control. I. Sebastian Muszynski, Jörg Stückler, Sven Behnke |
RO-MAN | 3 |
| 2012 | Lateral Disturbance Rejection for the Nao Robot
Juan José Alcaraz-Jiménez, Marcell Missura, Humberto Martínez Barberá, Sven Behnke |
RoboCup | 4 |
| 2012 | RoboCup 2012 Best Humanoid Award Winner NimbRo TeenSize
Marcell Missura, Cedrick Münstermann, Malte Mauelshagen, Michael Schreiber, Sven Behnke |
RoboCup | 5 |
| 2012 | NimbRo@Home: Winning Team of the RoboCup@Home Competition 2012
Jörg Stückler, Ishrat Badami, David Droeschel, Kathrin Gräve, Dirk Holz, Manus McElhone, Matthias Nieuwenhuisen, Michael Schreiber, Max Schwarz, Sven Behnke |
RoboCup | 10 |
| 2011 | Learning to interpret pointing gestures with a time-of-flight cameraabstractPointing gestures are a common and intuitive way to draw somebody's attention to a certain object. While humans can easily interpret robot gestures, the perception of human behavior using robot sensors is more difficult. David Droeschel, Jörg Stückler, Sven Behnke |
HRI | 3 |
| 2011 | Towards joint attention for a domestic service robot - person awareness and gesture recognition using Time-of-Flight camerasabstractJoint attention between a human user and a robot is essential for effective human-robot interaction. In this work, we propose an approach to person awareness and to the perception of showing and pointing gestures for a domestic service robot. In contrast to previous work, we do not require the person to be at a predefined position, but instead actively approach and orient towards the communication partner. For perceiving showing and pointing gestures and for estimating the pointing direction a Time-of-Flight camera is used. Estimated pointing directions and shown objects are matched to objects in the robot's environment. Both the perception of showing and pointing gestures as well as the accurary of estimated pointing directions have been evaluated in a set of different experiments. The results show that both gestures are adequatly perceived by the robot. Furthermore, our system achieves a higher accuracy in estimating the pointing direction than is reported in the literature for a stereo-based system. In addition, the overall system has been successfully tested in two international RoboCup@Home competitions and the 2010 ICRA Mobile Manipulation Challenge. David Droeschel, Jörg Stückler, Dirk Holz, Sven Behnke |
ICRA | 4 |
| 2011 | Efficient kinodynamic trajectory generation for wheeled robotsabstractPlanning dynamic motion is computationally demanding and thus can hardly be done in real-time onboard robots. In this paper, we present an analytic approximation to predict the dynamic state of wheeled robots with non-holonomic constraints, given a start state and a sequence of piecewise constant controls. Our approximations are accurate and fast to calculate. They can be used to replace numerical integrators in kinodynamic planning algorithms. The predictions are differentiable and allow us to utilize gradient descent methods to solve the inverse dynamics as well and generate trajectories connecting arbitrary points in state space. Marcell Missura, Sven Behnke |
ICRA | 2 |
| 2011 | Interest point detection in depth images through scale-space surface analysisabstractMany perception problems in robotics such as object recognition, scene understanding, and mapping are tackled using scale-invariant interest points extracted from intensity images. Since interest points describe only local portions of objects and scenes, they offer robustness to clutter, occlusions, and intra-class variation. In this paper, we present an efficient approximate algorithm to extract surface normal interest points (SNIPs) in corners and blob-like surface regions from depth images. The interest points are detected on characteristic scales that indicate their spatial extent. Our method is able to cope with irregularly sampled, noisy measurements which are typical to depth imaging devices. It also offers a trade-off between computational speed and accuracy which allows our approach to be applicable in a wide range of problem sets. We evaluate our approach on depth images of basic geometric shapes, more complex objects, and indoor scenes. Jörg Stückler, Sven Behnke |
ICRA | 2 |
| 2011 | Real-Time Plane Segmentation Using RGB-D Cameras
Dirk Holz, Stefan Holzer, Radu Bogdan Rusu, Sven Behnke |
RoboCup | 4 |
| 2011 | RoboCup 2011 Humanoid League Winners
Daniel D. Lee, Seung-Joon Yi, Stephen G. McGill, Sven Behnke, Marcell Missura, Hannes Schulz, Dennis W. Hong, Jeakweon Han, Michael A. Hopkins |
RoboCup | 5 |
| 2011 | Learning Visual Obstacle Detection Using Color Histogram Features
Saskia Metzler, Matthias Nieuwenhuisen, Sven Behnke |
RoboCup | 3 |
| 2011 | Local Multiresolution Path Planning in Soccer Games Based on Projected Intentions
Matthias Nieuwenhuisen, Ricarda Steffens, Sven Behnke |
RoboCup | 3 |
| 2011 | Real-Time Trajectory Generation by Offline Footstep Planning for a Humanoid Soccer Robot
Andreas Schmitz, Marcell Missura, Sven Behnke |
RoboCup | 3 |
| 2011 | Compliant Task-Space Control with Back-Drivable Servo Actuators
Jörg Stückler, Sven Behnke |
RoboCup | 2 |
| 2011 | Towards Robust Mobility, Flexible Object Manipulation, and Intuitive Multimodal Interaction for Domestic Service Robots
Jörg Stückler, David Droeschel, Kathrin Gräve, Dirk Holz, Jochen Kläß, Michael Schreiber, Ricarda Steffens, Sven Behnke |
RoboCup | 8 |
| 2011 | Exploiting local structure in Boltzmann machines
Hannes Schulz, Andreas C. Müller 0001, Sven Behnke |
Neurocomputing | 3 |
| 2010 | Exploiting local structure in stacked Boltzmann machines
Hannes Schulz, Andreas C. Müller 0001, Sven Behnke |
ESANN | 3 |
| 2010 | Intuitive multimodal interaction for service robotsabstractDomestic service tasks require three main skills from autonomous robots: robust navigation in indoor environments, flexible object manipulation, and intuitive communication with the users. In this report, we present the communication skills of our anthropomorphic service and communication robots Dynamaid and Robotinho. Both robots are equipped with an intuitive multimodal communication system, including speech synthesis and recognition, gestures and mimic. We evaluate our systems in the @Home league of the RoboCup competitions and in a museum tour guide scenario. Matthias Nieuwenhuisen, Jörg Stückler, Sven Behnke |
HRI | 3 |
| 2010 | Evaluation of Pooling Operations in Convolutional Architectures for Object Recognition
Dominik Scherer, Andreas C. Müller 0001, Sven Behnke |
ICANN (3) | 3 |
| 2010 | Accelerating Large-Scale Convolutional Neural Networks with Parallel Graphics Multiprocessors
Dominik Scherer, Hannes Schulz, Sven Behnke |
ICANN (3) | 3 |
| 2010 | Using Time-of-Flight cameras with active gaze control for 3D collision avoidanceabstractWe propose a 3D obstacle avoidance method for mobile robots. Besides the robot's 2D laser range finder, a Time-of-Flight camera is used to perceive obstacles that are not in the scan plane of the laser range finder. Existing approaches that employ Time-of-Flight cameras suffer from the limited field-of-view of the sensor. To overcome this issue, we mount the camera on the head of our anthropomorphic robot Dynamaid. This allows to change the gaze direction through the robot's pan-tilt neck and its torso yaw joint. The proposed obstacle detection method is robust against kinematic inaccuracies and noise in the range measurements. The gaze controller takes motion blur effects into account and controls the gaze depending on the robot's motion and the obstacles in its vicinity. In experiments, we demonstrate that our approach enables the robot to avoid obstacles that the laser range finder can not perceive. We also compare our active gaze control strategy with a fixed gaze orientation. David Droeschel, Dirk Holz, Jörg Stückler, Sven Behnke |
ICRA | 4 |
| 2010 | Sancta simplicitas - on the efficiency and achievable results of SLAM using ICP-based incremental registrationabstractThis paper presents an efficient combination of algorithms for SLAM in dynamic environments. The overall approach is based on range image registration using the ICP algorithm. Different extensions to this algorithm are used to incrementally construct point models of the robot's workspace. A simple heuristic allows for determining which points in a newly acquired range image are already contained in the point model and for adding only those points that provide new information. Furthermore, the means for dealing with environment dynamics are presented which allow for continuously conducting SLAM and updating the point model according to changes in a dynamic environment. The achievable results of the overall approach are compared to Rao-Blackwellized Particle Filters as a state-of-the-art solution to the SLAM problem and evaluated using a recently published benchmark by Burgard et al. (2009). Dirk Holz, Sven Behnke |
ICRA | 2 |
| 2010 | Improving indoor navigation of autonomous robots by an explicit representation of doorsabstractIn the last decades, tremendous progress has been made in the field of autonomous indoor navigation for mobile robots. However, these approaches assume the structural part of the environment to be completely static. In practice, movable parts of scenes, e.g. doors, frequently violate this assumption which leads to poor performance. Also, mobile manipulation capabilities can only be utilized, if the robot knows about the movability of objects. In this paper, we address an important part of these problems by the explicit representation of doors as door leaves and joints. We propose to augment standard approaches to navigation like 2D occupancy grid mapping and Monte-Carlo-Localization. Our algorithm detects doors during mapping and represents their movability adequately in the map. During localization, the state of doors is estimated from measurements while it is simultaneously used to improve localization robustness and accuracy. In experimental results we demonstrate superior performance of our method compared to a state-of-the-art approach to localization. Matthias Nieuwenhuisen, Jörg Stückler, Sven Behnke |
ICRA | 3 |
| 2010 | Topological features in locally connected RBMsabstractUnsupervised learning algorithms find ways to model latent structure present in the data. These latent structures can then serve as a basis for supervised classification methods. A common choice for unsupervised feature discovery is the Restricted Boltzmann Machine (RBM). Since the RBM is a general purpose learning machine, it is not particularly tailored for image data. Representations found by RBMs are consequently not image-like. Since it is essential to exploit the known topological structure for image analysis, it is desirable not to discard the topology property when learning new representations. Then, the same learning methods can be applied to the latent representation in a hierarchical manner. In this work, we propose a modification to the learning rule of locally connected RBMs, which ensures that topological image structure is preserved in the latent representation. To this end, we use a Gaussian kernel to transfer topological properties of the image space to the feature space. The learned model is then used as an initialization for a neural network trained to classify the images. We evaluate our approach on the MNIST and Caltech 101 datasets and demonstrate that we are able to learn topological feature maps. Andreas C. Müller 0001, Hannes Schulz, Sven Behnke |
IJCNN | 3 |
| 2010 | Multi-frequency Phase Unwrapping for Time-of-Flight camerasabstractTime-of-Flight (ToF) cameras gain depth information by emitting amplitude-modulated near-infrared light and measuring the phase shift between the emitted and the reflected signal. The phase shift is proportional to the object's distance modulo the wavelength of the modulation frequency. This results in a distance ambiguity. Distances larger than the wavelength are wrapped into the sensor's non-ambiguity range and cause spurious distance measurements. We apply Phase Unwrapping to reconstruct these wrapped measurements. Our approach is based on a probabilistic graphical model. We use loopy belief propagation to detect and infer the position of wrapped measurements. Besides depth discontinuities, our method utilizes multiple modulation frequencies to identify wrapped measurements. In experiments, we show that wrapped measurements are identified and corrected, even in situations where the scene shows steep slopes in the depth measurements. David Droeschel, Dirk Holz, Sven Behnke |
IROS | 3 |
| 2010 | Combining depth and color cues for scale- and viewpoint-invariant object segmentation and recognition using Random ForestsabstractIn this paper we present an approach to object segmentation and recognition that combines depth and color cues. We fuse information from color images with depth from a Time-of-Flight (ToF) camera to improve recognition performance under scale and viewpoint changes. Firstly, we use depth and local surface orientation extracted from the ToF image to normalize color and depth image features with regard to scale and viewpoint. Secondly, we incorporate local 3D shape features into the classifier. The use of a Random Forest classifier facilitates the seamless combination of depth and texture features. It also provides image segmentation through pixel-wise classification. We demonstrate our approach on a labeled dataset of seven object categories in table-top scenes and compare it with a vision-only approach. Jörg Stückler, Sven Behnke |
IROS | 2 |
| 2010 | Towards Semantic Scene Analysis with Time-of-Flight Cameras
Dirk Holz, Ruwen Schnabel, David Droeschel, Jörg Stückler, Sven Behnke |
RoboCup | 5 |
| 2010 | Designing Effective Humanoid Soccer Goalies
Marcell Missura, Tobias Wilken, Sven Behnke |
RoboCup | 3 |
| 2010 | Learning Footstep Prediction from Motion Capture
Andreas Schmitz, Marcell Missura, Sven Behnke |
RoboCup | 3 |
| 2010 | Utilizing the Structure of Field Lines for Efficient Soccer Robot Localization
Hannes Schulz, Weichao Liu, Jörg Stückler, Sven Behnke |
RoboCup | 4 |
| 2010 | Improving People Awareness of Service Robots by Semantic Scene Knowledge
Jörg Stückler, Sven Behnke |
RoboCup | 2 |
| 2009 | Utilizing reflection properties of surfaces to improve mobile robot localizationabstractA main difficulty that arises in the context of probabilistic localization is the design of an appropriate observation model, i.e., determining the likelihood of a sensor measurement given the pose of the robot and a map of the environment. Many successful approaches to localization rely on data provided by range sensors, e.g., laser range scanners. When using such data one normally has to deal with erroneous maximum-range readings that occur due to poor-reflecting surfaces. In general, these readings cannot be distinguished from readings obtained when no obstacle is within the measurement range of the sensor. Therefore, existing localization techniques treat these readings alike in the observation model. In this paper, we present a novel approach that explicitly considers the reflection properties of surfaces and thus the expectation of valid range measurements. In addition to the expected range measurement, we compute the probability of reflectance for a beam given the relative pose of the robot to the obstacle taking into account the angle of incidence of the beam. We estimate the reflection properties of surfaces using data collected with a mobile robot equipped with a laser range scanner. As we demonstrate in experiments carried out with a real robot, our technique leads to significantly improved localization results compared to a state-of-the-art observation model. Maren Bennewitz, Cyrill Stachniss, Sven Behnke, Wolfram Burgard |
ICRA | 3 |
| 2009 | The humanoid museum tour guide RobotinhoabstractWheeled tour guide robots have already been deployed in various museums or fairs worldwide. A key requirement for successful tour guide robots is to interact with people and to entertain them. Most of the previous tour guide robots, however, focused more on the involved navigation task than on natural interaction with humans. Humanoid robots, on the other hand, offer a great potential for investigating intuitive, multimodal interaction between humans and machines. In this paper, we present our mobile full-body humanoid tour guide robot Robotinho. We provide mechanical and electrical details and cover perception, the integration of multiple modalities for interaction, navigation control, and system integration aspects. The multimodal interaction capabilities of Robotinho have been designed and enhanced according to the questionnaires filled out by the people who interacted with the robot at previous public demonstrations. We present experiences we have made during experiments in which untrained users interacted with the robot. Felix Faber, Maren Bennewitz, Clemens Eppner, Attila Görög, Christoph Gonsior, Dominik Joho, Michael Schreiber, Sven Behnke |
RO-MAN | 8 |
| 2008 | How to learn accurate grid maps with a humanoidabstractHumanoids have recently become a popular research platform in the robotics community. Such robots offer various fields for new applications. However, they have several drawbacks compared to wheeled vehicles such as stability problems, limited payload capabilities, violation of the flat world assumption, and they typically provide only very rough odometry information, if at all. In this paper, we investigate the problem of learning accurate grid maps with humanoid robots. We present techniques to deal with some of the above-mentioned difficulties. We describe how an existing approach to the simultaneous localization and mapping (SLAM) problem can be adapted to robustly learn accurate maps with a humanoid equipped with a laser range finder. We present an experiment in which our mapping system builds a highly accurate map with a size of around 20 m by 20 m using data acquired with a humanoid in our office environment containing two loops. The resulting maps have a similar accuracy as maps built with a wheeled robot. Cyrill Stachniss, Maren Bennewitz, Giorgio Grisetti, Sven Behnke, Wolfram Burgard |
ICRA | 4 |
| 2008 | Orthogonal wall correction for visual motion estimationabstractA good motion model is a prerequisite for many approaches to simultaneous localization and mapping. Without an absolute reference, it is however difficult to prevent drift when estimating motion. To prevent orientation drift, our approach exploits typical features of indoor environments: Straight walls that are parallel or orthogonal to each other. Our idea is to detect walls in monocular depth measurements and to correct odometry obtained from matching successive images and from inertial measurements, such that the observed walls are aligned with the main orientation estimated from the map that is being built. The experimental results indicate that orientation drift can be prevented and orientation uncertainty can be reduced greatly when applying the proposed orthogonal wall correction. This can make the difference between reliable mapping and failure. Jörg Stückler, Sven Behnke |
ICRA | 2 |
| 2008 | Controlling the gaze direction of a humanoid robot with redundant jointsabstractDue to their high number of joints, humanoid robots typically have kinematic redundancies to achieve end- effector poses. Examples for such redundancies are the kinematic chains of pitch and yaw joints that allow the robot to turn towards a gaze target. Our humanoid communication robot currently uses its spine, its neck, and its eye joints to direct its cameras towards an object. In this paper, we propose a control strategy that considers three factors, namely tracking error, discomfort, defined at the joint level, and "effort" to control the pitch and yaw joints. Our strategy is based on gradient descent on a cost function. During the optimization, we use different step sizes to reflect the different inertia of the moved parts. Our control scheme produces human-like motions, where smaller, light-weight parts such as the eyes of the robot move quickly towards the target and then move back while the larger joints turn towards the target. We present experiments to evaluate the proposed strategy qualitatively and quantitatively. Felix Faber, Maren Bennewitz, Sven Behnke |
RO-MAN | 3 |
| 2007 | Pitch Estimation using Models of Voiced Speech on Three LevelsabstractWe present an algorithm for estimating the fundamental frequency in speech signals. Our approach incorporates models of voiced speech on three levels. First, we estimate the pitch for each time frame based on its harmonic structure using non-negative matrix factorization. The second level utilizes temporal pitch continuity to extract partial pitch contours. Thirdly, we incorporate statistics of the succession of voiced segments to aggregate partial contours to the final contour of an utterance. We evaluate our approach on the Keele database. The experimental results show the robustness of our method for noisy speech, and the good performance for clean speech in comparison with state-of-the-art algorithms. Dominik Joho, Maren Bennewitz, Sven Behnke |
ICASSP (4) | 3 |
| 2007 | Fundamental Frequency Estimation Based on Pitch-Scaled Harmonic FilteringabstractIn this paper, we present an algorithm for robustly estimating the fundamental frequency in speech signals. Our approach is based on pitch-scaled harmonic filtering (PSHF). Following PSHF, we perform a filtering in the frequency domain using the short-time Fourier transform in order to separate the harmonic and non-harmonic parts of the processed signal. We enhance the standard PSHF approach by using a range of window lengths and a cost function that is applied to each window size. This cost function takes into account the energy at the harmonic and non-harmonic frequency coefficients to estimate harmonic energy for a frame. By using energy peaks and applying a cost function that considers the change in pitch in subsequent frames, we then determine the final pitch contour. We evaluated our approach on the Keele database. As the experimental results demonstrate, our methods performs robustly for noisy speech and has a good performance for clean speech in comparison with state-of-the-art algorithms. Sergio Roa, Maren Bennewitz, Sven Behnke |
ICASSP (4) | 3 |
| 2007 | Fritz - A Humanoid Communication RobotabstractIn this paper, we present the humanoid communication robot Fritz. Our robot communicates with people in an intuitive, multimodal way. Fritz uses speech, facial expressions, eye-gaze, and gestures to interact with people. Depending on the audio-visual input, our robot shifts its attention between different persons in order to involve them into the conversation. He performs human-like arm gestures during the conversation and also uses pointing gestures generated with eyes, head, and arms to direct the attention of its communication partners towards objects of interest. To express its emotional state, the robot generates facial expressions and adapts the speech synthesis. We discuss experiences made during two public demonstrations of our robot. Maren Bennewitz, Felix Faber, Dominik Joho, Sven Behnke |
RO-MAN | 4 |
| 2006 | Online Trajectory Generation for Omnidirectional Biped WalkingabstractThis paper describes the online generation of trajectories for omnidirectional walking on two legs. The gait can be parameterized using walking direction, walking speed, and rotational speed. Our approach has a low computational complexity and can be implemented on small onboard computers. We tested the proposed approach using our humanoid robot Jupp. The competitions in the RoboCup soccer domain showed that omnidirectional walking has advantages when acting in dynamic environments Sven Behnke |
ICRA | 1 |
| 2006 | Instability Detection and Fall Avoidance for a Humanoid using Attitude Sensors and ReflexesabstractHumanoid robots are inherently unstable because their center of mass is high, compared to the support polygon's size. Bipedal walking currently works well only under controlled conditions with limited external disturbances. In less controlled dynamic environments, such as RoboCup soccer fields, external disturbances might be large. While some disturbances might be too large to prevent a fall, some disturbances can be dealt with by specific rescue behaviors. This paper proposes a method to detect instabilities that occur during omnidirectional walking. We model the readings of attitude sensors using sinusoids. The model takes the gait target vector into account. We estimate model parameters from a gait test sequence and detect deviations of the actual sensor readings from the model later on. These deviations are aggregated to an instability indicator that triggers one of two reflexes, based on indicator strength. For small instabilities the robot is slowing down, but continues walking. For stronger instabilities the robot stops and is brought into a stable posture with a low center of mass. Walking continues as soon as the instability disappears. We extensively evaluated our approach in simulation by disturbing the robot with a variety of impulses. The results indicate that our method is very effective. For smaller disturbances, the probability of a fall could be reduced to zero. Most of the medium-sized disturbances could also be rejected. For the evaluation with the real robot, we used a walking against a wall with different speeds and at various angles. Here the results show a similar outcome to the ones in the simulations Reimund Renner, Sven Behnke |
IROS | 2 |
| 2006 | Imitative Reinforcement Learning for Soccer Playing Robots
Tobias Latzke, Sven Behnke, Maren Bennewitz |
RoboCup | 2 |
| 2006 | Multi-cue Localization for Soccer Playing Humanoid Robots
Hauke Strasdat, Maren Bennewitz, Sven Behnke |
RoboCup | 3 |
| 2005 | Integrating vision and speech for conversations with multiple personsabstractAn essential capability for a robot designed to interact with humans is to show attention to the people in its surroundings. To enable a robot to involve multiple persons into interaction requires the maintenance of an accurate belief about the people in the environment. In this paper, we use a probabilistic technique to update the knowledge of the robot based on sensory input. In this way, the robot is able to reason about the uncertainty in its belief about people in the vicinity and is able to shift its attention between different persons. Even people who are not the primary conversational partners are included into the interaction. In practical experiments with a humanoid robot, we demonstrate the effectiveness of our approach. Maren Bennewitz, Felix Faber, Dominik Joho, Michael Schreiber, Sven Behnke |
IROS | 5 |
| 2005 | Playing Soccer with RoboSapien
Sven Behnke, Jürgen Müller 0001, Michael Schreiber |
RoboCup | 1 |
| 2005 | Toni: A Soccer Playing Humanoid Robot
Sven Behnke, Jürgen Müller 0001, Michael Schreiber |
RoboCup | 1 |
| 2005 | Face localization and tracking in the neural abstraction pyramid
Sven Behnke |
Neural Comput. Appl. | 1 |
| 2003 | Meter value recognition using locally connected hierarchical networks
Sven Behnke |
ESANN | 1 |
| 2003 | A two-stage system for meter value recognitionabstractThis paper describes a two-stage system for the recognition of postage meter values. A feed-forward neural abstraction pyramid is initialized in an unsupervised manner and trained in a supervised fashion to classify an entire digit block. It does not need prior digit segmentation. If the block recognition is not confident enough, a second stage tries to recognize single digits, taking into account the block classifier output for a neighboring digit as context. The system is evaluated on a large database. It can recognize meter values that are hard to read for humans. Sven Behnke |
ICIP (1) | 1 |
| 2003 | Discovering hierarchical speech features using convolutional non-negative matrix factorizationabstractDiscovering a representation that reflects the structure of a dataset is a first step for many inference and learning methods. This paper aims at finding a hierarchy of localized speech features that can be interpreted as parts. Non-negative matrix factorization (NMF) has been proposed recently for the discovery of parts-based localized additive representations. The author proposes a variant of this method, convolutional NMF, that enforces a particular local connectivity with shared weights. Analysis starts from a spectrogram. The hidden representations produced by convolutional NMF are input to the same analysis method at the next higher level. Repeated application of convolutional NMF yields a sequence of increasingly abstract representations. These speech representations are parts-based, where complex higher-level parts are defined in terms of less complex lower-level ones. Sven Behnke |
IJCNN | 1 |
| 2003 | Face Localization in the Neural Abstraction Pyramid
Sven Behnke |
KES | 1 |
| 2003 | Local Multiresolution Path Planning
Sven Behnke |
RoboCup | 1 |
| 2003 | Predicting Away Robot Control Latency
Sven Behnke, Anna Förster, Alexander Gloye, Raúl Rojas 0001, Mark Simon |
RoboCup | 1 |
| 2002 | Learning Face Localization Using Hierarchical Recurrent Networks
Sven Behnke |
ICANN | 1 |
| 2001 | Learning Iterative Image Reconstruction
Sven Behnke |
IJCAI | 1 |
| 2001 | An Omnidirectional Vision System That Finds and Tracks Color Edges and Blobs
Felix von Hundelshausen, Sven Behnke, Raúl Rojas 0001 |
RoboCup | 2 |
| 2001 | FU-Fighters 2001 (Global Vision)
Raúl Rojas 0001, Sven Behnke, Achim Liers, Lars Knipping |
RoboCup | 2 |
| 2001 | FU-Fighters Omni 2001 (Local Vision)
Raúl Rojas 0001, Felix von Hundelshausen, Sven Behnke, Bernhard Frötschl |
RoboCup | 3 |
| 2001 | Learning Iterative Image Reconstruction in the Neural Abstraction PyramidabstractSuccessful image reconstruction requires the recognition of a scene and the generation of a clean image of that scene. We propose to use recurrent neural networks for both analysis and synthesis. The networks have a hierarchical architecture that represents images in multiple scales with different degrees of abstraction. The mapping between these representations is mediated by a local connection structure. We supply the networks with degraded images and train them to reconstruct the originals iteratively. This iterative reconstruction makes it possible to use partial results as context information to resolve ambiguities. We demonstrate the power of the approach using three examples: superresolution, fill-in of occluded parts, and noise removal/contrast enhancement. We also reconstruct images from sequences of degraded images. Sven Behnke |
Int. J. Comput. Intell. Appl. | 1 |
| 2000 | FU-Fighters 2000
Raúl Rojas 0001, Sven Behnke, Lars Knipping, Bernhard Frötschl |
RoboCup | 2 |
| 2000 | Robust Real Time Color Tracking
Mark Simon, Sven Behnke, Raúl Rojas 0001 |
RoboCup | 2 |
| 2000 | Recognition of Handwritten ZIP Codes in a Real-World Non-Standard-Letter Sorting System
Marcus Pfister, Sven Behnke, Raúl Rojas 0001 |
Appl. Intell. | 2 |
| 1999 | Hebbian learning and competition in the neural abstraction pyramidabstractThe neural abstraction pyramid is a hierarchical neural architecture for image interpretation that is inspired by the principles of information processing found in the visual cortex. In this paper we present an unsupervised learning algorithm for its connectivity based on Hebbian weight updates and competition. The algorithm yields a sequence of feature detectors that produce increasingly abstract representations of the image content. These representations are distributed and sparse, and facilitate the interpretation of the image. We apply the algorithm to a dataset of handwritten digits, starting from local contrast detectors. The emerging feature detectors correspond to step edges, lines, strokes, curves, and digit shapes. They can be used to reliably classify the digits. Sven Behnke |
IJCNN | 1 |
| 1999 | Using Hierarchical Dynamical Systems to Control Reactive Behavior
Sven Behnke, Bernhard Frötschl, Raúl Rojas 0001, Peter Ackers, Wolf Lindstrot, Manuel de Melo, Andreas Schebesch, Mark Simon, Martin Sprengel, Oliver Tenchio |
RoboCup | 1 |
| 1999 | FU-Fighters Team Description
Sven Behnke, Bernhard Frötschl, Raúl Rojas 0001, Peter Ackers, Wolf Lindstrot, Manuel de Melo, Andreas Schebesch, Mark Simon, Martin Sprengel, Oliver Tenchio |
RoboCup | 1 |
| 1998 | Competitive neural trees for pattern classificationabstractThis paper presents competitive neural trees (CNeT's) for pattern classification. The CNeT contains m-ary nodes and grows during learning by using inheritance to initialize new nodes. At the node level, the CNeT employs unsupervised competitive learning. The CNeT performs hierarchical clustering of the feature vectors presented to it as examples, while its growth is controlled by forward pruning. Because of the tree structure, the prototype in the CNeT close to any example can be determined by searching only a fraction of the tree. This paper introduces different search methods for the CNeT, which are utilized for training as well as for recall. The CNeT is evaluated and compared with existing classifiers on a variety of pattern classification problems. Sven Behnke, Nicolaos B. Karayiannis |
IEEE Trans. Neural Networks | 1 |