EDBT 2026 Demo / reviewers in the wild / expert
Jörg Stückler
dblp:99/3327 · also Joerg Stueckler
· DBLP profile ↗
65ranked-venue papers
16as first author
14since 2021 · last 2025
0000-0002-2328-4363ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 14 first-author · 12 since 2021Systems, architecture and hardware · 25 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Guiding Diffusion-Based Articulated Object Generation by Partial Point Cloud Alignment and Physical Plausibility Constraints
Jens U. Kreber, Jörg Stückler |
ICCV | 2 |
| 2025 | Incremental Few-Shot Adaptation for Non-Prehensile Object Manipulation Using Parallelizable Physics SimulatorsabstractFew-shot adaptation is an important capability for intelligent robots that perform tasks in open-world settings such as everyday environments or flexible production. In this paper, we propose a novel approach for non-prehensile manipulation which incrementally adapts a physics-based dynamics model for model-predictive control (MPC). The model prediction is aligned with a few examples of robot-object interactions collected with the MPC. This is achieved by using a parallelizable rigid-body physics simulation as dynamic world model and sampling-based optimization of the model parameters. In turn, the optimized dynamics model can be used for MPC using efficient sampling-based optimization. We evaluate our fewshot adaptation approach in object pushing experiments in simulation and with a real robot. Fabian Baumeister, Lukas Mack, Jörg Stückler |
ICRA | 3 |
| 2025 | Visuo-Tactile Object Pose Estimation for a Multi-Finger Robot Hand With Low-Resolution in-Hand Tactile SensingabstractAccurate 3D pose estimation of grasped objects is an important prerequisite for robots to perform assembly or in-hand manipulation tasks, but object occlusion by the robot's own hand greatly increases the difficulty of this perceptual task. Here, we propose that combining visual information and proprioception with binary, low-resolution tactile contact measurements from across the interior surface of an articulated robotic hand can mitigate this issue. The visuo-tactile object-pose-estimation problem is formulated probabilistically in a factor graph. The pose of the object is optimized to align with the three kinds of measurements using a robust cost function to reduce the influence of visual or tactile outlier readings. The advantages of the proposed approach are first demonstrated in simulation: a custom 15-DoF robot hand with one binary tactile sensor per link grasps 17 YCB objects while observed by an RGB-D camera. This low-resolution inhand tactile sensing significantly improves object-pose estimates under high occlusion and also high visual noise. We also show these benefits through grasping tests with a preliminary real version of our tactile hand, obtaining reasonable visuo-tactile estimates of object pose at approximately 13.3 Hz on average. Lukas Mack, Felix Grüninger, Benjamin A. Richardson, Regine Lendway, Katherine J. Kuchenbecker, Jörg Stückler |
ICRA | 6 |
| 2024 | Physics-Based Rigid Body Object Tracking and Friction Filtering From RGB-D VideosabstractPhysics-based understanding of object interactions from sensory observations is an essential capability in augmented reality and robotics. It enables to capture the properties of a scene for simulation and control. In this paper, we propose a novel approach for real-to-sim which tracks rigid objects in 3D from RGB-D images and infers physical properties of the objects. We use a differentiable physics simulation as state-transition model in an Extended Kalman Filter which can model contact and friction for arbitrary mesh-based shapes and in this way estimate physically plausible trajectories. We demonstrate that our approach can filter position, orientation, velocities, and concurrently can estimate the coefficient of friction of the objects. We analyze our approach on various sliding scenarios in synthetic image sequences of single objects and colliding objects. We also demonstrate and evaluate our approach on a real-world dataset. We make our novel benchmark datasets publicly available to foster future research in this novel problem setting and comparison with our method. Rama Krishna Kandukuri, Michael Strecke, Jörg Stückler |
3DV | 3 |
| 2024 | Online Calibration of a Single-Track Ground Vehicle Dynamics Model by Tight Fusion with Visual-Inertial OdometryabstractWheeled mobile robots need the ability to estimate their motion and the effect of their control actions for navigation planning. In this paper, we present ST-VIO, a novel approach which tightly fuses a single-track dynamics model for wheeled ground vehicles with visual-inertial odometry (VIO). Our method calibrates and adapts the dynamics model online to improve the accuracy of forward prediction conditioned on future control inputs. The single-track dynamics model approximates wheeled vehicle motion under specific control inputs on flat ground using ordinary differential equations. We use a singularity-free and differentiable variant of the single-track model to enable seamless integration as dynamics factor into VIO and to optimize the model parameters online together with the VIO state variables. We validate our method with real-world data in both indoor and outdoor environments with different terrain types and wheels. In experiments, we demonstrate that ST-VIO can not only adapt to wheel or ground changes and improve the accuracy of prediction under new control inputs, but can even improve tracking accuracy. Jörg Stückler |
ICRA | 2 |
| 2024 | Event-Based Non-rigid Reconstruction of Low-Rank Parametrized Deformations from ContoursabstractAbstract Visual reconstruction of fast non-rigid object deformations over time is a challenge for conventional frame-based cameras. In recent years, event cameras have gained significant attention due to their bio-inspired properties, such as high temporal resolution and high dynamic range. In this paper, we propose a novel approach for reconstructing such deformations using event measurements. Under the assumption of a static background, where all events are generated by the motion, our approach estimates the deformation of objects from events generated at the object contour in a probabilistic optimization framework. It associates events to mesh faces on the contour and maximizes the alignment of the line of sight through the event pixel with the associated face. In experiments on synthetic and real data of human body motion, we demonstrate the advantages of our method over state-of-the-art optimization and learning-based approaches for reconstructing the motion of human arms and hands. In addition, we propose an efficient event stream simulator to synthesize realistic event data for human motion. Yuxuan Xue 0001, Stefan Leutenegger, Jörg Stückler |
Int. J. Comput. Vis. | 4 |
| 2023 | Visual-Inertial and Leg Odometry Fusion for Dynamic LocomotionabstractImplementing dynamic locomotion behaviors on legged robots requires a high-quality state estimation module. Especially when the motion includes flight phases, state-of-the-art approaches fail to produce reliable estimation of the robot posture, in particular base height. In this paper, we propose a novel approach for combining visual-inertial odometry (VIO) with leg odometry in an extended Kalman filter (EKF) based state estimator. The VIO module uses a stereo camera and IMU to yield low-drift 3D position and yaw orientation and drift-free pitch and roll orientation of the robot base link in the inertial frame. However, these values have a considerable amount of latency due to image processing and optimization, while the rate of update is quite low which is not suitable for low-level control. To reduce the latency, we predict the VIO state estimate at the rate of the IMU measurements of the VIO sensor. The EKF module uses the base pose and linear velocity predicted by VIO, fuses them further with a second high-rate IMU and leg odometry measurements, and produces robot state estimates with a high frequency and small latency suitable for control. We integrate this lightweight estimation framework with a nonlinear model predictive controller and show successful implementation of a set of agile locomotion behaviors, including trotting and jumping at varying horizontal speeds, on a torque-controlled quadruped robot. Victor Dhédin, Shahram Khorshidi, Lukas Mack, Adithya Kumar Chinnakkonda Ravi, Avadesh Meduri, Paarth Shah, Felix Grimminger, Ludovic Righetti, Majid Khadiv, Jörg Stückler |
ICRA | 11 |
| 2023 | Learning-based Relational Object Matching Across ViewsabstractIntelligent robots require object-level scene understanding to reason about possible tasks and interactions with the environment. Moreover, many perception tasks such as scene reconstruction, image retrieval, or place recognition can benefit from reasoning on the level of objects. While keypoint-based matching can yield strong results for finding correspondences for images with small to medium view point changes, for large view point changes, matching semantically on the object-level becomes advantageous. In this paper, we propose a learning-based approach which combines local keypoints with novel object-level features for matching object detections between RGB images. We train our object-level matching features based on appearance and inter-frame and cross-frame spatial relations between objects in an associative graph neural network. We demonstrate our approach in a large variety of views on realistically rendered synthetic images. Our approach compares favorably to previous state-of-the-art object-level matching approaches and achieves improved performance over a pure keypoint-based approach for large view-point changes. Cathrin Elich, Iro Armeni, Martin R. Oswald, Marc Pollefeys, Jörg Stückler |
ICRA | 5 |
| 2022 | Event-based Non-Rigid Reconstruction from Contours
Yuxuan Xue 0001, Stefan Leutenegger, Jörg Stückler |
BMVC | 4 |
| 2022 | Weakly supervised learning of multi-object 3D scene decompositions using deep shape priorsabstractRepresenting scenes at the granularity of objects is a prerequisite for scene understanding and decision making. We propose PriSMONet, a novel approach based on Prior Shape knowledge for learning Multi-Object 3D scene decomposition and representations from single images. Our approach learns to decompose images of synthetic scenes with multiple objects on a planar surface into its constituent scene objects and to infer their 3D properties from a single view. A recurrent encoder regresses a latent representation of 3D shape, pose and texture of each object from an input RGB image. By differentiable rendering, we train our model to decompose scenes from RGB-D images in a self-supervised way. The 3D shapes are represented continuously in function-space as signed distance functions which we pre-train from example shapes in a supervised way. These shape priors provide weak supervision signals to better condition the challenging overall learning task. We evaluate the accuracy of our model in inferring 3D scene layout, demonstrate its generative capabilities, assess its generalization to real images, and point out benefits of the learned representation. Cathrin Elich, Martin R. Oswald, Marc Pollefeys, Jörg Stückler |
Comput. Vis. Image Underst. | 4 |
| 2022 | Physical Representation Learning and Parameter Identification from Video Using Differentiable PhysicsabstractAbstract Representation learning for video is increasingly gaining attention in the field of computer vision. For instance, video prediction models enable activity and scene forecasting or vision-based planning and control. In this article, we investigate the combination of differentiable physics and spatial transformers in a deep action conditional video representation network. By this combination our model learns a physically interpretable latent representation and can identify physical parameters. We propose supervised and self-supervised learning methods for our architecture. In experiments, we consider simulated scenarios with pushing, sliding and colliding objects, for which we also analyze the observability of the physical properties. We demonstrate that our network can learn to encode images and identify physical properties like mass and friction from videos and action sequences. We evaluate the accuracy of our training methods, and demonstrate the ability of our method to predict future video frames from input images and actions. Rama Krishna Kandukuri, Jan Achterhold, Michael Möller 0001, Jörg Stückler |
Int. J. Comput. Vis. | 4 |
| 2021 | DiffSDFSim: Differentiable Rigid-Body Dynamics With Implicit Shapes
Michael Strecke, Jörg Stückler |
3DV | 2 |
| 2021 | Explore the Context: Optimal Data Collection for Context-Conditional Dynamics ModelsabstractIn this paper, we learn dynamics models for parametrized families of dynamical systems with varying properties. The dynamics models are formulated as stochastic processes conditioned on a latent context variable which is inferred from observed transitions of the respective system. The probabilistic formulation allows us to compute an action sequence which, for a limited number of environment interactions, optimally explores the given system within the parametrized family. This is achieved by steering the system through transitions being most informative for the context variable. We demonstrate the effectiveness of our method for exploration on a non-linear toy-problem and two well-known reinforcement learning environments. Jan Achterhold, Jörg Stückler |
AISTATS | 2 |
| 2021 | Tracking 6-DoF Object Motion from Events and FramesabstractEvent cameras are promising devices for low latency tracking and high-dynamic range imaging. In this paper, we propose a novel approach for 6 degree-of-freedom (6-DoF) object motion tracking that combines measurements of event and frame-based cameras. We formulate tracking from high rate events with a probabilistic generative model of the event measurement process of the object. On a second layer, we refine the object trajectory in slower rate image frames through direct image alignment. We evaluate the accuracy of our approach in several object tracking scenarios with synthetic data, and also perform experiments with real data. Jörg Stückler |
ICRA | 2 |
| 2020 | Learning to Adapt Multi-View Stereo by Self-Supervision
Arijit Mallick, Jörg Stückler, Hendrik P. A. Lensch |
BMVC | 2 |
| 2020 | Where Does It End? - Reasoning About Hidden Surfaces by Object Intersection ConstraintsabstractDynamic scene understanding is an essential capability in robotics and VR/AR. In this paper we propose Co-Section, an optimization-based approach to 3D dynamic scene reconstruction, which infers hidden shape information from intersection constraints. An object-level dynamic SLAM frontend detects, segments, tracks and maps dynamic objects in the scene. Our optimization backend completes the shapes using hull and intersection constraints between the objects. In experiments, we demonstrate our approach on real and synthetic dynamic scene datasets. We also assess the shape completion performance of our method quantitatively. To the best of our knowledge, our approach is the first method to incorporate such physical plausibility constraints on object intersections for shape completion of dynamic objects in an energy minimization framework. Michael Strecke, Jörg Stückler |
CVPR | 2 |
| 2020 | DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape EstimationabstractScene understanding from images is a challenging problem encountered in autonomous driving. On the object level, while 2D methods have gradually evolved from computing simple bounding boxes to delivering finer grained results like instance segmentations, the 3D family is still dominated by estimating 3D bounding boxes. In this paper, we propose a novel approach to jointly infer the 3D rigid-body poses and shapes of vehicles from a stereo image pair using shape priors. Unlike previous works that geometrically align shapes to point clouds from dense stereo reconstruction, our approach works directly on images by combining a photometric and a silhouette alignment term in the energy function. An adaptive sparse point selection scheme is proposed to efficiently measure the consistency with both terms. In experiments, we show superior performance of our method on 3D pose and shape estimation over the previous geometric approach and demonstrate that our method can also be applied as a refinement step and significantly boost the performances of several state-of-the-art deep learning based 3D object detectors. All related materials and demonstration videos are available at the project page https://vision.in.tum.de/research/vslam/direct-shape. Rui Wang 0037, Nan Yang 0007, Jörg Stückler, Daniel Cremers |
ICRA | 3 |
| 2019 | EM-Fusion: Dynamic Object-Level SLAM With Probabilistic Data AssociationabstractThe majority of approaches for acquiring dense 3D environment maps with RGB-D cameras assumes static environments or rejects moving objects as outliers. The representation and tracking of moving objects, however, has significant potential for applications in robotics or augmented reality. In this paper, we propose a novel approach to dynamic SLAM with dense object-level representations. We represent rigid objects in local volumetric signed distance function (SDF) maps, and formulate multi-object tracking as direct alignment of RGB-D images with the SDF representations. Our main novelty is a probabilistic formulation which naturally leads to strategies for data association and occlusion handling. We analyze our approach in experiments and demonstrate that our approach compares favorably with the state-of-the-art methods in terms of robustness and accuracy. Michael Strecke, Jörg Stückler |
ICCV | 2 |
| 2018 | Direct Sparse Odometry with Rolling Shutter
David Schubert, Nikolaus Demmel, Vladyslav Usenko, Jörg Stückler, Daniel Cremers |
ECCV (8) | 4 |
| 2018 | Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry
Nan Yang 0007, Rui Wang 0037, Jörg Stückler, Daniel Cremers |
ECCV (8) | 3 |
| 2018 | The TUM VI Benchmark for Evaluating Visual-Inertial OdometryabstractVisual odometry and SLAM methods have a large variety of applications in domains such as augmented reality or robotics. Complementing vision sensors with inertial measurements tremendously improves tracking accuracy and robustness, and thus has spawned large interest in the development of visual-inertial (VI) odometry approaches. In this paper, we propose the TUM VI benchmark, a novel dataset with a diverse set of sequences in different scenes for evaluating VI odometry. It provides camera images with 1024×1024 resolution at 20 Hz, high dynamic range and photometric calibration. An IMU measures accelerations and angular velocities on 3 axes at 200 Hz, while the cameras and IMU sensors are time-synchronized in hardware. For trajectory evaluation, we also provide accurate pose ground truth from a motion capture system at high frequency (120 Hz) at the start and end of the sequences which we accurately aligned with the camera and IMU measurements. The full dataset with raw and calibrated data is publicly available. We also evaluate state-of-the-art VI odometry approaches on our dataset. David Schubert, Thore Goll, Nikolaus Demmel, Vladyslav Usenko, Jörg Stückler, Daniel Cremers |
IROS | 5 |
| 2017 | Semi-Supervised Deep Learning for Monocular Depth Map PredictionabstractSupervised deep learning often suffers from the lack of sufficient training data. Specifically in the context of monocular depth map prediction, it is barely possible to determine dense ground truth depth images in realistic dynamic outdoor environments. When using LiDAR sensors, for instance, noise is present in the distance measurements, the calibration between sensors cannot be perfect, and the measurements are typically much sparser than the camera images. In this paper, we propose a novel approach to depth map prediction from monocular images that learns in a semi-supervised way. While we use sparse ground-truth depth for supervised learning, we also enforce our deep network to produce photoconsistent dense depth maps in a stereo setup using a direct image alignment loss. In experiments we demonstrate superior performance in depth map prediction from single images compared to the state-of-the-art methods. Yevhen Kuznietsov, Jörg Stückler, Bastian Leibe |
CVPR | 2 |
| 2017 | Keyframe-based visual-inertial online SLAM with relocalizationabstractComplementing images with inertial measurements has become one of the most popular approaches to achieve highly accurate and robust real-time camera pose tracking. In this paper, we present a keyframe-based approach to visual-inertial simultaneous localization and mapping (SLAM) for monocular and stereo cameras. Our visual-inertial SLAM system is based on a real-time capable visual-inertial odometry method that provides locally consistent trajectory and map estimates. We achieve global consistency in the estimate through online loop-closing and non-linear optimization. Furthermore, our system supports relocalization in a map that has been previously obtained and allows for continued SLAM operation. We evaluate our approach in terms of accuracy, relocalization capability and run-time efficiency on public indoor benchmark datasets and on newly recorded outdoor sequences. We demonstrate state-of-the-art performance of our system compared to a visual-inertial odometry method and baseline visual SLAM approaches in recovering the trajectory of the camera. Anton Kasyanov, Francis Engelmann, Jörg Stückler, Bastian Leibe |
IROS | 3 |
| 2017 | Multi-view deep learning for consistent semantic mapping with RGB-D camerasabstractVisual scene understanding is an important capability that enables robots to purposefully act in their environment. In this paper, we propose a novel deep neural network approach to predict semantic segmentation from RGB-D sequences. The key innovation is to train our network to predict multi-view consistent semantics in a self-supervised way. At test time, its semantics predictions can be fused more consistently in semantic keyframe maps than predictions of a network trained on individual views. We base our network architecture on a recent single-view deep learning approach to RGB and depth fusion for semantic object-class segmentation and enhance it with multi-scale loss minimization. We obtain the camera trajectory using RGB-D SLAM and warp the predictions of RGB-D images into ground-truth annotated frames in order to enforce multi-view consistency during training. At test time, predictions from multiple views are fused into keyframes. We propose and analyze several methods for enforcing multi-view consistency during training and testing. We evaluate the benefit of multi-view consistency training and demonstrate that pooling of deep features and fusion over multiple views outperforms single-view baselines on the NYUDv2 benchmark for semantic segmentation. Our end-to-end trained network achieves state-of-the-art performance on the NYUDv2 dataset in single-view segmentation as well as multi-view semantic fusion. Lingni Ma, Jörg Stückler, Christian Kerl, Daniel Cremers |
IROS | 2 |
| 2017 | SAMP: Shape and Motion Priors for 4D Vehicle ReconstructionabstractInferring the pose and shape of vehicles in 3D from a movable platform still remains a challenging task due to the projective sensing principle of cameras, difficult surface properties e.g. reflections or transparency, and illumination changes between images. In this paper, we propose to use 3D shape and motion priors to regularize the estimation of the trajectory and the shape of vehicles in sequences of stereo images. We represent shapes by 3D signed distance functions and embed them in a low-dimensional manifold. Our optimization method allows for imposing a common shape across all image observations along an object track. We employ a motion model to regularize the trajectory to plausible object motions. We evaluate our method on the KITTI dataset and show state-of-the-art results in terms of shape reconstruction and pose estimation accuracy. Francis Engelmann, Jörg Stückler, Bastian Leibe |
WACV | 2 |
| 2016 | Unsupervised Learning of Shape-Motion Patterns for Objects in Urban Street Scenes
Dirk Klostermann, Aljosa Osep, Jörg Stückler, Bastian Leibe |
BMVC | 3 |
| 2016 | CPA-SLAM: Consistent plane-model alignment for direct RGB-D SLAMabstractPlanes are predominant features of man-made environments which have been exploited in many mapping approaches. In this paper, we propose a real-time capable RGB-D SLAM system that consistently integrates frame-to-keyframe and frame-to-plane alignment. Our method models the environment with a global plane model and - besides direct image alignment - it uses the planes for tracking and global graph optimization. This way, our method makes use of the dense image information available in keyframes for accurate short-term tracking. At the same time it uses a global model to reduce drift. Both components are integrated consistently in an expectation-maximization framework. In experiments, we demonstrate the benefits our approach and its state-of-the-art accuracy on challenging benchmarks. Lingni Ma, Christian Kerl, Jörg Stückler, Daniel Cremers |
ICRA | 3 |
| 2016 | Direct visual-inertial odometry with stereo camerasabstractWe propose a novel direct visual-inertial odometry method for stereo cameras. Camera pose, velocity and IMU biases are simultaneously estimated by minimizing a combined photometric and inertial energy functional. This allows us to exploit the complementary nature of vision and inertial data. At the same time, and in contrast to all existing visual-inertial methods, our approach is fully direct: geometry is estimated in the form of semi-dense depth maps instead of manually designed sparse keypoints. Depth information is obtained both from static stereo - relating the fixed-baseline images of the stereo camera - and temporal stereo - relating images from the same camera, taken at different points in time. We show that our method outperforms not only vision-only or loosely coupled approaches, but also can achieve more accurate results than state-of-the-art keypoint-based methods on different datasets, including rapid motion and significant illumination changes. In addition, our method provides high-fidelity semi-dense, metric reconstructions of the environment, and runs in real-time on a CPU. Vladyslav Usenko, Jakob J. Engel, Jörg Stückler, Daniel Cremers |
ICRA | 3 |
| 2016 | Scene flow propagation for semantic mapping and object discovery in dynamic street scenesabstractScene understanding is an important prerequisite for vehicles and robots that operate autonomously in dynamic urban street scenes. For navigation and high-level behavior planning, the robots not only require a persistent 3D model of the static surroundings-equally important, they need to perceive and keep track of dynamic objects. In this paper, we propose a method that incrementally fuses stereo frame observations into temporally consistent semantic 3D maps. In contrast to previous work, our approach uses scene flow to propagate dynamic objects within the map. Our method provides a persistent 3D occupancy as well as semantic belief on static as well as moving objects. This allows for advanced reasoning on objects despite noisy single-frame observations and occlusions. We develop a novel approach to discover object instances based on the temporally consistent shape, appearance, motion, and semantic cues in our maps. We evaluate our approaches to dynamic semantic mapping and object discovery on the popular KITTI benchmark and demonstrate improved results compared to single-frame methods. Deyvid Kochanov, Aljosa Osep, Jörg Stückler, Bastian Leibe |
IROS | 3 |
| 2015 | Motion Cooperation: Smooth Piece-wise Rigid Scene Flow from RGB-D ImagesabstractWe propose a novel joint registration and segmentation approach to estimate scene flow from RGB-D images. Instead of assuming the scene to be composed of a number of independent rigidly-moving parts, we use non-binary labels to capture non-rigid deformations at transitions between the rigid parts of the scene. Thus, the velocity of any point can be computed as a linear combination (interpolation) of the estimated rigid motions, which provides better results than traditional sharp piecewise segmentations. Within a variational framework, the smooth segments of the scene and their corresponding rigid velocities are alternately refined until convergence. A K-means-based segmentation is employed as an initialization, and the number of regions is subsequently adapted during the optimization process to capture any arbitrary number of independently moving objects. We evaluate our approach with both synthetic and real RGB-D images that contain varied and large motions. The experiments show that our method estimates the scene flow more accurately than the most recent works in the field, and at the same time provides a meaningful segmentation of the scene based on 3D motion. Mariano Jaimez, Mohamed Souiai, Jörg Stückler, Javier González 0001, Daniel Cremers |
3DV | 3 |
| 2015 | Super-resolution Keyframe Fusion for 3D Modeling with High-Quality TexturesabstractWe propose a novel fast and robust method for obtaining 3D models with high-quality appearance using commodity RGB-D sensors. Our method uses a direct key frame-based SLAM front end to consistently estimate the camera motion during the scan. The aligned images are fused into a volumetric truncated signed distance function representation, from which we extract a mesh. For obtaining a high-quality appearance model, we additionally deblur the low-resolution RGB-D frames using filtering techniques and fuse them into super-resolution key frames. The meshes are textured from these sharp super-resolution key frames employing a texture mapping approach. In experiments, we demonstrate that our method achieves superior quality in appearance compared to other state-of-the-art approaches. Robert Maier 0001, Jörg Stückler, Daniel Cremers |
3DV | 2 |
| 2015 | Reconstructing Street-Scenes in Real-Time from a Driving CarabstractMost current approaches to street-scene 3D reconstruction from a driving car to date rely on 3D laser scanning or tedious offline computation from visual images. In this paper, we compare a real-time capable 3D reconstruction method using a stereo extension of large-scale direct SLAM (LSD-SLAM) with laser-based maps and traditional stereo reconstructions based on processing individual stereo frames. In our reconstructions, small-baseline comparison over several subsequent frames are fused with fixed-baseline disparity from the stereo camera setup. These results demonstrate that our direct SLAM technique provides an excellent compromise between speed and accuracy, generating visually pleasing and globally consistent semi-dense reconstructions of the environment in real-time on a single CPU. Vladyslav Usenko, Jakob J. Engel, Jörg Stückler, Daniel Cremers |
3DV | 3 |
| 2015 | Dense Continuous-Time Tracking and Mapping with Rolling Shutter RGB-D CamerasabstractWe propose a dense continuous-time tracking and mapping method for RGB-D cameras. We parametrize the camera trajectory using continuous B-splines and optimize the trajectory through dense, direct image alignment. Our method also directly models rolling shutter in both RGB and depth images within the optimization, which improves tracking and reconstruction quality for low-cost CMOS sensors. Using a continuous trajectory representation has a number of advantages over a discrete-time representation (e.g. camera poses at the frame interval). With splines, less variables need to be optimized than with a discrete representation, since the trajectory can be represented with fewer control points than frames. Splines also naturally include smoothness constraints on derivatives of the trajectory estimate. Finally, the continuous trajectory representation allows to compensate for rolling shutter effects, since a pose estimate is available at any exposure time of an image. Our approach demonstrates superior quality in tracking and reconstruction compared to approaches with discrete-time or global shutter assumptions. Christian Kerl, Jörg Stückler, Daniel Cremers |
ICCV | 2 |
| 2015 | Large-scale direct SLAM with stereo camerasabstractWe propose a novel Large-Scale Direct SLAM algorithm for stereo cameras (Stereo LSD-SLAM) that runs in real-time at high frame rate on standard CPUs. In contrast to sparse interest-point based methods, our approach aligns images directly based on the photoconsistency of all high-contrast pixels, including corners, edges and high texture areas. It concurrently estimates the depth at these pixels from two types of stereo cues: Static stereo through the fixed-baseline stereo camera setup as well as temporal multi-view stereo exploiting the camera motion. By incorporating both disparity sources, our algorithm can even estimate depth of pixels that are under-constrained when only using fixed-baseline stereo. Using a fixed baseline, on the other hand, avoids scale-drift that typically occurs in pure monocular SLAM.We furthermore propose a robust approach to enforce illumination invariance, capable of handling aggressive brightness changes between frames - greatly improving the performance in realistic settings. In experiments, we demonstrate state-of-the-art results on stereo SLAM benchmarks such as Kitti or challenging datasets from the EuRoC Challenge 3 for micro aerial vehicles. Jakob J. Engel, Jörg Stückler, Daniel Cremers |
IROS | 2 |
| 2015 | Real-time object detection, localization and verification for fast robotic depalletizingabstractDepalletizing is a challenging task for manipulation robots. Key to successful application are not only robustness of the approach, but also achievable cycle times in order to keep up with the rest of the process. In this paper, we propose a system for depalletizing and a complete pipeline for detecting and localizing objects as well as verifying that the found object does not deviate from the known object model, e.g., if it is not the object to pick. In order to achieve high robustness (e.g., with respect to different lighting conditions) and generality with respect to the objects to pick, our approach is based on multi-resolution surfel models. All components (both software and hardware) allow operation at high frame rates and, thus, allow for low cycle times. In experiments, we demonstrate depalletizing of automotive and other prefabricated parts with both high reliability (w.r.t. success rates) and efficiency (w.r.t. low cycle times). Dirk Holz, Angeliki Topalidou-Kyniazopoulou, Jörg Stückler, Sven Behnke |
IROS | 3 |
| 2015 | Efficient Dense Rigid-Body Motion Segmentation and Estimation in RGB-D Video
Jörg Stückler, Sven Behnke |
Int. J. Comput. Vis. | 1 |
| 2015 | Corrigendum to "Multi-resolution surfel maps for efficient dense 3D modeling and tracking" [J. Visual Communication and Image Representation 25 (1) (2014) 137-147]
Jörg Stückler, Sven Behnke |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Mobile teleoperation interfaces with adjustable autonomy for personal service robotsabstractPersonal service robots require a comprehensive set of perception, control, and planning skills to perform everyday tasks autonomously. While achieving full autonomy is an ongoing research topic, first real-world applications of personal robots may come into reach, if state-of-the-art autonomous capabilities are combined with the intelligence of the users in a complementary way. We report on handheld user interfaces for personal robots that allow for teleoperating the robot on three levels of autonomy: body, skill, and task control. On the higher levels, autonomous behavior of the robot relieves the user from significant workload. If autonomous execution fails, or autonomous functionality is not provided by the robot system, the user can select a lower level of autonomy to solve a task. The benefits of providing adjustable autonomy in teleoperation have been successfully demonstrated at [email protected] competitions. Max Schwarz, Jörg Stückler, Sven Behnke |
HRI | 2 |
| 2014 | Local multi-resolution representation for 6D motion estimation and mapping with a continuously rotating 3D laser scannerabstractMicro aerial vehicles (MAV) pose a challenge in designing sensory systems and algorithms due to their size and weight constraints and limited computing power. We present an efficient 3D multi-resolution map that we use to aggregate measurements from a lightweight continuously rotating laser scanner. We estimate the robot's motion by means of visual odometry and scan registration, aligning consecutive 3D scans with an incrementally built map. By using local multi-resolution, we gain computational efficiency by having a high resolution in the near vicinity of the robot and a lower resolution with increasing distance from the robot, which correlates with the sensor's characteristics in relative distance accuracy and measurement density. Compared to uniform grids, local multi-resolution leads to the use of fewer grid cells without loosing information and consequently results in lower computational costs. We efficiently and accurately register new 3D scans with the map in order to estimate the motion of the MAV and update the map in-flight. In experiments, we demonstrate superior accuracy and efficiency of our registration approach compared to state-of-the-art methods such as GICP. Our approach builds an accurate 3D obstacle map and estimates the vehicle's trajectory in real-time. David Droeschel, Jörg Stückler, Sven Behnke |
ICRA | 2 |
| 2014 | Efficient deformable registration of multi-resolution surfel maps for object manipulation skill transferabstractEndowing mobile manipulation robots with skills to use objects and tools often involves the programming or training on specific object instances. To apply this knowledge to novel instances from the same class of objects, a robot requires generalization capabilities for control as well as perception. In this paper, we propose an efficient approach to deformable registration of RGB-D images that enables robots to transfer skills between object instances. Our method provides a dense deformation field between the current image and an object model which allows for estimating local rigid transformations on the object's surface. Since we define grasp and motion strategies as poses and trajectories with respect to the object models, these strategies can be transferred to novel instances through local transformations derived from the deformation field. In experiments, we demonstrate the accuracy and runtime efficiency of our registration method. We also report on the use of our skill transfer approach in a public demonstration. Jörg Stückler, Sven Behnke |
ICRA | 1 |
| 2014 | Cosero, Find My Keys! Object Localization and Retrieval Using Bluetooth Low Energy Tags
David Schwarz, Max Schwarz, Jörg Stückler, Sven Behnke |
RoboCup | 3 |
| 2014 | Multi-resolution surfel maps for efficient dense 3D modeling and tracking
Jörg Stückler, Sven Behnke |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | Efficient Dense 3D Rigid-Body Motion Segmentation in RGB-D VideoabstractMotion is a fundamental segmentation cue in video.Many current approaches segment 3D motion in monocular or stereo image sequences, mostly relying on sparse interest points or being dense but computationally demanding.We propose an efficient expectation-maximization (EM) framework for dense 3D segmentation of moving rigid parts in RGB-D video.Our approach segments two images into pixel regions that undergo coherent 3D rigid-body motion.Our formulation treats background and foreground objects equally and poses no further assumptions on the motion of the camera or the objects than rigidness.While our EM-formulation is not restricted to a specific image representation, we supplement it with efficient image representation and registration for rapid segmentation of RGB-D video.In experiments we demonstrate that our approach recovers segmentation and 3D motion at good precision. Jörg Stückler, Sven Behnke |
BMVC | 1 |
| 2013 | Combining contour and shape primitives for object detection and pose estimation of prefabricated partsabstractMan-made objects such as mechanical construction parts can typically be described as a composition of shape primitives like cylinders, planes, cones and spheres. We propose a robust method for the detection and pose estimation of such objects in 3D point clouds. Our main contribution is to enhance a probabilistic graph-matching approach that detects objects using 3D shape primitives with distinct 2D primitives such as circular contours. With this extension, our method copes with difficult occlusion situations and can be applied for object manipulation in complex scenarios such as grasping from a pile or bin-picking. We demonstrate the performance of our approach in a comparison with a state-of-the-art feature-based method for objects of generic shape and a primitive-based approach using only 3D shapes and no contours. Alexander Berner, Jun Li 0042, Dirk Holz, Jörg Stückler, Sven Behnke, Reinhard Klein |
ICIP | 4 |
| 2013 | Mobile bin picking with an anthropomorphic service robotabstractGrasping individual objects from an unordered pile in a box has been investigated in static scenarios so far. In this paper, we demonstrate bin picking with an anthropomorphic mobile robot. To this end, we extend global navigation techniques by precise local alignment with a transport box. Objects are detected in range images using a shape primitive-based approach. Our approach learns object models from single scans and employs active perception to cope with severe occlusions. Grasps and arm motions are planned in an efficient local multiresolution height map. All components are integrated and evaluated in a bin picking and part delivery task. Matthias Nieuwenhuisen, David Droeschel, Dirk Holz, Jörg Stückler, Alexander Berner, Jun Li 0042, Reinhard Klein, Sven Behnke |
ICRA | 4 |
| 2013 | Hierarchical Object Discovery and Dense Modelling From Motion Cues in RGB-D Video
Jörg Stückler, Sven Behnke |
IJCAI | 1 |
| 2013 | Increasing Flexibility of Mobile Manipulation and Intuitive Human-Robot Interaction in RoboCup@Home
Jörg Stückler, David Droeschel, Kathrin Gräve, Dirk Holz, Michael Schreiber, Angeliki Topalidou-Kyniazopoulou, Max Schwarz, Sven Behnke |
RoboCup | 1 |
| 2012 | Model Learning and Real-Time Tracking Using Multi-Resolution Surfel MapsabstractFor interaction with its environment, a robot is required to learn models of objects and to perceive these models in the livestreams from its sensors. In this paper, we propose a novel approach to model learning and real-time tracking. We extract multi-resolution 3D shape and texture representations from RGB-D images at high frame-rates. An efficient variant of the iterative closest points algorithm allows for registering maps in real-time on a CPU. Our approach learns full-view models of objects in a probabilistic optimization framework in which we find the best alignment between multiple views. Finally, we track the pose of the camera with respect to the learned model by registering the current sensor view to the model. We evaluate our approach on RGB-D benchmarks and demonstrate its accuracy, efficiency, and robustness in model learning and tracking. We also report on the successful public demonstration of our approach in a mobile manipulation task. Jörg Stückler, Sven Behnke |
AAAI | 1 |
| 2012 | Semantic mapping using object-class segmentation of RGB-D imagesabstractFor task planning and execution in unstructured environments, a robot needs the ability to recognize and localize relevant objects. When this information is made persistent in a semantic map, it can be used, e. g., to communicate with humans. In this paper, we propose a novel approach to learning such maps. Our approach registers measurements of RGB-D cameras by means of simultaneous localization and mapping. We employ random decision forests to segment object classes in images and exploit dense depth measurements to obtain scale-invariance. Our object recognition method integrates shape and texture seamlessly. The probabilistic segmentation from multiple views is filtered in a voxel-based 3D map using a Bayesian framework. We report on the quality of our object-class segmentation method and demonstrate the benefits in accuracy when fusing multiple views in a semantic map. Jörg Stückler, Nenad Biresev, Sven Behnke |
IROS | 1 |
| 2012 | Adjustable autonomy for mobile teleoperation of personal service robotsabstractAbstract — Controlling personal service robots is an impor-tant, complex task which must be accomplished by the users themselves. In this paper, we propose a novel user interface for personal service robots that allows the user to teleoperate the robot using handheld computers. The user can adjust the autonomy of the robot between three levels: body control, skill control, and task control. We design several user interfaces for teleoperation on these levels. On the higher levels, au-tonomous behavior of the robot relieves the user from significant workload. If the autonomous execution fails, or autonomous functionality is not provided by the robot system, the user can select a lower level of autonomy, e.g., direct body control, to solve a task. In a qualitative user study we evaluate usability aspects of our teleoperation interface with our domestic service robots Cosero and Dynamaid. We demonstrate the benefits of providing adjustable manual and autonomous control. I. Sebastian Muszynski, Jörg Stückler, Sven Behnke |
RO-MAN | 2 |
| 2012 | NimbRo@Home: Winning Team of the RoboCup@Home Competition 2012
Jörg Stückler, Ishrat Badami, David Droeschel, Kathrin Gräve, Dirk Holz, Manus McElhone, Matthias Nieuwenhuisen, Michael Schreiber, Max Schwarz, Sven Behnke |
RoboCup | 1 |
| 2011 | Learning to interpret pointing gestures with a time-of-flight cameraabstractPointing gestures are a common and intuitive way to draw somebody's attention to a certain object. While humans can easily interpret robot gestures, the perception of human behavior using robot sensors is more difficult. David Droeschel, Jörg Stückler, Sven Behnke |
HRI | 2 |
| 2011 | Towards joint attention for a domestic service robot - person awareness and gesture recognition using Time-of-Flight camerasabstractJoint attention between a human user and a robot is essential for effective human-robot interaction. In this work, we propose an approach to person awareness and to the perception of showing and pointing gestures for a domestic service robot. In contrast to previous work, we do not require the person to be at a predefined position, but instead actively approach and orient towards the communication partner. For perceiving showing and pointing gestures and for estimating the pointing direction a Time-of-Flight camera is used. Estimated pointing directions and shown objects are matched to objects in the robot's environment. Both the perception of showing and pointing gestures as well as the accurary of estimated pointing directions have been evaluated in a set of different experiments. The results show that both gestures are adequatly perceived by the robot. Furthermore, our system achieves a higher accuracy in estimating the pointing direction than is reported in the literature for a stereo-based system. In addition, the overall system has been successfully tested in two international RoboCup@Home competitions and the 2010 ICRA Mobile Manipulation Challenge. David Droeschel, Jörg Stückler, Dirk Holz, Sven Behnke |
ICRA | 2 |
| 2011 | Interest point detection in depth images through scale-space surface analysisabstractMany perception problems in robotics such as object recognition, scene understanding, and mapping are tackled using scale-invariant interest points extracted from intensity images. Since interest points describe only local portions of objects and scenes, they offer robustness to clutter, occlusions, and intra-class variation. In this paper, we present an efficient approximate algorithm to extract surface normal interest points (SNIPs) in corners and blob-like surface regions from depth images. The interest points are detected on characteristic scales that indicate their spatial extent. Our method is able to cope with irregularly sampled, noisy measurements which are typical to depth imaging devices. It also offers a trade-off between computational speed and accuracy which allows our approach to be applicable in a wide range of problem sets. We evaluate our approach on depth images of basic geometric shapes, more complex objects, and indoor scenes. Jörg Stückler, Sven Behnke |
ICRA | 1 |
| 2011 | Compliant Task-Space Control with Back-Drivable Servo Actuators
Jörg Stückler, Sven Behnke |
RoboCup | 1 |
| 2011 | Towards Robust Mobility, Flexible Object Manipulation, and Intuitive Multimodal Interaction for Domestic Service Robots
Jörg Stückler, David Droeschel, Kathrin Gräve, Dirk Holz, Jochen Kläß, Michael Schreiber, Ricarda Steffens, Sven Behnke |
RoboCup | 1 |
| 2010 | Intuitive multimodal interaction for service robotsabstractDomestic service tasks require three main skills from autonomous robots: robust navigation in indoor environments, flexible object manipulation, and intuitive communication with the users. In this report, we present the communication skills of our anthropomorphic service and communication robots Dynamaid and Robotinho. Both robots are equipped with an intuitive multimodal communication system, including speech synthesis and recognition, gestures and mimic. We evaluate our systems in the @Home league of the RoboCup competitions and in a museum tour guide scenario. Matthias Nieuwenhuisen, Jörg Stückler, Sven Behnke |
HRI | 2 |
| 2010 | Using Time-of-Flight cameras with active gaze control for 3D collision avoidanceabstractWe propose a 3D obstacle avoidance method for mobile robots. Besides the robot's 2D laser range finder, a Time-of-Flight camera is used to perceive obstacles that are not in the scan plane of the laser range finder. Existing approaches that employ Time-of-Flight cameras suffer from the limited field-of-view of the sensor. To overcome this issue, we mount the camera on the head of our anthropomorphic robot Dynamaid. This allows to change the gaze direction through the robot's pan-tilt neck and its torso yaw joint. The proposed obstacle detection method is robust against kinematic inaccuracies and noise in the range measurements. The gaze controller takes motion blur effects into account and controls the gaze depending on the robot's motion and the obstacles in its vicinity. In experiments, we demonstrate that our approach enables the robot to avoid obstacles that the laser range finder can not perceive. We also compare our active gaze control strategy with a fixed gaze orientation. David Droeschel, Dirk Holz, Jörg Stückler, Sven Behnke |
ICRA | 3 |
| 2010 | Improving indoor navigation of autonomous robots by an explicit representation of doorsabstractIn the last decades, tremendous progress has been made in the field of autonomous indoor navigation for mobile robots. However, these approaches assume the structural part of the environment to be completely static. In practice, movable parts of scenes, e.g. doors, frequently violate this assumption which leads to poor performance. Also, mobile manipulation capabilities can only be utilized, if the robot knows about the movability of objects. In this paper, we address an important part of these problems by the explicit representation of doors as door leaves and joints. We propose to augment standard approaches to navigation like 2D occupancy grid mapping and Monte-Carlo-Localization. Our algorithm detects doors during mapping and represents their movability adequately in the map. During localization, the state of doors is estimated from measurements while it is simultaneously used to improve localization robustness and accuracy. In experimental results we demonstrate superior performance of our method compared to a state-of-the-art approach to localization. Matthias Nieuwenhuisen, Jörg Stückler, Sven Behnke |
ICRA | 2 |
| 2010 | Combining depth and color cues for scale- and viewpoint-invariant object segmentation and recognition using Random ForestsabstractIn this paper we present an approach to object segmentation and recognition that combines depth and color cues. We fuse information from color images with depth from a Time-of-Flight (ToF) camera to improve recognition performance under scale and viewpoint changes. Firstly, we use depth and local surface orientation extracted from the ToF image to normalize color and depth image features with regard to scale and viewpoint. Secondly, we incorporate local 3D shape features into the classifier. The use of a Random Forest classifier facilitates the seamless combination of depth and texture features. It also provides image segmentation through pixel-wise classification. We demonstrate our approach on a labeled dataset of seven object categories in table-top scenes and compare it with a vision-only approach. Jörg Stückler, Sven Behnke |
IROS | 1 |
| 2010 | Towards Semantic Scene Analysis with Time-of-Flight Cameras
Dirk Holz, Ruwen Schnabel, David Droeschel, Jörg Stückler, Sven Behnke |
RoboCup | 4 |
| 2010 | Utilizing the Structure of Field Lines for Efficient Soccer Robot Localization
Hannes Schulz, Weichao Liu, Jörg Stückler, Sven Behnke |
RoboCup | 3 |
| 2010 | Improving People Awareness of Service Robots by Semantic Scene Knowledge
Jörg Stückler, Sven Behnke |
RoboCup | 1 |
| 2008 | Orthogonal wall correction for visual motion estimationabstractA good motion model is a prerequisite for many approaches to simultaneous localization and mapping. Without an absolute reference, it is however difficult to prevent drift when estimating motion. To prevent orientation drift, our approach exploits typical features of indoor environments: Straight walls that are parallel or orthogonal to each other. Our idea is to detect walls in monocular depth measurements and to correct odometry obtained from matching successive images and from inertial measurements, such that the observed walls are aligned with the main orientation estimated from the map that is being built. The experimental results indicate that orientation drift can be prevented and orientation uncertainty can be reduced greatly when applying the proposed orthogonal wall correction. This can make the difference between reliable mapping and failure. Jörg Stückler, Sven Behnke |
ICRA | 1 |
| 2005 | Successful Search and Rescue in Simulated Disaster Areas
Alexander Kleiner, Michael Brenner 0001, Tobias Bräuer, Christian Dornhege, Moritz Göbelbecker, Matthias Luber, Johann Prediger, Jörg Stückler, Bernhard Nebel |
RoboCup | 8 |