VLDB 2026 Research / reviewers in the wild / expert
Markus Vincze
dblp:23/6056
· DBLP profile ↗
158ranked-venue papers
12as first author
24since 2021 · last 2026
0000-0002-2799-491XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 110 · 9 first-author · 14 since 2021Systems, architecture and hardware · 63 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 56 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 17 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCOPE: Semantic Cross-Attention Conditioning for Category-Level Object Pose Estimation
Peter Hönig, Jean-Baptiste Weibel, Stefan Thalhammer, Matthias Hirschmanner, Markus Vincze, Andreas Holzinger |
Image Vis. Comput. | 5 |
| 2025 | ODYSSEE: Oyster Detection Yielded by Sensor Systems on Edge ElectronicsabstractOysters are a vital keystone species in coastal ecosystems, providing significant economic, environmental, and cultural benefits. As the importance of oysters grows, so does the relevance of autonomous systems for their detection and monitoring. However, current monitoring strategies often rely on destructive methods. While manual identification of oysters from video footage is non-destructive, it is time-consuming, requires expert input, and is further complicated by the challenges of the underwater environment. To address these challenges, we propose a novel pipeline using stable diffusion to augment a collected real dataset with photorealistic synthetic data. This method enhances the dataset used to train a YOLOv10-based vision model. The model is then deployed and tested on an edge platform; Aqua2, an Autonomous Underwater Vehicle (AUV), achieving a state-of-the-art 0.657 mAP@50 for oyster detection. Xiaomin Lin 0002, Vivek Mange, Arjun Suresh, Bernhard Neuberger, Aadi Palnitkar, Brendan Campbell, Kleio Baxevani, Jeremy Mallette, Alhim Vera, Markus Vincze, Ioannis M. Rekleitis, Herbert G. Tanner, Yiannis Aloimonos |
ICRA | 11 |
| 2025 | Shape-Biased Texture Agnostic Representations for Improved Textureless and Metallic Object Detection and 6D Pose EstimationabstractRecent advances in machine learning have greatly benefited object detection and 6D pose estimation. However, textureless and metallic objects still pose a significant challenge due to few visual cues and the texture bias of CNNs. To address this issue, we propose a strategy for inducing a shape bias to CNN training. In particular, by randomizing textures applied to object surfaces during data rendering, we create training data without consistent textural cues. This methodology allows for seamless integration into existing data rendering engines, and results in negligible computational overhead for data rendering and network training. Our findings demonstrate that the shape bias we induce via randomized texturing, improves over existing approaches using style transfer. We evaluate with five detectors and two pose estimators. For three object detectors and for pose estimation in general, estimation accuracy improves for textureless and metallic objects. Additionally we show that our approach increases the pose estimation accuracy in the presence of image noise and strong illumination changes. Code available at https://github.com/hoenigpeter/randomized_texturing. Peter Hönig, Stefan Thalhammer, Jean-Baptiste Weibel, Matthias Hirschmanner, Markus Vincze |
WACV | 5 |
| 2025 | CAGT: Sim-to-Real Depth Completion With Interactive Embedding Aggregation and Geometry Awareness for Transparent ObjectsabstractRobust depth completion of transparent objects would be beneficial for industrial automation such as vision-based robotic grasping and manipulation. However, although some methods try to learn a compact intra-layer feature representation with the boost of the attention mechanism or the vision Transformer, they ignore the neglected corner regions and sparse geometry information that are important for accurate depth completion. To tackle these issues, we propose a novel sim-to-real transferable model, named CAGT, with interactive embedding aggregation and geometry awareness to reconstruct severely sparse depth maps of transparent objects in this paper. We design a Depth-clue Interaction Aggregation Module (DIAM) to enhance the Transformer’s ability to extract boundary corner features and thus supplement depth clues. Then, we propose a Geometric Information Augmentation Module (GIAM) to fuse the geometry-aware feature containing shape and surface details. Moreover, we introduce a contrastive learning mechanism to facilitate the sim-to-real generalization of the completion model. Extensive experiment results on two challenging datasets, ClearGrasp and TransCG, demonstrate that our proposed CAGT can obtain superior performance over the state-of-the-art methods. We also demonstrate that CAGT can improve the grasp accuracy of transparent objects by a robotic grasping generalization experiment. The code and supplementary video will be available at:https://github.com/xingshuojing/CAGT. Xingshuo Jing, Kun Qian 0005, Markus Vincze |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | ZS6D: Zero-shot 6D Object Pose Estimation using Vision TransformersabstractAs robotic systems increasingly encounter complex and unconstrained real-world scenarios, there is a demand to recognize diverse objects. The state-of-the-art 6D object pose estimation methods rely on object-specific training and therefore do not generalize to unseen objects. Recent novel object pose estimation methods are solving this issue using task-specific fine-tuned CNNs for deep template matching. This adaptation for pose estimation still requires expensive data rendering and training procedures. MegaPose for example is trained on a dataset consisting of two million images showing 20,000 different objects to reach such generalization capabilities. To overcome this shortcoming we introduce ZS6D, for zero-shot novel object 6D pose estimation. Visual descriptors, extracted using pre-trained Vision Transformers (ViT), are used for matching rendered templates against query images of objects and for establishing local correspondences. These local correspondences enable deriving geometric correspondences and are used for estimating the object's 6D pose with RANSAC- based PnP. This approach showcases that the image descriptors extracted by pre-trained ViTs are well-suited to achieve a notable improvement over two state-of-the-art novel object 6D pose estimation methods, without the need for task-specific fine-tuning. Experiments are performed on LMO, YCBV, and TLESS. In comparison to MegaPose, we improve the Average Recall on all three datasets and compared to OSOP we improve on two datasets. The code is available at https://github.com/PhilippAuss/ZS6D. Philipp Ausserlechner, David Haberger, Stefan Thalhammer, Jean-Baptiste Weibel, Markus Vincze |
ICRA | 5 |
| 2024 | EdgeSoil 2.0 - Soil Analyzer Using Convolutional Neural Network and Camera Imaging for Agricultural RoboticsabstractSoil is the most important building element of agriculture and its analysis is crucial for healthy plants and a high crop yield. But apart from its importance, soil analysis is a tedious and time-consuming task. This paper presents EdgeSoil 2.0, a non-invasive, accurate, and real-time robotic system for soil pH prediction, a key parameter of soil status for farmers. The EdgeSoil 2.0 predicts the pH value of the soil in real-time, using a live video stream from a webcam with an average of 7 FPS. The method is suitable to be implemented on edge devices necessary for the application: we are using a mobile robot with the NVIDIA Jetson Nano module which is running a pH-estimator trained with a Convolutional Neural Network (CNN) on a novel dataset we built for this purpose. Predictions are performed while the robot is moving over the plowed field before the planting process starts. In order to achieve the best performance, we train the pH-estimator with different input modalities and validate each result using Mean Squared Error (MSE) and Standard Deviation (SD). We are able to achieve accurate results with the MSE value of 0.08, the SD value of 0.15, and with testing results from the field showing up to ± 0.3 deviation from the GT value during prediction, which is sufficient to comply with agricultural standards. Roni Kasemi, Lara Lammer, Stefan Thalhammer, Markus Vincze |
ICRA | 4 |
| 2024 | NRDF - Neural Region Descriptor Fields as Implicit ROI Representation for Robotic 3D Surface ProcessingabstractTo automate 3D surface processing across diverse category-level objects it is imperative to represent process-related region of interest (P-ROI), which is not obtained with conventional keypoint or semantic part correspondences. To resolve this issue, we propose Neural Region Descriptor Fields (NRDF) for achieving unsupervised dense 3D surface region correspondence such that arbitrary ROI is retrieved for a new instance of a known category of object. We utilize the NRDF representation as a medium to facilitate one-shot P-ROI level process knowledge transfer. Recent developments in implicit 3D object representations have focused on keypoint or part correspondences, which have resulted in applications like robotic grasping and manipulation. However, explicit one-shot P-ROI correspondence, and its application for 3D surface process knowledge transfer, is treated for the first time in this work, to the best of our knowledge. The evaluation results show that the proposed approach outperforms the dense correspondence baselines in implicit shape representation and the capacity to retrieve matching arbitrary ROIs. In addition, we validate the practicality of our proposed system in a real-world robotic surface processing application. Our code is available at https://github.com/Profactor/Neural-Region-Descriptor-Fields. Anish Pratheepkumar, Markus Ikeda, Michael Hofmann 0006, Fabian Widmoser, Andreas Pichler, Markus Vincze |
IROS | 6 |
| 2024 | Real-time 6-DoF Pose Estimation by an Event-based Camera using Active LED MarkersabstractReal-time applications for autonomous operations depend largely on fast and robust vision-based localization systems. Since image processing tasks require processing large amounts of data, the computational resources often limit the performance of other processes. To overcome this limitation, traditional marker-based localization systems are widely used since they are easy to integrate and achieve reliable accuracy. However, classical marker-based localization systems significantly depend on standard cameras with low frame rates, which often lack accuracy due to motion blur. In contrast, event-based cameras provide high temporal resolution and a high dynamic range, which can be utilized for fast localization tasks, even under challenging visual conditions. This paper proposes a simple but effective event-based pose estimation system using active LED markers (ALM) for fast and accurate pose estimation. The proposed algorithm is able to operate in real time with a latency below 0.5 ms while maintaining output rates of 3 kHz. Experimental results in static and dynamic scenarios are presented to demonstrate the performance of the proposed approach in terms of computational speed and absolute accuracy, using the OptiTrack system as the basis for measurement. Moreover, we demonstrate the feasibility of the proposed approach by deploying the hardware, i.e., the event-based camera and ALM, and the software in a real quadcopter application. Our project page is available at: almpose.github.io Gerald Ebmer, Adam Loch, Minh Nhat Vu, Roberto Mecca, Germain Haessig, Christian Hartl-Nesic, Markus Vincze, Andreas Kugi |
WACV | 7 |
| 2024 | Challenges for Monocular 6-D Object Pose Estimation in RoboticsabstractObject pose estimation is a core perception task that enables, for example, object manipulation and scene understanding. The widely available, inexpensive, and high-resolution RGB sensors and CNNs that allow for fast inference make monocular approaches especially well-suited for robotics applications. We observe that previous surveys establish the state of the art for varying modalities, single- and multiview settings, and datasets and metrics that consider a multitude of applications. We argue, however, that those works' broad scope hinders the identification of open challenges that are specific to monocular approaches and the derivation of promising future challenges for their application in robotics. By providing a unified view on recent publications from both robotics and computer vision, we find that occlusion handling, pose representations, and formalizing and improving category-level pose estimation are still fundamental challenges that are highly relevant for robotics. Moreover, to further improve robotic performance, large object sets, novel objects, refractive materials, and uncertainty estimates are central and largely unsolved open challenges. In order to address them, ontological reasoning, deformability handling, scene-level reasoning, realistic datasets, and the ecological footprint of algorithms need to be improved. Stefan Thalhammer, Dominik Bauer, Peter Hönig, Jean-Baptiste Weibel, José García Rodríguez 0001, Markus Vincze |
IEEE Trans. Robotics | 6 |
| 2023 | 3D-DAT: 3D-Dataset Annotation Toolkit for Robotic VisionabstractRobots operating in the real world are expected to detect, classify, segment, and estimate the pose of objects to accomplish their task. Modern approaches using deep learning not only require large volumes of data but also pixel-accurate annotations in order to evaluate the performance and therefore safety of these algorithms. At present, publicly available tools for annotating data are scarce and those that are available rely on depth sensors, which excludes their use for transparent, metallic, and general non-Lambertian objects. To address this issue, we present a novel method for creating valuable datasets that can be used in these more difficult cases. Our key contribution is a purely RGB-based scene-level annotation approach that uses a neural radiance field-based method to automatically align objects. A set of user studies demonstrates the accuracy and speed of our approach over a purely manual or depth sensor assisted pipeline. We provide an open-source implementation of each component and a ROS-based recorder for capturing data with a eye-in-hand robot system. Code will be made available at https://github.com/markus-suchi/3D-DAT. Markus Suchi, Bernhard Neuberger, Amanzhol Salykov, Jean-Baptiste Weibel, Tim Patten, Markus Vincze |
ICRA | 6 |
| 2023 | TrackAgent: 6D Object Tracking via Reinforcement Learning
Konstantin Röhrl, Dominik Bauer, Tim Patten, Markus Vincze |
ICVS | 4 |
| 2023 | COPE: End-to-end trainable Constant Runtime Object Pose EstimationabstractState-of-the-art object pose estimation handles multiple instances in a test image by using multi-model formulations: detection as a first stage and then separately trained networks per object for 2D-3D geometric correspondence prediction as a second stage. Poses are subsequently estimated using the Perspective-n-Points algorithm at runtime. Unfortunately, multi-model formulations are slow and do not scale well with the number of object instances involved. Recent approaches show that direct 6D object pose estimation is feasible when derived from the aforementioned geometric correspondences. We present an approach that learns an intermediate geometric representation of multiple objects to directly regress 6D poses of all instances in a test image. The inherent end-to-end trainability overcomes the requirement of separately processing individual object instances. By calculating the mutual Intersection-over-Unions, pose hypotheses are clustered into distinct instances, which achieves negligible runtime overhead with respect to the number of object instances. Results on multiple challenging standard datasets show that the pose estimation performance is superior to single-model state-of-the-art approaches despite being more than ~35 times faster. We additionally provide an analysis showing real-time applicability (> 24 fps) for images where more than 90 object instances are present. Further results show the advantage of supervising geometric correspondence-based object pose estimation with the 6D pose. Stefan Thalhammer, Tim Patten, Markus Vincze |
WACV | 3 |
| 2023 | Self-supervised Vision Transformers for 3D pose estimation of novel objectsabstractObject pose estimation is important for object manipulation and scene understanding. In order to improve the general applicability of pose estimators, recent research focuses on providing estimates for novel objects, that is, objects unseen during training. Such works use deep template matching strategies to retrieve the closest template connected to a query image, which implicitly provides object class and pose. Despite the recent success and improvements of Vision Transformers over CNNs for many vision tasks, the state of the art uses CNN-based approaches for novel object pose estimation. This work evaluates and demonstrates the differences between self-supervised CNNs and Vision Transformers for deep template matching. In detail, both types of approaches are trained using contrastive learning to match training images against rendered templates of isolated objects. At test time such templates are matched against query images of known and novel objects under challenging settings, such as clutter, occlusion and object symmetries, using masked cosine similarity. The presented results not only demonstrate that Vision Transformers improve matching accuracy over CNNs but also that for some cases pre-trained Vision Transformers do not need fine-tuning to achieve the improvement. Furthermore, we highlight the differences in optimization and network architecture when comparing these two types of networks for deep template matching. Stefan Thalhammer, Jean-Baptiste Weibel, Markus Vincze, José García Rodríguez 0001 |
Image Vis. Comput. | 3 |
| 2022 | GigaDepth: Learning Depth from Structured Light with Branching Neural Networks
Simon Schreiberhuber, Jean-Baptiste Weibel, Tim Patten, Markus Vincze |
ECCV (33) | 4 |
| 2022 | A New VR Kitchen Environment for Recording Well Annotated Object Interaction TasksabstractThis paper presents the Virtual Annotated Cooking Environment (VACE), a new open-source virtual reality dataset (https://sites.google.com/view/vacedataset) and simulator (https://github.com/michaelkoller/vacesimulator) for object inter-action tasks in a rich kitchen environment. We use the Unity-based VR simulator to create thoroughly annotated video se-quences of a virtual human avatar performing food preparation activities. Based on the MPII Cooking 2 dataset, it enables the recreation of recipes for meals such as sandwiches, pizzas, fruit salads and smaller activity sequences such as cutting vegetables. For complex recipes, multiple samples are present, following different orderings of valid partially ordered plans. The dataset includes an RGB and depth camera view, bounding boxes, object masks segmentation, human joint poses and object poses, as well as ground truth interaction data in the form of temporally labeled semantic predicates (holding, on, in, colliding, moving, cutting). In our effort to make the simulator accessible as an open-source tool, researchers are able to expand the setting and annotation to create additional data samples. Michael Koller 0002, Tim Patten, Markus Vincze |
HRI | 3 |
| 2022 | Event-based high-speed low-latency fiducial marker trackingabstractMotion and dynamic environments, especially under challenging lighting conditions, are still an open issue in the field of computer vision. In this paper, we propose an online, end-to-end pipeline for real-time, low latency, 6 degrees-of-freedom pose estimation and tracking of fiducial markers. We employ the high-speed abilities of event-based sensors to directly refine spatial transformations. Furthermore, we introduce a novel two-way verification process for detecting tracking errors by backtracking the estimated pose, allowing to evaluate the quality of our tracking. This approach allows us to achieve pose estimation with an average latency lower than 3 ms and with an average error lower than 5 mm. Adam Loch, Germain Haessig, Markus Vincze |
ICMV | 3 |
| 2022 | SporeAgent: Reinforced Scene-level Plausibility for Object Pose RefinementabstractObservational noise, inaccurate segmentation and ambiguity due to symmetry and occlusion lead to inaccurate object pose estimates. While depth- and RGB-based pose refinement approaches increase the accuracy of the resulting pose estimates, they are susceptible to ambiguity in the observation as they consider visual alignment. We propose to leverage the fact that we often observe static, rigid scenes. Thus, the objects therein need to be under physically plausible poses. We show that considering plausibility reduces ambiguity and, in consequence, allows poses to be more accurately predicted in cluttered environments. To this end, we extend a recent RL-based registration approach towards iterative refinement of object poses. Experiments on the LINEMOD and YCB-VIDEO datasets demonstrate the state-of-the-art performance of our depth-based refinement approach. Code is available at github.com/dornik/sporeagent. Dominik Bauer, Tim Patten, Markus Vincze |
WACV | 3 |
| 2022 | Real-time limb tracking in single depth images based on circle matching and line fittingabstractAbstract Modern lower limb prostheses neither measure nor incorporate healthy residual leg information for intent recognition or device control. In order to increase robustness and reduce misclassification of devices like these, we propose a vision-based solution for real-time 3D human contralateral limb tracking (CoLiTrack). An inertial measurement unit and a depth camera are placed on the side of the prosthesis. The system is capable of estimating the shank axis of the healthy leg. Initially, the 3D input is transformed into a stabilized coordinate system. By splitting the subsequent shank estimation problem into two less computationally intensive steps, the computation time is significantly reduced: First, an iterative closest point algorithm is applied to fit circular models against 2D projections. Second, the random sample consensus method is used to determine the final shank axis. In our study, three experiments were conducted to validate the static, the dynamic and the real-world performance of our CoLiTrack approach. The shank angle can be tracked at 20 Hz for one sixth of the entire human gait cycle with an angle estimation error below $$2.8\pm 2.1^{\circ }$$ 2.8 ± 2 . 1 ∘ . Our promising results demonstrate the robustness of the novel CoLiTrack approach to make “next-generation prostheses” more user-friendly, functional and safe. Michael Tschiedel, Michael Friedrich Russold, Eugenijus Kaniusas, Markus Vincze |
Vis. Comput. | 4 |
| 2021 | ReAgent: Point Cloud Registration Using Imitation and Reinforcement LearningabstractPoint cloud registration is a common step in many 3D computer vision tasks such as object pose estimation, where a 3D model is aligned to an observation. Classical registration methods generalize well to novel domains but fail when given a noisy observation or a bad initialization. Learning-based methods, in contrast, are more robust but lack in generalization capacity. We propose to consider iterative point cloud registration as a reinforcement learning task and, to this end, present a novel registration agent (ReAgent). We employ imitation learning to initialize its discrete registration policy based on a steady expert policy. Integration with policy optimization, based on our proposed alignment reward, further improves the agent’s registration performance. We compare our approach to classical and learning-based registration methods on both ModelNet40 (synthetic) and ScanObjectNN (real data) and show that our ReAgent achieves state-of-the-art accuracy. The lightweight architecture of the agent, moreover, enables reduced inference time as compared to related approaches. Code is available at github.com/dornik/reagent. Dominik Bauer, Tim Patten, Markus Vincze |
CVPR | 3 |
| 2021 | PyraPose: Feature Pyramids for Fast and Accurate Object Pose Estimation under Domain ShiftabstractObject pose estimation enables robots to understand and interact with their environments. Training with synthetic data is necessary in order to adapt to novel situations. Unfortunately, pose estimation under domain shift, i.e., training on synthetic data and testing in the real world, is challenging. Deep learning-based approaches currently perform best when using encoder-decoder networks but typically do not generalize to new scenarios with different scene characteristics. We argue that patch-based approaches, instead of encoder-decoder networks, are more suited for synthetic-to-real transfer because local to global object information is better represented. To that end, we present a novel approach based on a specialized feature pyramid network to compute multi-scale features for creating pose hypotheses on different feature map resolutions in parallel. Our single-shot pose estimation approach is evaluated on multiple standard datasets and outperforms the state of the art by up to ∼35 %. We also perform grasping experiments in the real world to demonstrate the advantage of using synthetic data to generalize to novel environments. Stefan Thalhammer, Markus Leitner, Tim Patten, Markus Vincze |
ICRA | 4 |
| 2021 | Measuring the Sim2Real Gap in 3D Object Classification for Different 3D Data Representation
Jean-Baptiste Weibel, Rainer Rohrböck, Markus Vincze |
ICVS | 3 |
| 2021 | UnrealROX+: An Improved Tool for Acquiring Synthetic Data from Virtual 3D EnvironmentsabstractSynthetic data generation has become essential in last years for feeding data-driven algorithms, which surpassed traditional techniques performance in almost every computer vision problem. Gathering and labelling the amount of data needed for these data-hungry models in the real world may become unfeasible and error-prone, while synthetic data give us the possibility of generating huge amounts of data with pixel-perfect annotations. However, most synthetic datasets lack from enough realism in their rendered images. In that context UnrealROX generation tool was presented in 2019, allowing to generate highly realistic data, at high resolutions and framerates, with an efficient pipeline based on Unreal Engine, a cutting-edge videogame engine. UnrealROX enabled robotic vision researchers to generate realistic and visually plausible data with full ground truth for a wide variety of problems such as class and instance semantic segmentation, object detection, depth estimation, visual grasping, and navigation. Nevertheless, its workflow was very tied to generate image sequences from a robotic on-board camera, making hard to generate data for other purposes. In this work, we present UnrealROX+, an improved version of UnrealROX where its decoupled and easy-to-use data acquisition system allows to quickly design and generate data in a much more flexible and customizable way. Moreover, it is packaged as an Unreal plug-in, which makes it more comfortable to use with already existing Unreal projects, and it also includes new features such as generating albedo or a Python API for interacting with the virtual environment from Deep Learning frameworks. Pablo Martinez-Gonzalez, Sergiu Ovidiu-Oprea, John Alejandro Castro-Vargas, Alberto Garcia-Garcia, Sergio Orts, José García Rodríguez 0001, Markus Vincze |
IJCNN | 7 |
| 2021 | Object Learning for 6D Pose Estimation and Grasping from RGB-D Videos of In-hand ManipulationabstractObject models are highly useful for robots as they enable tasks such as detection, pose estimation and manipulation. However, models are not always easily available, especially in real-world domains of operation such as peoples’ homes. This work presents a pipeline to generate high-quality object reconstructions from human in-hand manipulation to alleviate the necessity of specialised or expensive hardware. Missing data, due to occlusion or unseen sides, is explicitly handled by incorporating shape completion. We demonstrate the usability of the reconstructions by applying a model-based as well as a CNN-based object pose estimator that is trained on synthetic images by employing state-of-the-art texture synthesis. Using our pipeline to cheaply generate object models and synthetic RGB images for training, we achieve competitive performance compared to baselines that require an elaborate set-up to construct models or large amounts of annotated data. Object grasping is also enabled by learning with the reconstructions in simulation, then executing with a real robot. These evaluations show that our reconstructions are comparable to those made under near-perfect conditions and enable 6D object pose estimation as well as real-world grasping. Tim Patten, Kiru Park, Markus Leitner, Kevin Wolfram, Markus Vincze |
IROS | 5 |
| 2021 | Investigating Transparency Methods in a Robot Word-Learning System and Their Effects on Human Teaching BehaviorsabstractRobots need to understand words for references in social spaces (e.g., objects, locations, actions). Grounded language learning systems aim to learn these words from observing a human tutor. Teaching a robot is difficult for naive users due to the discrepancy between the users' mental model and the actual state of the robot. We present a grounded word-learning system with the Pepper robot which learns object and action labels and investigate two extensions geared towards increasing the system’s transparency. The first extension utilizes deictic gestures (pointing and gaze) to communicate knowledge about object names and to further request new labels. The second extension shows the current state of the lexicon on the robot’s tablet. We performed a user study (n=32) to investigate the effects of the transparency methods on learning performance and teaching behavior. In a quantitative analysis, we did not see a significant performance increase for the two extensions. However, users reported higher perception of control and perceived learning success the better they knew the current state of the learning system. In a qualitative analysis, we investigated the participants' teaching behaviors and identified factors that inhibited the learning process. Among other things, we found increased interactive behavior of users when the robot displayed deictic gestures. We saw that human tutors simplified their utterances over time to adapt to the perceived capabilities of the robot. The tablet was most helpful for users to understand what the robot had already learned. Still, learning was impaired in all conditions, when the human input substantially deviated from the form required by the learning system. Matthias Hirschmanner, Stephanie Gross, Setareh Zafari, Brigitte Krenn, Friedrich Neubarth, Markus Vincze |
RO-MAN | 6 |
| 2020 | Neural Object Learning for 6D Pose Estimation Using a Few Cluttered Images
Kiru Park, Tim Patten, Markus Vincze |
ECCV (4) | 3 |
| 2020 | Robust and Efficient Object Change Detection by Combining Global Semantic Information and Local Geometric VerificationabstractIdentifying new, moved or missing objects is an important capability for robot tasks such as surveillance or maintaining order in homes, offices and industrial settings. However, current approaches do not distinguish between novel objects or simple scene readjustments nor do they sufficiently deal with localization error and sensor noise. To overcome these limitations, we combine the strengths of global and local methods for efficient detection of novel objects in 3D reconstructions of indoor environments. Global structure, determined from 3D semantic information, is exploited to establish object candidates. These are then locally verified by comparing isolated geometry to a reference reconstruction provided by the task. We evaluate our approach on a novel dataset containing different types of rooms with 31 scenes and 260 annotated objects. Experiments show that our proposed approach significantly outperforms baseline methods. Edith Langer, Tim Patten, Markus Vincze |
IROS | 3 |
| 2020 | Positive-unlabeled learning for open set domain adaptation
Mohammad Reza Loghmani, Markus Vincze, Tatiana Tommasi |
Pattern Recognit. Lett. | 2 |
| 2020 | Results of Field Trials with a Mobile Service Robot for Older Adults in 16 Private HouseholdsabstractIn this article, we present results obtained from field trials with the Hobbit robotic platform, an assistive, social service robot aiming at enabling prolonged independent living of older adults in their own homes. Our main contribution lies within the detailed results on perceived safety, usability, and acceptance from field trials with autonomous robots in real homes of older users. In these field trials, we studied how 16 older adults (75 plus) lived with autonomously interacting service robots over multiple weeks. Robots have been employed for periods of months previously in home environments for older people, and some have been tested with manipulation abilities, but this is the first time a study has tested a robot in private homes that provided the combination of manipulation abilities, autonomous navigation, and non-scheduled interaction for an extended period of time. This article aims to explore how older adults interact with such a robot in their private homes. Our results show that all users interacted with Hobbit daily, rated most functions as well working, and reported that they believe that Hobbit will be part of future elderly care. We show that Hobbit’s adaptive behavior approach towards the user increasingly eased the interaction between the users and the robot. Our trials reveal the necessity to move into actual users’ homes, as only there, we encounter real-world challenges and demonstrate issues such as misinterpretation of actions during non-scripted human-robot interaction. Markus Bajones, David Fischinger, Astrid Weiss, Paloma de la Puente, Daniel Wolf, Markus Vincze, Tobias Körtner, Markus Weninger, Konstantinos E. Papoutsakis, Damien Michel, Ammar Qammaz, Paschalis Panteleris, Michalis Foukarakis, Ilia Adami, Danae Ioannidi, Asterios Leonidis, Margherita Antona, Antonis A. Argyros, Peter Mayer 0002, Paul Panek, Håkan Eftring, Susanne Frennert |
ACM Trans. Hum. Robot Interact. | 6 |
| 2019 | SyDPose: Object Detection and Pose Estimation in Cluttered Real-World Depth Images Trained using Only Synthetic DataabstractObject pose estimation is an important problem in robotics because it supports scene understanding and enables subsequent grasping and manipulation. Many methods, including modern deep learning approaches, exploit known object models, however, in industry these are difficult and expensive to obtain. 3D CAD models, on the other hand, are often readily available. Consequently, training a deep architecture for pose estimation exclusively from CAD models leads to a considerable decrease of the data creation effort. While this has been shown to work well for feature-and template-based approaches, real-world data is still required for pose estimation in clutter using deep learning. We use synthetically created depth data with domain-relevant background randomized noise heuristics to train an end-to-end, multi-task network, for pose estimation. We simultaneously detect, classify and estimate the poses of texture-less objects in cluttered real-world depth images of an arbitrary amount of objects. We present the results of our experiments with the LineMOD and the Occlusion dataset. Stefan Thalhammer, Tim Patten, Markus Vincze |
3DV | 3 |
| 2019 | A Pilot Study on Determining the Relation Between Gaze Aversion and Interaction ExperienceabstractPrevious work in HHI and HRI demonstrates the impact of gaze on the human interaction experience (IE). In this paper, we discuss an experimental design that should enable measuring the influence of the gaze aversion ratio (GAR) on the users' IE with a social robot. We assume gaze behavior studied in HHI to be restrictive for autonomous social robots, limiting the time available to robots for perception tasks besides HRI. Our goal is to determine if a deviation from human gaze behavior is accepted by users in HRI. With an in-between experimental design we evaluate the effect of varied GAR on IE and behavioral measures. A pilot study with 9 participants suggests that averting the gaze for longer time spans is favorable. Michael Koller 0002, Dominik Bauer, Jesse de Pagter, Guglielmo Papagni, Markus Vincze |
HRI | 5 |
| 2019 | Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationabstractEstimating the 6D pose of objects using only RGB images remains challenging because of problems such as occlusion and symmetries. It is also difficult to construct 3D models with precise texture without expert knowledge or specialized scanning devices. To address these problems, we propose a novel pose estimation method, Pix2Pose, that predicts the 3D coordinates of each object pixel without textured models. An auto-encoder architecture is designed to estimate the 3D coordinates and expected errors per pixel. These pixel-wise predictions are then used in multiple stages to form 2D-3D correspondences to directly compute poses with the PnP algorithm with RANSAC iterations. Our method is robust to occlusion by leveraging recent achievements in generative adversarial training to precisely recover occluded parts. Furthermore, a novel loss function, the transformer loss, is proposed to handle symmetric objects by guiding predictions to the closest symmetric pose. Evaluations on three different benchmark datasets containing symmetric and occluded objects show our method outperforms the state of the art using only RGB images. Kiru Park, Tim Patten, Markus Vincze |
ICCV | 3 |
| 2019 | Multi-Task Template Matching for Object Detection, Segmentation and Pose Estimation Using Depth ImagesabstractTemplate matching has been shown to accurately estimate the pose of a new object given a limited number of samples. However, pose estimation of occluded objects is still challenging. Furthermore, many robot application domains encounter texture-less objects for which depth images are more suitable than color images. In this paper, we propose a novel framework, Multi-Task Template Matching (MTTM), that finds the nearest template of a target object from a depth image while predicting segmentation masks and a pose transformation between the template and a detected object in the scene using the same feature map of the object region. The proposed feature comparison network computes segmentation masks and pose predictions by comparing feature maps of templates and cropped features of a scene. The segmentation result from this network improves the robustness of the pose estimation by excluding points that do not belong to the object. Experimental results show that MTTM outperforms baseline methods for segmentation and pose estimation of occluded objects despite using only depth images. Kiru Park, Tim Patten, Johann Prankl, Markus Vincze |
ICRA | 4 |
| 2019 | ScalableFusion: High-resolution Mesh-based Real-time 3D ReconstructionabstractDense 3D reconstructions generate globally consistent data of the environment suitable for many robot applications. Current RGB-D based reconstructions, however, only maintain the color resolution equal to the depth resolution of the used sensor. This firmly limits the precision and realism of the generated reconstructions. In this paper we present a real-time approach for creating and maintaining a surface reconstruction in as high as possible geometrical fidelity with full sensor resolution for its colorization (or surface texture). A multi-scale memory management process and a Level of Detail scheme enable equally detailed reconstructions to be generated at small scales, such as objects, as well as large scales, such as rooms or buildings. We showcase the benefit of this novel pipeline with a PrimeSense RGB-D camera as well as combining the depth channel of this camera with a high resolution global shutter camera. Further experiments show that our memory management approach allows us to scale up to larger domains that are not achievable with current state-of-the-art methods. Simon Schreiberhuber, Johann Prankl, Tim Patten, Markus Vincze |
ICRA | 4 |
| 2019 | EasyLabel: A Semi-Automatic Pixel-wise Object Annotation Tool for Creating Robotic RGB-D DatasetsabstractDeveloping robot perception systems for recognizing objects in the real world requires computer vision algorithms to be carefully scrutinized with respect to the expected operating domain. This demands large quantities of ground truth data to rigorously evaluate the performance of algorithms. This paper presents the EasyLabel tool for easily acquiring high-quality ground truth annotation of objects at pixel-level in densely cluttered scenes. In a semi-automatic process, complex scenes are incrementally built and EasyLabel exploits depth changes to extract precise object masks at each step. We use this tool to generate the Object Cluttered Indoor Dataset (OCID) that captures diverse settings of objects, background, context, sensor to scene distance, viewpoint angle and lighting conditions. OCID is used to perform a systematic comparison of existing object segmentation methods. The baseline comparison supports the need for pixel- and object-wise annotation to progress robot vision towards realistic applications. This insight reveals the usefulness of EasyLabel and OCID to better understand the challenges that robots face in the real world. Markus Suchi, Tim Patten, David Fischinger, Markus Vincze |
ICRA | 4 |
| 2019 | Robust 3D Object Classification by Combining Point Pair Features and Graph ConvolutionabstractObject classification is an important capability for robots as it provides vital semantic information that underpin most practical high-level tasks. Classic handcrafted features, such as point pair features, have demonstrated their robustness for this task. Combining these features with modern deep learning methods provide discriminative features that are rotation invariant and robust to various sources of noise. In this work, we aim to improve the descriptiveness of point pair features while retaining their robustness. We propose a method to achieve more structured sampling of pairs and combine this information through the use of graph convolutional networks. We introduce a novel attention model based on a repeatable local reference frame. Experiments show that our approach significantly improves the state of the art for object classification on large scale reconstruction such as the Stanford 3D indoor dataset and ScanNet and obtains competitive accuracy on the artificial dataset ModelNet. Jean-Baptiste Weibel, Tim Patten, Markus Vincze |
ICRA | 3 |
| 2019 | Leveraging Symmetries to Improve Object Detection and Pose Estimation from Range Data
Sergey V. Alexandrov, Tim Patten, Markus Vincze |
ICVS | 3 |
| 2019 | Monte Carlo Tree Search on Directed Acyclic Graphs for Object Pose Verification
Dominik Bauer, Tim Patten, Markus Vincze |
ICVS | 3 |
| 2019 | Tillage Machine Control Based on a Vision System for Soil Roughness and Soil Cover Estimation
Peter Riegler-Nurscher, Johann Prankl, Markus Vincze |
ICVS | 3 |
| 2018 | Introducing storytelling to educational robotic activitiesabstractCreativity is a skill that has been recognized as one of the 21stcentury skills. Likewise robotics has been recognized as technology with several features to enthrall children and be used to teach a variety of topics (e.g. Mathematics and programming). This paper presents a study done to verify the impact of introducing a storytelling session in an activity with children from 6 to 18 years old in Austria. A total of 196 participants participated in the workshops, held in the campus of the university. Quantitative and qualitative data were collected and analyzed. The results showed that the most difficult task for old participants was the collaboration between groups created. Participants also did not mention the use of creativity during the design and implementation of the story. Instead participants referred to the work done with robots and technology. Moreover, participants think that working with robots is interesting and fun. Julian M. Angel Fernandez, Markus Vincze |
EDUCON | 2 |
| 2018 | Multi-View 3D Entangled Forest for Semantic Segmentation and MappingabstractApplications that provide location related services need to understand the environment in which humans live such that verbal references and human interaction are possible. We formulate this semantic labelling task as the problem of learning the semantic labels from the perceived 3D structure. In this contribution we propose a batch approach and a novel multi-view frame fusion technique to exploit multiple views for improving the semantic labelling results. The batch approach works offline and is the direct application of an existing single-view method to scene reconstructions with multiple views. The multi-view frame fusion works in an incremental fashion accumulating the single-view results, hence allowing the online multi-view semantic segmentation of single frames and the offline reconstruction of semantic maps. Our experiments show the superiority of the approaches based on our fusion scheme, which leads to a more accurate semantic labelling. Morris Antonello, Daniel Wolf, Johann Prankl, Stefano Ghidoni, Emanuele Menegatti, Markus Vincze |
ICRA | 6 |
| 2018 | Recognizing Objects in-the-Wild: Where do we Stand?abstractThe ability to recognize objects is an essential skill for a robotic system acting in human-populated environments. Despite decades of effort from the robotic and vision research communities, robots are still missing good visual perceptual systems, preventing the use of autonomous agents for realworld applications. The progress is slowed down by the lack of a testbed able to accurately represent the world perceived by the robot in-the-wild. In order to fill this gap, we introduce a large-scale, multi-view object dataset collected with an RGB-D camera mounted on a mobile robot. The dataset embeds the challenges faced by a robot in a real-life application and provides a useful tool for validating object recognition algorithms. Besides describing the characteristics of the dataset, the paper evaluates the performance of a collection of well-established deep convolutional networks on the new dataset and analyzes the transferability of deep representations from Web images to robotic data. Despite the promising results obtained with such representations, the experiments demonstrate that object classification with real-life robotic data is far from being solved. Finally, we provide a comparative study to analyze and highlight the open challenges in robot vision, explaining the discrepancies in the performance. Mohammad Reza Loghmani, Barbara Caputo, Markus Vincze |
ICRA | 3 |
| 2018 | Towards Autonomous Auto Calibration of Unregistered RGB-D Setups: The Benefit of Plane PriorsabstractIn the last few years novel color and depth (RGB-D) sensors have greatly pushed robot perception. To enable a precise pixel-wise fusion of color and depth information good calibration is needed. The calibration determines the intrinsic parameters, the extrinsic parameters, and corrects for depth errors. While classic calibration approaches involve a dedicated calibration target and a trained expert, the autonomous calibration of such camera systems for robots operating in unknown environments is still an open problem. It demands for robust methods that do not need an expert to set up or tune the algorithm. Hence, we present a robust calibration algorithm that utilizes structure from motion (SfM) reconstructions as a calibration target and incorporates plane priors in the optimization to improve the convergence behavior and improve the calibration robustness. We evaluate our method against the state of the art performing over 300 experiments on ten different datasets, and show a significant improvement of the calibration accuracy. Georg Halmetschlager-Funek, Johann Prankl, Markus Vincze |
IROS | 3 |
| 2018 | Action Selection for Interactive Object Segmentation in ClutterabstractRobots operating in human environments are often required to recognise, grasp and manipulate objects. Identifying the locations of objects amongst their complex surroundings is therefore an important capability. However, when environments are unstructured and cluttered, as is typical for indoor human environments, reliable and accurate object segmentation is not always possible because the scene representation is often incomplete or ambiguous. We overcome the limitations of static object segmentation by enabling a robot to directly interact with the scene with non-prehensile actions. Our method does not rely on object models to infer object existence. Rather, interaction induces scene motion and this provides an additional clue for associating observed parts to the same object. We use a probabilistic segmentation framework in order to identify segmentation uncertainty. This uncertainty is then used to guide a robot while it manipulates the scene. Our probabilistic segmentation approach recursively updates the segmentation given the motion cues and the segmentation is monitored during interaction, thus providing online feedback. Experiments performed with RGB-D data show that the additional source of information from motion enables more certain object segmentation that was otherwise ambiguous. We then show that our interaction approach based on segmentation uncertainty maintains higher quality segmentation than competing methods with increasing clutter. Tim Patten, Michael Zillich, Markus Vincze |
IROS | 3 |
| 2018 | Grounded Word Learning on a Pepper RobotabstractIn this demonstration, we will showcase realtime grounded language learning on the humanoid robot Pepper. In particular, learning word-object and word-action mapping from cross-modal data, where simple actions, such as take, put and push, are shown to the robot by a human tutor. The visual demonstration is accompanied by verbal descriptions of the performed actions, such as I take the box and put it next to the bottle. Learning was realized on the humanoid robot Pepper using the Google Speech API for speech to text and the robot's camera system for object tracking. Matthias Hirschmanner, Stephanie Gross, Brigitte Krenn, Friedrich Neubarth, Martin Trapp 0001, Markus Vincze |
IVA | 6 |
| 2017 | High Dynamic Range SLAM with Map-Aware Exposure Time ControlabstractThe research in dense online 3D mapping is mostly focused on the geometrical accuracy and spatial extent of the reconstructions. Their color appearance is often neglected, leading to inconsistent colors and noticeable artifacts. We rectify this by extending a state-of-the-art SLAM system to accumulate colors in HDR space. We replace the simplistic pixel intensity averaging scheme with HDR color fusion rules tailored to the incremental nature of SLAM and a noise model suitable for off-the-shelf RGB-D cameras. Our main contribution is a map-aware exposure time controller. It makes decisions based on the global state of the map and predicted camera motion, attempting to maximize the information gain of each observation. We report a set of experiments demonstrating the improved texture quality and advantages of using the custom controller that is tightly integrated in the mapping loop. Sergey V. Alexandrov, Johann Prankl, Michael Zillich, Markus Vincze |
3DV | 4 |
| 2017 | Multi-label Point Cloud Annotation by Selection of Sparse Control PointsabstractThis paper presents a user-friendly approach for multi-label point cloud annotation. The method requires the user to select sparse control points belonging to the objects through a mouse-based interface. Multiple control points may be assigned to the same label. The software utilizes the selected control points to perform a segmentation algorithm on the neighborhood graph, based on shortest path tree. The user is provided a real-time feedback about the result, and can correct segmentation errors. In contrast to previous work the method supports multi-label annotation of unorganized point clouds. The method has been evaluated by multiple users and compared with a standard rectangle-based selection technique. Results indicate that the proposed method is perceived as easier to use, and that it allows a faster segmentation even in complex scenarios with occlusions. Riccardo Monica, Jacopo Aleotti, Michael Zillich, Markus Vincze |
3DV | 4 |
| 2017 | Designing Emotionally Expressive Robots: A Comparative Study on the Perception of Communication ModalitiesabstractSocially assistive agents, be it virtual avatars or robots, need to engage in social interactions with humans and express their internal emotional states, goals, and desires. In this work, we conducted a comparative study to investigate how humans perceive emotional cues expressed by humanoid robots through five communication modalities (face, head, body, voice, locomotion) and examined whether the degree of a robot's human-like embodiment affects this perception. In an online survey, we asked people to identify emotions communicated by Pepper - a highly human-like robot and Hobbit - a robot with abstract humanlike features. A qualitative and quantitative data analysis confirmed the expressive power of the face, but also demonstrated that body expressions or even simple head and locomotion movements could convey emotional information. These findings suggest that emotion recognition accuracy varies as a function of the modality, and a higher degree of anthropomorphism does not necessarily lead to a higher level of recognition accuracy. Our results further the understanding of how people respond to single communication modalities and have implications for designing recognizable multimodal expressions for robots. Christiana Tsiourti, Astrid Weiss, Katarzyna Wac, Markus Vincze |
HAI | 4 |
| 2017 | Learning the Floor Type for Automated Detection of Dirt Spots for Robotic Floor Cleaning Using Gaussian Mixture Models
Andreas Grünauer, Georg Halmetschlager-Funek, Johann Prankl, Markus Vincze |
ICVS | 4 |
| 2017 | RGB-D fusion enhancement by mode filter for surfel cloud segmentationabstractThis paper presents an algorithm for surfel color and position enhancement from RGB-D data acquired across multiple image frames. Surfel-based reconstruction algorithms associate each RGB-D frame pixel to a surfel in the model. As the reconstruction progresses, surfel color and position are the average of all observations. Our proposed algorithm is designed to enhance position discontinuities and to produce sharper colors, to facilitate subsequent segmentation steps on the 3D model. During reconstruction, several colors and positions are tracked for each surfel. Only at the end of reconstruction phase the most frequent value is chosen through a Winner Takes All policy. The result has been compared to the standard averaging policy of reconstruction algorithms. Experiments have been performed using both Flood Fill and Supervoxel-LCCP segmentation and by applying two segmentation evaluation metrics. Results show that the proposed method is suitable to enhance a surfel-based model for object segmentation purposes. Riccardo Monica, Michael Zillich, Markus Vincze, Jacopo Aleotti |
IROS | 3 |
| 2017 | Determining the effect of programming language in educational robotic activitiesabstractRobotics has been suggested as a field of high potential in education and with high expectancy to impact teaching from kindergarten to university. This paper presents a study conducted with the purpose to have a better understanding of the impact of programming languages among participants of a workshop. An activity, which encompasses ten exercises, was designed and we used three different programming languages (i.e. visual, blocky and text) to program the Thymio robot. A total of six workshops were held using this activity two for each programming language. Qualitative and quantitative data were collected in each workshop. The results suggest that despite the programming language used, participants enjoyed working with robots. Moreover participants with previous experience on programming prefer more advance programming languages. Julian M. Angel Fernandez, Markus Vincze |
RO-MAN | 2 |
| 2016 | Guided Matching Based on Statistical Optical Flow for Fast and Robust Correspondence Analysis
Josef Maier, Martin Humenberger, Markus Murschitz, Oliver Zendel 0001, Markus Vincze |
ECCV (7) | 5 |
| 2016 | Results of a Real World Trial with a Mobile Social Service Robot for Older AdultsabstractRobots are an increasingly discussed solution for assistance of seniors. Importance of testing natural interaction therefore becomes crucial. This paper presents first results of a study with an autonomous mobile social service robot prototype that was deployed in 18 private households of senior adults aged 75 years and older for a total of 371 days. Findings show that utility met the users' expectations. However, the robot was rather seen as a toy instead of being supportive for independent living. Furthermore, despite of an emergency function of the robot, perceived safety did not increase. Reasons for this might be the good health conditions of our users, a lack of technological robustness and slow performance of the prototype. However, users believed that a market ready version of the robot would be vital for supporting people who are more fragile and more socially isolated. Jürgen Pripfl, Tobias Körtner, Daliah Batko-Klein, Denise Hebesberger, Markus Weninger, Christoph Gisinger, Susanne Frennert, Håkan Eftring, Margarita Antona, Ilia Adami, Astrid Weiss, Markus Bajones, Markus Vincze |
HRI | 13 |
| 2016 | A multi-modal RGB-D object recognizerabstractIn this paper we propose a multi-modal object recognition system that uses a two-step hypothesis verification approach to improve runtime efficiency. The system uses local and global appearance and shape features, generating many possibly competing hypotheses, which are then verified such that the scene can be optimally explained in terms of recognized object models. The introduced modification in this time consuming step reduces runtime considerably, while maintaining recognition performance. We evaluate recognition performance for various feature extraction modalities on the publicly available Willow Garage RGB-D dataset and show runtime improvements of a factor 2 to 10. Thomas Faulhammer, Michael Zillich, Johann Prankl, Markus Vincze |
ICPR | 4 |
| 2016 | Calibration and correction of vignetting effects with an application to 3D mappingabstractCheap RGB-D sensors are ubiquitous in robotics. They typically contain a consumer-grade color camera that suffers from significant optical nonlinearities, often referred to as vignetting effects. For example, in Asus Xtion Live Pro cameras the pixels in the corners are two times darker than those in the center of the image. This deteriorates the visual appearance of 3D maps built with such cameras. We propose a simple calibration method that only requires a sheet of white paper as a calibration object and allows to reliably recover the vignetting response of a camera. We demonstrate calibration results for multiple popular RGB-D sensors and show that removal of vignetting effects using a nonparametric response model results in improved color coherence of the reconstructed maps. Furthermore, we show how to effectively compensate color variations caused by automatic white balance and exposure time control of the camera. Sergey V. Alexandrov, Johann Prankl, Michael Zillich, Markus Vincze |
IROS | 4 |
| 2016 | A Global Hypothesis Verification Framework for 3D Object Recognition in ClutterabstractPipelines to recognize 3D objects despite clutter and occlusions usually end up with a final verification stage whereby recognition hypotheses are validated or dismissed based on how well they explain sensor measurements. Unlike previous work, we propose a Global Hypothesis Verification (GHV) approach which regards all hypotheses jointly so as to account for mutual interactions. GHV provides a principled framework to tackle the complexity of our visual world by leveraging on a plurality of recognition paradigms and cues. Accordingly, we present a 3D object recognition pipeline deploying both global and local 3D features as well as shape and color. Thereby, and facilitated by the robustness of the verification process, diverse object hypotheses can be gathered and weak hypotheses need not be suppressed too early to trade sensitivity for specificity. Experiments demonstrate the effectiveness of our proposal, which significantly improves over the state-of-art and attains ideal performance (no false negatives, no false positives) on three out of the six most relevant and challenging benchmark datasets. Aitor Aldoma, Federico Tombari, Luigi Di Stefano, Markus Vincze |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Mattie: a simple educational platform for children to realize their first robot prototypeabstractThe landscape of robotics platforms for education is broad and mainly focused on technical problem solving skills. The different application areas serve for adding some variety regarding skill sets and areas of interest, however design or product development skills are often neglected. Mattie is a simple robotic platform for children to realize their first robotic product ideas. The design is kept very simple. Materials are affordable and easily available. The aim is to guide children to build their first robot prototype from scratch by learning product development and scientific working. The concept around the Mattie robot incorporates different perspectives like product, behavior, marketing or design besides technology and engineering. This way, robotics can appeal to those children who are not interested in engineering or programming right away. We have successfully used the Mattie robot in classroom workshop settings with seven different classes (6th, 7th and 9th grades). The feedback from teachers and students aged 11-18 is very positive. Matthias Hirschmanner, Lara Lammer, Markus Vincze |
IDC | 3 |
| 2015 | Temporal integration of feature correspondences for enhanced recognition in cluttered and dynamic environmentsabstractWe propose a method for recognizing rigid object instances in RGB-D point clouds by accumulating low-level information from keypoint correspondences over multiple observations. Compared to existing multi-view approaches, we make fewer assumptions on the recognition problem, dealing with cluttered and partially dynamic environments as well as covering a wide range of objects. Evaluation on the publicly available TUW and Willow datasets showed that our method achieves state-of-the-art recognition performance for challenging sequences of static environments and a significant improvement for environments partially changing during the observation. Thomas Faulhammer, Aitor Aldoma, Michael Zillich, Markus Vincze |
ICRA | 4 |
| 2015 | Saliency-based object discovery on RGB-D data with a late-fusion approachabstractWe present a novel method based on saliency and segmentation to generate generic object candidates from RGB-D data. Our method uses saliency as a cue to roughly estimate the location and extent of the objects present in the scene. Salient regions are used to glue together the segments obtained from over-segmenting the scene by either color or depth segmentation algorithms, or by a combination of both. We suggest a late-fusion approach that first extracts segments from color and depth independently before fusing them to exploit that the data is complementary. Furthermore, we investigate several mechanisms for ranking the object candidates. We evaluate on one publicly available dataset and on one challenging sequence with a high degree of clutter. The results show that we are able to retrieve most objects in real-world indoor scenes and clearly outperform other state-of-the art methods. Germán Martín García, Ekaterina Potapova, Thomas Werner, Michael Zillich, Markus Vincze, Simone Frintrop |
ICRA | 5 |
| 2015 | Fast semantic segmentation of 3D point clouds using a dense CRF with learned parametersabstractIn this paper, we present an efficient semantic segmentation framework for indoor scenes operating on 3D point clouds. We use the results of a Random Forest Classifier to initialize the unary potentials of a densely interconnected Conditional Random Field, for which we learn the parameters for the pairwise potentials from training data. These potentials capture and model common spatial relations between class labels, which can often be observed in indoor scenes. We evaluate our approach on the popular NYU Depth datasets, for which it achieves superior results compared to the current state of the art. Exploiting parallelization and applying an efficient CRF inference method based on mean field approximation, our framework is able to process full resolution Kinect point clouds in half a second on a regular laptop, more than twice as fast as comparable methods. Daniel Wolf, Johann Prankl, Markus Vincze |
ICRA | 3 |
| 2015 | The 5-Step Plan - Empowered Children's Robotic Product Ideas
Lara Lammer, Astrid Weiss, Markus Vincze |
INTERACT (2) | 3 |
| 2015 | Fast and accurate normal estimation by efficient 3d edge detectionabstractAccurate surface normal computation is one of the most basic and important tasks for 3d perception. While much progress has been made in speeding up normal estimation algorithms and improving their accuracy, a significant inaccuracy still remains even with modern implementations, which is the correct determination of surface normals close to non-differentiable surface edges. Current algorithms tend to amalgamate neighborhood points from independent surfaces yielding normals that neither fit well to the one nor the other surface. This paper introduces a fast and accurate 3d edge detection algorithm suitable to detect discontinuities both in depth and on surfaces with nearly 90% accuracy at rates beyond 30 Hz. Based on this method, we demonstrate how established normal estimation algorithms can be extended for edge-awareness. Additionally, a new edge-aware, fast, accurate, and robust normal estimation approach is described which exploits the data structures computed for 3d edge detection and estimates normals at 23 Hz. We assess the performance of all proposed methods and compare them with other state-of-the-art approaches. Richard Bormann, Joshua Hampp, Martin Hägele, Markus Vincze |
IROS | 4 |
| 2015 | RGB-D object modelling for object recognition and trackingabstractThis work presents a flexible system to reconstruct 3D models of objects captured with an RGB-D sensor. A major advantage of the method is that unlike other modelling tools, our reconstruction pipeline allows the user to acquire a full 3D model of the object. This is achieved by acquiring several partial 3D models in different sessions-each individual session presenting the object of interest in different configurations that reveal occluded parts of the object - that are automatically merged together to reconstruct a full 3D model. In addition, the 3D models acquired by our system can be directly used by state-of-the-art object instance recognition and object tracking modules, providing object-perception capabilities to complex applications requiring these functionalities (e.g. human-object interaction analysis, robot grasping, etc.). The system does not impose constraints in the appearance of objects (textured, untextured) nor in the modelling setup (moving camera with static object or turn-table setups with static camera). The proposed reconstruction system has been used to model a large number of objects resulting in metrically accurate and visually appealing 3D models. Johann Prankl, Aitor Aldoma, Alexander Svejda, Markus Vincze |
IROS | 4 |
| 2014 | Direct Optimization of T-Splines Based on Multiview StereoabstractWe propose a multi-view stereo reconstruction method in which the surface is represented by CAD-compatible T-splines. Our method hinges on the principle of is geometric analysis, formulating an energy functional that can be directly computed in terms of the T-spline basis. Paying attention to the idiosyncracies of this basis, we derive an analytic formula for the gradient of the functional which is then used in photo-consistency optimization. The numbers of degrees of freedom our model requires is drastically reduced compared to the state of the art. Gains in efficiency can firstly be attributed to the fact that T-splines are particularly suited for adaptive refinement. Secondly, evaluation of the proposed energy functional is highly parallelizable as demonstrated by means of a T-spline-specific GPU implementation. Our experiments indicate the superiority of T-spline surfaces over the widely-used triangular meshes in terms of memory efficiency and numerical stability, without relying on dedicated regularizers. Thomas Morwald, Jonathan Balzer, Markus Vincze |
3DV | 3 |
| 2014 | Socially assistive robots for the aging population: are we trapped in stereotypes?abstractRobots caring for the older population in care facilities and at home is an ongoing theme in HRI research. Research projects on this topic exist all over the globe in the USA, Europe, and Asia. All of these projects have the overall ambitious goal to increase the well-being of older adults and to enable them to stay at home as long as possible. In this workshop we want to reflect whether the HRI community is trapped in stereotypes when it comes to socially assistive robots for older adults' Therefore we want to gather and compare findings from user needs analysis, user evaluation studies, as well as interaction scenarios and functionalities of existing care robots. Are our results suggesting similar scenarios? Do older end users in all countries have similar needs and desires when it comes to assistive robots? What are the challenges and opportunities for future assistive robots (maybe for those we develop for ourselves when we belong to the older population?) also on an ethical and legal level? In this workshop we want to escape the stereotype trap what socially assistive robots should do. Can socially assistive robots solve the aging population problem on a societal and individual level? Are older people in general technology opponents? Will robotic helpers be accepted in the home as long as they pretend to be social actors. Astrid Weiss, Jenay M. Beer, Takanori Shibata, Markus Vincze |
HRI | 4 |
| 2014 | Designing a service robot for public space: an "action and experiences" - approachabstractWhen we think of service robots for public spaces, we often consider human-like systems that socially interact with users in short-term and dynamically changing scenarios. Many design assumptions exist for this type of robot, but what about for a service robot that needs to interact with people in a very task/role-oriented manner? This paper presents a research-through-design approach, which explores the idea of a "luggage-carrying robot guide" for train stations which supports travellers. Contrary to the classical paper describing the requirements, design, implementation, and evaluation of the robot, this work presents an exploratory design study. An interdisciplinary team composed of two industrial designers, a roboticist, and a social scientist performed a study with users to establish a design space for the aforementioned service robot. Astrid Weiss, Markus Bader, Markus Vincze, Gert Hasenhütl, Stefan Moritsch |
HRI | 3 |
| 2014 | Don't bother me: users' reactions to different robot disturbing behaviorsabstractWhen living together in a household with a socially assistive robot, it can happen that the robot disturbs its owner by offering a service. One might argue that a social robot should act according to the social norms people expect of each other, but still then disturbances of daily routines are a challenging endeavor to address. We conducted a preliminary user study in which we explored four different disturbing behaviors the socially assistive robot HOBBIT showed, while the user was focusing on a different primary task. We used the BEHAVE measurement set to evaluate the attitudinal and behavioral responses of the users, which disturbance distracted the user the most from his/her primary task and how the disturbance affected the overall attitudinal response towards the robot. Interestingly, our results showed that the disturbing behavior did not heavily negatively impact the assessment of the robot and that not all types of disturbance did distract the users with the same intensity. Astrid Weiss, Markus Vincze, Paul Panek, Peter Mayer 0002 |
HRI | 2 |
| 2014 | Workshop on attention models in robotics: visual systems for better HRIabstractAttention is a concept of human perception that enables human subjects to select the potentially relevant parts out of the huge amount of sensory data and that enables interactions with other human subjects by sharing attention with each other. These abilities are also of large interest for autonomous robots, therefore, interest in modeling concepts of human attention computationally has increased strongly in the robotics community during the last decade. Especially in human-robot interaction, the ability to detect what a human partner is attending to and to act in a similar way to enable intuitive communication, are important skills for a robotic system. Michael Zillich, Simone Frintrop, Fiora Pirri, Ekaterina Potapova, Markus Vincze |
HRI | 5 |
| 2014 | 4D Space-Time Mereotopogeometry-Part Connectivity Calculus for Visual Object RepresentationabstractRegion Connectivity Calculus (RCC) can be used to define the formal grammar describing the relationship between image regions. While RCC provides a possible framework for representation of object part constellations leading to object recognition, little has been done in the direction of RCC for 3D images. Almost all prior RCC representations, such as RCC5/ RCC8/ RCC23/RCC62 are oriented towards 2D projections. While the recently introduced RCC-3D does address this limitation to a certain extent, it heavily depends on other 2D RCC frameworks and as such is limited -- it provides no representation for orthogonally aligned or staggered object parts. It also does not provide convenient representations for shape related pose information (such as "horizontally aligned along principal axis" or "vertically aligned" etc.). While it is possible to use oriented matroids for projected 3D representations using cocircuits and chirotopes, these are again sub-optimal given that there is considerable information loss (through dimensionality reduction) and multiple object models map to the same structures. In this paper, we introduce a new hierarchical graph based 3D Region/ Surface/ Object/ Part Connectivity Calculus (OCC/PCC), given the domain of affordance based equivalence recognition. We call our PCC -- 4D Space Time Mereo-topo-geometry (4D-STMTG) and the equivalent OCC as 4D Space Time Joint Topo-geometry (4D-STJTG). The modeling of simple objects using the calculus is demonstrated in practical scenarios. Comparisons of the proposed calculus with respect to the state-of-art RCC-3D is also presented, demonstrating the flexibility, suitability and superiority of 4D-STMTG. Karthik Mahesh Varadarajan, Markus Vincze |
ICPR | 2 |
| 2014 | Attention-driven object detection and segmentation of cluttered table scenes using 2.5D symmetryabstractThe task of searching and grasping objects in cluttered scenes, typical of robotic applications in domestic environments requires fast object detection and segmentation. Attentional mechanisms provide a means to detect and prioritize processing of objects of interest. In this work, we combine a saliency operator based on symmetry with a segmentation method based on clustering locally planar surface patches, both operating on 2.5D point clouds (RGB-D images) as input data to yield a novel approach to table-top scene segmentation. Evaluation on indoor table-top scenes containing man-made objects clustered in piles and dumped in a box show that our approach to selection of attention points significantly improves performance of state-of-the-art attention-based segmentation methods. Ekaterina Potapova, Karthik Mahesh Varadarajan, Andreas Richtsfeld, Michael Zillich, Markus Vincze |
ICRA | 5 |
| 2014 | Automation of "ground truth" annotation for multi-view RGB-D object instance recognition datasetsabstractAiming at reducing the labour intensity associated with the acquisition of ground truth annotations for object instance recognition datasets, this paper discusses a novel multi-view recognition method to automate the annotation (object instances and associated poses) of individual images in multi-view RGB-D datasets. In combination with recent single-view object recognition techniques, the supplementary information provided by multiple vantage points results in a rich and integrated representation of the environment, in the form of a 3D reconstructed scene as well as object hypotheses therein. We argue that such a representation facilitates improved recognition to an extent that the recovered results, obtained by means of a suitable 3D hypotheses verification stage, closely resemble the ground truth of the scene under consideration. On two large datasets, totalling more than 3500 object instances, our method yields 99.1% and 93.2% correct automatic annotations. These results corroborate our approach for the task at hand. Aitor Aldoma, Thomas Faulhammer, Markus Vincze |
IROS | 3 |
| 2014 | RGB-D sensor setup for multiple tasks of home robots and experimental resultsabstractWhile navigation based on 2D laser data is well understood, the application of robots at home environments requires seeing more than a slice of the world. RGB-D cameras have been used to perceive the full scenes and solutions exist consuming extensive computing power. We propose a setup with two RGB-D cameras that covers the need for conflicting requirements regarding localization, obstacle avoidance, object search and recognition, and gesture recognition. We show that this setup provides sufficient data to enable navigation at homes and we present how ROS modules can be configured to use virtual RGB-D scans instead of laser data for operation in real-time (10Hz). Finally, we present first results of exploiting this versatile setup for a home service robot that picks up things from the floor to prevent potential falls of its future users. Paloma de la Puente, Markus Bajones, Peter Einramhof, Daniel Wolf, David Fischinger, Markus Vincze |
IROS | 6 |
| 2014 | Learning of perceptual grouping for object segmentation on RGB-D dataabstractObject segmentation of unknown objects with arbitrary shape in cluttered scenes is an ambitious goal in computer vision and became a great impulse with the introduction of cheap and powerful RGB-D sensors. We introduce a framework for segmenting RGB-D images where data is processed in a hierarchical fashion. After pre-clustering on pixel level parametric surface patches are estimated. Different relations between patch-pairs are calculated, which we derive from perceptual grouping principles, and support vector machine classification is employed to learn Perceptual Grouping. Finally, we show that object hypotheses generation with Graph-Cut finds a globally optimal solution and prevents wrong grouping. Our framework is able to segment objects, even if they are stacked or jumbled in cluttered scenes. We also tackle the problem of segmenting objects when they are partially occluded. The work is evaluated on publicly available object segmentation databases and also compared with state-of-the-art work of object segmentation. Andreas Richtsfeld, Thomas Morwald, Johann Prankl, Michael Zillich, Markus Vincze |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | Designing adaptive roles for socially assistive robots: a new method to reduce technological determinism and role stereotypesabstractSocial roles are a design option for robots that behave in accordance with user expectations. We believe that robots have to exceed stereotypical role behaviors and dynamically provide roles that suit the people's living conditions in order to achieve long-term acceptance. We are introducing a new user-focused design method to develop social role repertoires for adaptive human-robot interaction (HRI). The method consists of five sequential steps: (1) user group and application scenario identification; (2) acquisition of users' mental associations; (3) derivation of role traits; (4) prioritization of these traits; and (5) synthesis of an adaptive social role repertoire. We tested our method with two specific user groups: elder adults living at home and those living in care facilities. The results reveal basic role concepts and specific preference clusters in each user group. The empirically based clusters are suitable for the parameterizations and development of robots with adaptive social roles. Andreas Huber, Lara Lammer, Astrid Weiss, Markus Vincze |
J. Hum. Robot Interact. | 4 |
| 2013 | Compressive Distance Classifier Correlation FilterabstractCompressed Sensing (CS) is seen as the pathway to increase the efficiency of sensor systems such as MRI, SAR and SAS while avoiding the huge costs and related processing accompanying high-resolution data acquisition. While there has been a surge in the number of sensor systems and related algorithms using CS, target/object recognition in the sensing domain which offers numerous advantages, is a rather nascent field. The state-of-the-art in this field includes the Smashed Filter (SF), which is a reduced dimensionality maximum likelihood classifier. Nevertheless, the accuracy of the filter remains low for practical applications, especially with variations in scale, translation and rotation in the test data. This paper offers a new type of filter - called the Compressive Distance Classifier Correlation Filter (CDCCF), which applies a transformation in the CS domain thereby increasing the distance between intra-class correlation peaks while reducing the distance between inter-class correlation peaks and is based on the Restricted Isometry Property (RIP) of the compressed manifold and the Johnson Lindenstrauss Lemma. Results presented show that the accuracy of the CDCCF filter is about 70% on a 12 class test data set, which is over a two-fold increase in accuracy over the SF. Confusion matrices, measures of ROC, Mean Average Precision and Accuracy demonstrate the robust performance of the algorithm over SF across different compressive sampling resolutions. Karthik Mahesh Varadarajan, Markus Vincze |
ICIP | 2 |
| 2013 | Multimodal cue integration through Hypotheses Verification for RGB-D object recognition and 6DOF pose estimationabstractThis paper proposes an effective algorithm for recognizing objects and accurately estimating their 6DOF pose in scenes acquired by a RGB-D sensor. The proposed method is based on a combination of different recognition pipelines, each exploiting the data in a diverse manner and generating object hypotheses that are ultimately fused together in an Hypothesis Verification stage that globally enforces geometrical consistency between model hypotheses and the scene. Such a scheme boosts the overall recognition performance as it enhances the strength of the different recognition pipelines while diminishing the impact of their specific weaknesses. The proposed method outperforms the state-of-the-art on two challenging benchmark datasets for object recognition comprising 35 object models and, respectively, 176 and 353 scenes. Aitor Aldoma, Federico Tombari, Johann Prankl, Andreas Richtsfeld, Luigi Di Stefano, Markus Vincze |
ICRA | 6 |
| 2013 | Learning grasps for unknown objects in cluttered scenesabstractIn this paper, we propose a method for grasping unknown objects from piles or cluttered scenes, given a point cloud from a single depth camera. We introduce a shape-based method - Symmetry Height Accumulated Features (SHAF) - that reduces the scene description complexity such that the use of machine learning techniques becomes feasible. We describe the basic Height Accumulated Features and the Symmetry Features and investigate their quality using an F-score metric. We discuss the gain from Symmetry Features for grasp classification and demonstrate the expressive power of Height Accumulated Features by comparing it to a simple height based learning method. In robotic experiments of grasping single objects, we test 10 novel objects in 150 trials and show significant improvement of 34% over a state-of-the-art method, achieving a success rate of 92%. An improvement of 29% over the competitive method was achieved for a task of clearing a table with 5 to 10 objects and overall 90 trials. Furthermore we show that our approach is easily adaptable for different manipulators by running our experiments on a second platform. David Fischinger, Markus Vincze |
ICRA | 2 |
| 2013 | Geometric data abstraction using B-splines for range image segmentationabstractWith the availability of cheap and powerful RGB-D sensors interest in 3D point cloud based methods has drastically increased. One common prerequisite of these methods is to abstract away from raw point cloud data, e.g. to planar patches, to reduce the amount of data and to handle noise and clutter. We present a novel method to abstract RGB-D sensor data to parametric surface models described by B-spline surfaces and associated boundaries. Data is first pre-segmented into smooth patches before B-spline surfaces are fitted. The best surface representations of these patches are selected in a merging procedure. Furthermore, we show how curve fitting estimates smooth boundaries and improves the given sensor information compared to hand-labelled ground truth annotation when using colour in addition to depth information. All parts of the framework are open-source1and are evaluated on the object segmentation database (OSD) also available online, showing accuracy and usability of the proposed methods. Thomas Morwald, Andreas Richtsfeld, Johann Prankl, Michael Zillich, Markus Vincze |
ICRA | 5 |
| 2013 | Probabilistic Cue Integration for Real-Time Object Pose Tracking
Johann Prankl, Thomas Morwald, Michael Zillich, Markus Vincze |
ICVS | 4 |
| 2013 | Anytime Perceptual Grouping of 2D Features into 3D Basic Shapes
Andreas Richtsfeld, Michael Zillich, Markus Vincze |
ICVS | 3 |
| 2013 | Parallel Deep Learning with Suggestive Activation for Object Category Recognition
Karthik Mahesh Varadarajan, Markus Vincze |
ICVS | 2 |
| 2013 | Automatic in-pipe robot centering from 3D to 2D controller simplificationabstractAfter 50 years the connections between fresh water pipes (800–1200mm diameter) need to be repaired due to aging and dissolution of the filling material. Only in Vienna 3000km of pipes need to be improved, which requires a robotic solution. The main challenge is to accurately align the robot axis with the pipe axis to enable the rotary motion of the maintenance tool. The tool system for cleaning and sealing is mounted on the maintenance unit of the robot consisting of six wheeled-legs. These legs extend to the irregular cast-iron pipe and set the robot structure eccentric to the pipe's center. In order to center the maintenance unit, distance sensors on the legs allow to adapt to the noncircular shape of the pipe. Correcting the leg extension allows to obtain better positioning of the cleaning tool. Luis A. Mateos, Marcos Rodríguez Dominguez, Markus Vincze |
IROS | 3 |
| 2013 | Spontaneous Reorientation for Self-localization
Markus Bader, Markus Vincze |
RoboCup | 2 |
| 2013 | Interactive object modelling based on piecewise planar surface patchesabstractDetecting elements such as planes in 3D is essential to describe objects for applications such as robotics and augmented reality. While plane estimation is well studied, table-top scenes exhibit a large number of planes and methods often lock onto a dominant plane or do not estimate 3D object structure but only homographies of individual planes. In this paper we introduce MDL to the problem of incrementally detecting multiple planar patches in a scene using tracked interest points in image sequences. Planar patches are reconstructed and stored in a keyframe-based graph structure. In case different motions occur, separate object hypotheses are modelled from currently visible patches and patches seen in previous frames. We evaluate our approach on a standard data set published by the Visual Geometry Group at the University of Oxford [24] and on our own data set containing table-top scenes. Results indicate that our approach significantly improves over the state-of-the-art algorithms. Johann Prankl, Michael Zillich, Markus Vincze |
Comput. Vis. Image Underst. | 3 |
| 2013 | Gaussian-weighted Jensen-Shannon divergence as a robust fitness function for multi-model fittingabstractModel fitting is a fundamental component in computer vision for salient data selection, feature extraction and data parameterization. Conventional approaches such as the RANSAC family show limitations when dealing with data containing multiple models, high percentage of outliers or sample selection bias, commonly encountered in computer vision applications. In this paper, we present a novel model evaluation function based on Gaussian-weighted Jensen–Shannon divergence, and integrate into a particle swarm optimization (PSO) framework using ring topology. We avoid two problems from which most regression algorithms suffer, namely the requirements to specify inlier noise scale and the number of models. The novel evaluation method is generic and does not require any estimation of inlier noise. The continuous and meta-heuristic exploration facilitates estimation of each individual model while delivering the number of models automatically. Tests on datasets comprised of inlier noise and a large percentage of outliers (more than 90 % of the data) demonstrate that the proposed framework can efficiently estimate multiple models without prior information. Superior performance in terms of processing time and robustness to inlier noise is also demonstrated with respect to state of the art methods. Kai Zhou 0003, Karthik Mahesh Varadarajan, Michael Zillich, Markus Vincze |
Mach. Vis. Appl. | 4 |
| 2012 | Local 3D Symmetry for Visual Saliency in 2.5D Point Clouds
Ekaterina Potapova, Michael Zillich, Markus Vincze |
ACCV (1) | 3 |
| 2012 | AfNet: The Affordance Network
Karthik Mahesh Varadarajan, Markus Vincze |
ACCV (1) | 2 |
| 2012 | A Global Hypotheses Verification Method for 3D Object Recognition
Aitor Aldoma, Federico Tombari, Luigi Di Stefano, Markus Vincze |
ECCV (3) | 4 |
| 2012 | Attention-driven segmentation of cluttered 3D scenes
Ekaterina Potapova, Michael Zillich, Markus Vincze |
ICPR | 3 |
| 2012 | Implementation of Gestalt principles for object segmentation
Andreas Richtsfeld, Michael Zillich, Markus Vincze |
ICPR | 3 |
| 2012 | Semantic saliency using k-TR theory of visual perception
Karthik Mahesh Varadarajan, Markus Vincze |
ICPR | 2 |
| 2012 | RGB and depth intra-frame Cross-Compression for low bandwidth 3D video
Karthik Mahesh Varadarajan, Kai Zhou 0003, Markus Vincze |
ICPR | 3 |
| 2012 | Robust multiple model estimation with Jensen-Shannon Divergence
Kai Zhou 0003, Karthik Mahesh Varadarajan, Michael Zillich, Markus Vincze |
ICPR | 4 |
| 2012 | Supervised learning of hidden and non-hidden 0-order affordances and detection in real scenesabstractThe ability to perceive possible interactions with the environment is a key capability of task-guided robotic agents. An important subset of possible interactions depends solely on the objects of interest and their position and orientation in the scene. We call these object-based interactions 0-order affordances and divide them among non-hidden and hidden whether the current configuration of an object in the scene renders its affordance directly usable or not. Conversely to other works, we propose that detecting affordances that are not directly perceivable increase the usefulness of robotic agents with manipulation capabilities, so that by appropriate manipulation they can modify the object configuration until the seeked affordance becomes available. In this paper we show how 0-order affordances depending on the geometry of the objects and their pose can be learned using a supervised learning strategy on 3D mesh representations of the objects allowing the use of the whole object geometry. Moreover, we show how the learned affordances can be detected in real scenes obtained with a low-cost depth sensor like the Microsoft Kinect through object recognition and 6D0F pose estimation and present results for both learning on meshes and detection on real scenes to demonstrate the practical application of the presented approach. Aitor Aldoma, Federico Tombari, Markus Vincze |
ICRA | 3 |
| 2012 | 3DNet: Large-scale object class recognition from CAD modelsabstract3D object and object class recognition gained momentum with the arrival of low-cost RGB-D sensors and enables robotics tasks not feasible years ago. Scaling object class recognition to hundreds of classes still requires extensive time and many objects for learning. To overcome the training issue, we introduce a methodology for learning 3D descriptors from synthetic CAD-models and classification of never-before-seen objects at the first glance, where classification rates and speed are suited for robotics tasks. We provide this in 3DNet (3d-net.org), a free resource for object class recognition and 6DOF pose estimation from point cloud data. 3DNet provides a large-scale hierarchical CAD-model databases with increasing numbers of classes and difficulty with 10, 50, 100 and 200 object classes together with evaluation datasets that contain thousands of scenes captured with a RGB-D sensor. 3DNet further provides an open-source framework based on the Point Cloud Library (PCL) for testing new descriptors and benchmarking of state-of-the-art descriptors together with pose estimation procedures to enable robotics tasks such as search and grasping. Walter Wohlkinger, Aitor Aldoma, Radu Bogdan Rusu, Markus Vincze |
ICRA | 4 |
| 2012 | Empty the basket - a shape based learning approach for grasping piles of unknown objectsabstractThis paper presents a novel approach to emptying a basket filled with a pile of objects. Form, size, position, orientation and constellation of the objects are unknown. Additional challenges are to localize the basket and treat it as an obstacle, and to cope with incomplete point cloud data. There are three key contributions. First, we introduce Height Accumulated Features (HAF) which provide an efficient way of calculating grasp related feature values. The second contribution is an extensible machine learning system for binary classification of grasp hypotheses based on raw point cloud data. Finally, a practical heuristic for selection of the most robust grasp hypothesis is introduced. We evaluate our system in experiments where a robot was required to autonomously empty a basket with unknown objects on a pile. Despite the challenging scenarios, our system succeeded each time. David Fischinger, Markus Vincze |
IROS | 2 |
| 2012 | Segmentation of unknown objects in indoor environmentsabstractWe present a framework for segmenting unknown objects in RGB-D images suitable for robotics tasks such as object search, grasping and manipulation. While handling single objects on a table is solved, handling complex scenes poses considerable problems due to clutter and occlusion. After pre-segmentation of the input image based on surface normals, surface patches are estimated using a mixture of planes and NURBS (non-uniform rational B-splines) and model selection is employed to find the best representation for the given data. We then construct a graph from surface patches and relations between pairs of patches and perform graph cut to arrive at object hypotheses segmented from the scene. The energy terms for patch relations are learned from user annotated training data, where support vector machines (SVM) are trained to classify a relation as being indicative of two patches belonging to the same object. We show evaluation of the relations and results on a database of different test sets, demonstrating that the approach can segment objects of various shapes in cluttered table top scenes. Andreas Richtsfeld, Thomas Morwald, Johann Prankl, Michael Zillich, Markus Vincze |
IROS | 5 |
| 2012 | Attention driven grasping for clearing a heap of objectsabstractGeneration of grasps for automated object manipulation in cluttered scenarios presents major challenges for various modules of the pipeline such as 2D/3D visual processing, 3D modeling, grasp hypothesis generation, grasp planning and path planning. In this paper, we present a solution framework for solving a complex instance of the problem - represented by a heap of unknown and unstructured objects in a bounded environment - in our case, a box; with the goal being removing all objects in the box using an attention driven object modeling approach to cognitive grasp planning. The focus of the algorithm delves on Grasping by Components (GBC), with a prioritization scheme derived from scene based attention and attention driven segmentation. In order to overcome the traditional challenge of segmentation performing poorly in cluttered scenes, we employ a novel active segmentation approach suited to our scenario. While the attention module helps prioritize objects in the heap and salient regions, the GBC scheme segments out parts and generates grasp hypotheses for each part. GBC is a very important component of any scalable and holistic grasping system since it abstracts point cloud object data with parametric shapes and no apriori knowledge (such as 3D models) is required. Earlier work in 3D model building (such as CAD based, simple geometries, bounding boxes, Superquadrics etc.) have depended on precise shape and pose recognition as well as exhaustive training to learn or exhaustive searching in grasp space to generate good grasp hypotheses. These methods are not scalable for real-time scenarios, complex shapes and unknown environments - key challenges in robotic grasping. In order to alleviate this concern, we present a novel parametric algorithm to estimate grasp points and approach vectors from the 3D parametric shape model, along with innovative schemes to optimize the computation of the parametric models as well as to refine the generated grasp hypotheses based on the scene information to aid path planning. We present evaluation of our complex grasping pipeline for cluttered heaps through a series of test sequences involving removal of objects from a box, along with evaluations for our attention mechanisms, active segmentation, 3D model fitting optimizations and quality of our grasp hypotheses. Karthik Mahesh Varadarajan, Ekaterina Potapova, Markus Vincze |
IROS | 3 |
| 2012 | AfRob: The affordance network ontology for robotsabstractAfNet, The Affordance Network is an open affordance computing initiative that provides affordance knowledge ontologies for common household articles in terms of affordance features using surface forms termed as afbits (affordance bits). AfNet currently offers 68 base affordance features (25 structural, 10 material, 33 grasp), providing over 200 object category definitions in terms of 4000 afbits. Symbol grounding algorithms for these affordance features enable recognition of objects in visual (RGB-D) data. While AfNet is built as a generic visual knowledge ontology for recognition, it is well suited for deployment on domestic robots. In this paper, we describe AfRob, an extension of AfNet for robotic applications. AfRob builds upon AfNet by imbibing semantic context and mapping for holistic recognition and manipulation of objects in domestic environments. AfRob also offers modules to enable robots to interact and grasp objects through the generation of grasp affordances. The paper also details the inference mechanisms that adapt AfNet for robots in domestic contexts. Results demonstrate the efficiency of the affordance driven approach to holistic visual processing. Karthik Mahesh Varadarajan, Markus Vincze |
IROS | 2 |
| 2012 | Web mining driven object locality knowledge acquisition for efficient robot behaviorabstractAs an important information resource, visual perception has been widely employed for various indoor mobile robots. The common-sense knowledge about object locality (CSOL), e.g. a cup is usually located on the table top rather than on the floor and vice versa for a trash bin, is a very helpful context information for a robotic visual search task. In this paper, we propose an online knowledge acquisition mechanism for discovering CSOL, thereby facilitating a more efficient and robust robotic visual search. The proposed mechanism is able to create conceptual knowledge with the information acquired from the largest and the most diverse medium - the Internet. Experiments using an indoor mobile robot demonstrate the efficiency of our approach as well as reliability of goal-directed robot behaviour. Kai Zhou 0003, Michael Zillich, Hendrik Zender, Markus Vincze |
IROS | 4 |
| 2012 | Theoretic foundations of situating cognitive vision in robots and cognitive systemsabstractCognitive robots are supposed to execute different tasks, adapt to changes in the environment and cope with new situations. They shall operate in real-world settings. Hence, visual perception of the objects in the environment is an essential function. However, on the one hand there are robots using colour as dominant cue to cope with simple objects and on the other hand there are advanced computer vision methods that search through large data bases of images. What is usually missing to a robot are the higher level semantics of its world. To overcome this gap we analyse the specific requirement to the vision system given the context of cognitive robotics. We then propose a new approach to robot vision, which we termed situated vision in order to emphasise the situated and embodied aspects inherent in this domain. We are discussing requirements for such a situated vision system and provide a theory of cognitive functions that seem important to overcome limitations that appear when fixating only on the information that can be found in the visual input. Finally, we are explaining our understanding for the need of a common ontology on which these functions operate. Matthias J. Schlemmer, Markus Vincze |
J. Exp. Theor. Artif. Intell. | 2 |
| 2012 | Seam Following for Automated Industrial Fiber Mat StitchingabstractThis paper presents a method for automatic seam following of two overlapping carbon fiber mats based on laser scans. We introduce a novel approach that combines one existing and two newly developed edge detection methods in a two out of three voting scheme to obtain high edge tracking robustness. The experimental results demonstrate the feasibility of a fully automated, sensor-guided robotic stitching process. The seam can be located within 1.0 mm at a detection rate of 99.3%. Mario Richtsfeld, Markus Vincze |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2011 | Combining Plane Estimation with Shape Detection for Holistic Scene Understanding
Kai Zhou 0003, Andreas Richtsfeld, Karthik Mahesh Varadarajan, Michael Zillich, Markus Vincze |
ACIVS | 5 |
| 2011 | Predicting the unobservable Visual 3D tracking with a probabilistic motion modelabstractVisual tracking of an object can provide a powerful source of feedback information during complex robotic manipulation operations, especially those in which there may be uncertainty about which new object pose may result from a planned manipulative action. At the same time, robotic manipulation can provide a challenging environment for visual tracking, with occlusions of the object by other objects or by the robot itself, and sudden changes in object pose that may be accompanied by motion blur. Recursive filtering techniques use motion models for predictor-corrector tracking, but the simple models typically used often fail to adequately predict the complex motions of manipulated objects. We show how statistical machine learning techniques can be used to train sophisticated motion predictors, which incorporate additional information by being conditioned on the planned manipulative action being executed. We then show how these learned predictors can be used to propagate the particles of a particle filter from one predictor-corrector step to the next, enabling a visual tracking algorithm to maintain plausible hypotheses about the location of an object, even during severe occlusion and other difficult conditions. We demonstrate the approach in the context of robotic push manipulation, where a 5-axis robot arm equipped with a rigid finger applies a series of pushes to an object, while it is tracked by a vision algorithm using a single camera. Thomas Morwald, Marek Sewer Kopicki, Rustam Stolkin, Jeremy L. Wyatt, Sebastian Zurek, Michael Zillich, Markus Vincze |
ICRA | 7 |
| 2011 | Robust single view room structure segmentation in Manhattan-like environments from stereo visionabstractIn this paper we propose a novel approach for the robust segmentation of room structure using Manhattan world assumption i.e. the frequently observed dominance of three mutually orthogonal vanishing directions in man-made environments. First, separate histograms are generated for the Cartesian major axis, i.e. X, Y and Z, on stereo data with an arbitrary roll, pitch and yaw rotation. Using the traditional Markov particle filters and minimal entropy as metric on the histograms, we are able to estimate the camera orientation with respect to orthogonal structure. Once the orientation is estimated we extract a hypotheses of the room structure by exploiting 2D histograms using mean shift clustering techniques as rough estimate for a pre-segmentation of voxels i.e. plane orientation and position. We apply superpixel over segmentation on the colour input to achieve a dense segmentation. The over segmentation and pre-segmented voxels are combined using graph-cuts for a not a-priori known number of final plane segments with a α-expansion graph cut variant proposed by Delong et al. with polynomial runtime. We show the robustness of our approach with respect to noise in real world data. Sven Olufs, Markus Vincze |
ICRA | 2 |
| 2011 | 3D piecewise planar object model for robotics manipulationabstractMan-made environments are abundant with planar surfaces which have attractive properties for robotics manipulation tasks and are a prerequisite for a variety of vision tasks. This work presents automatic on-line 3D object model acquisition assuming a robot to manipulate the object. Objects are represented with piecewise planar surfaces in a spatio-temporal graph. Planes once detected as homographies are tracked and serve as priors in subsequent images. After reconstruction of the planes the 3D motion is analyzed and initial object hypotheses are created. In case planes start moving independently a split event is triggered, the spatio-temporal object graph is traced back and visible planes as well as occluded planes are assigned to the most probable split object. The novelty of this framework is to formalize Multi-body Structure-and-Motion (MSaM), that is, to segment interest point tracks into different rigid objects and compute the multiple-view geometry of each object, with Minimal Description Length (MDL) based on model selection of planes in an incremental manner. Thus, object models are built from planes, which directly can be used for robotic manipulation. Johann Prankl, Michael Zillich, Markus Vincze |
ICRA | 3 |
| 2011 | Learning What Matters: Combining Probabilistic Models of 2D and 3D Saliency Cues
Ekaterina Potapova, Michael Zillich, Markus Vincze |
ICVS | 3 |
| 2011 | Knowledge Representation and Inference for Grasp Affordances
Karthik Mahesh Varadarajan, Markus Vincze |
ICVS | 2 |
| 2011 | Shape-based depth image to 3D model matching and classification with inter-view similarityabstractObject recognition and especially object class recognition is and will be a key capability in home robotics when robots have to tackle manipulation tasks and grasp new objects or just have to search for objects. The goal is to have a robot classify 'never before seen objects' at first occurrence in a single view in a fast and robust manner. The classification task can be seen as a matching problem, finding the most appropriate 3D model and view with respect to a given depth image. We introduce a single-view shape model based classification approach using RGB-D sensors and a novel matching procedure for depth image to 3D model matching leading inherently to object classification. Utilizing the inter-view similarity of the 3D models for enhanced matching, the average precision of our descriptors is increased of up to 15% resulting in high classification accuracy. The presented adaptation of 3D shape descriptors to 2.5D data enables us to calculate the features in real time, directly from the 3D points of the sensor, without any calculation of normals or generating a mesh from it which is typical of state-of-art methods. Furthermore, we introduce a semi-automatic, user-centric approach to utilize the Internet for acquiring the required training data in the form of 3D models which significantly reduces the time for teaching new categories. Walter Wohlkinger, Markus Vincze |
IROS | 2 |
| 2011 | Coherent spatial abstraction and stereo line detection for robotic visual attentionabstractAttention operators based on 2D image cues (such as color, texture) are well known and discussed extensively in the vision literature but are not ideally suited for robotic applications. In such contexts it is the 3D structure of scene elements that makes them interesting or not. We show how a bottom-up exploration mechanism that fuses 2D saliency-based conspicuity with spatial abstraction resulting from the coherent plane estimation and stereo line detection is well suited for typical indoor robotics tasks. This spatial abstraction is performed by a joint probabilistic model which takes the interaction of stereo line detection and 3D supporting plane estimation into consideration. By maximizing the probability of the joint model, our method facilitates reduction of false-positive stereo line detection and refines the estimation of supporting surface simultaneously. Experiments demonstrate that our approach provides more accurate and plausible attention. Kai Zhou 0003, Andreas Richtsfeld, Michael Zillich, Markus Vincze |
IROS | 4 |
| 2011 | Knowing your limits - self-evaluation and prediction in object recognitionabstractAllowing a robot to acquire 3D object models autonomously not only requires robust feature detection and learning methods but also mechanisms for guiding learning and assessing learning progress. In this paper we present probabilistic measures for observed detection success, predicted detection success and the completeness of learned models, where learning is incremental and online. This allows the robot to decide when to add a new keyframe to its view-based object model, where to look next in order to complete the model, predicting the probability of successful object detection given the model trained so far as well as knowing when to stop learning. Michael Zillich, Johann Prankl, Thomas Morwald, Markus Vincze |
IROS | 4 |
| 2011 | Mental state and behavior inference using Mirror Neuron System architecture for traffic/driver monitoringabstractTraffic psychology presents interesting avenues towards the development of Intelligent Transportation Systems (ITS). Analysis of driver state, emotion and behavior are important components of traffic psychology. While being comprehensive in terms of theoretical frameworks, these analyses lack neurobiological computational models for evaluation. In this paper, we develop computational models for driver state and behavior, also known as Mental State Inference (MSI) based on the Mirror Neuron System (MNS) architecture. The integrated system combines neurobiological models with computer vision techniques for traffic monitoring from surveillance video leading to MSI and event recognition. Evaluation of the system is carried out in terms of actual, psychophysical as well as neurobiological criteria on both simulated and real data. Results demonstrate event and mental state recognition convergence within 0.5 normalized time units for the designed event models on synthetic and real data. The model is also robust to perturbations and is aligned to behavior expected at the psychophysical level. Karthik Mahesh Varadarajan, Kai Zhou 0003, Markus Vincze |
Intelligent Vehicles Symposium | 3 |
| 2011 | Room-structure estimation in Manhattan-like environments from dense 2½D range data using minumum entropy and histogramsabstractIn this paper we propose a novel approach for the robust estimation of room structure using Manhattan world assumption i.e. the frequently observed dominance of three mutually orthogonal vanishing directions in man-made environments. First, separate histograms are generated for every major axis, i.e. X, Y and Z, on stereo data with an arbitrary roll, pitch and yaw rotation. These histograms are maintained in the fashion of quadtrees. Using the traditional Markov particle filters and minimal entropy as metric on the histograms, we are able to estimate the camera orientation with respect to orthogonal structure. Once the orientation is estimated we extract hypothesis of the room structure by exploiting 2D histograms, i.e. X/Y, Z/Y, Z/X, using mean shift clustering techniques. Finally, the hypotheses are evaluated with the real data and false hypothesis are pruned. We also show the robustness of our approach with respect to noise in real world data. Sven Olufs, Markus Vincze |
WACV | 2 |
| 2010 | Incremental Model Selection for Detection and Tracking of Planar SurfacesabstractMan-made environments are abundant with planar surfaces which have attractive properties and are a prerequisite for a variety of vision tasks. This paper presents an incremental model selection method to detect piecewise planar surfaces, where planes once detected are tracked and serve as priors in subsequent images. The novelty of this approach is to formalize model selection for plane detection with Minimal Description Length (MDL) in an incremental manner. In each iteration tracked planes and new planes computed from randomly sampled interest points are evaluated, the hypotheses which best explain the scene are retained, and their supporting points are marked so that in the next iteration random sampling is guided to unexplained points. Hence, the remaining finer scene details can be represented. We show in a quantitative evaluation that this new method competes with state of the art algorithms while it is more flexible to incorporate prior knowledge from tracking. Johann Prankl, Michael Zillich, Bastian Leibe, Markus Vincze |
BMVC | 4 |
| 2010 | Real-time depth diffusion for 3D surface reconstructionabstractRange data obtained from conventional stereo-cameras employing dense stereo matching algorithms typically contain a high amount of noise, especially under poor illumination conditions. Furthermore, lack of reliable depth estimates in low-texture regions can result in poor 3D surface reconstruction. Anisotropic diffusion algorithms have been used recently in stereo matching, depth estimation and 3D surface reconstruction. However, these algorithms typically have long execution times, preventing real-time operation on resource constrained systems and robots. Moreover, most of these techniques suffer from excessive smoothing at depth discontinuities resulting in loss of structure, especially in areas where the 2D image does not provide structural cues to guide the depth diffusion. These algorithms are also unsuitable for diffusion of extremely sparse depth data such as in the case of homogenous surfaces. This paper addresses these issues by novel denoising and diffusion techniques. The results presented demonstrate the run-time efficiency and fidelity of reconstructed depth surfaces. Karthik Mahesh Varadarajan, Markus Vincze |
ICIP | 2 |
| 2010 | 3D room modeling and doorway detection from indoor stereo imagery using feature guided piecewise depth diffusionabstractTraditional indoor 3D structural environment modeling algorithms employ schemes such as clustering of dense point clouds for parameterization and identification of the 3D surfaces. RANSAC based plane fitting is one common approach in this regard. Alternatively, extensions to feature based stereo have also been used, mainly focusing on 3D line descriptions, along with techniques such as half-plane detection, real-plane or facade reconstruction, plane sweeping etc. Noise in the range data, especially in low texture regions, accidental line/plane grouping under lack of cues for visibility tests, presence of depth edges or discontinuities that are not visible in the 2D image and difficulties in adaptively estimating metrics for clustering can hamper efficiency of practical systems. In order to counter these issues, we propose a novel framework fusing 2D local and global features such as edges, texture and regions, with geometry information obtained from range data for reliable 3D indoor scene representation. The strength of the approach is derived from the novel depth diffusion and segmentation algorithms resulting in superior surface characterization as opposed to traditional feature based stereo or RANSAC based plane fitting approaches. These algorithms have also been heavily optimized to enable real-time deployments on personal, domestic and rehabilitation robots. Karthik Mahesh Varadarajan, Markus Vincze |
IROS | 2 |
| 2010 | A fast stereo matching algorithm suitable for embedded real-time systems
Martin Humenberger, Christian Zinner, Wilfried Kubinger, Markus Vincze |
Comput. Vis. Image Underst. | 5 |
| 2010 | Model-based 3D object detection
Georg Biegelbauer, Markus Vincze, Walter Wohlkinger |
Mach. Vis. Appl. | 2 |
| 2010 | Efficient borehole detection from single scan data
Markus Vincze, Walter Wohlkinger, Georg Biegelbauer |
Mach. Vis. Appl. | 1 |
| 2009 | Point Cloud Segmentation Based on Radial Reflection
Mario Richtsfeld, Markus Vincze |
CAIP | 2 |
| 2009 | Consistent Interpretation of Image Sequences to Improve Object Models on the Fly
Johann Prankl, Martin Antenreiter, Peter Auer, Markus Vincze |
ICVS | 4 |
| 2009 | Selecting good corners for structure and motion recovery using a time-of-flight cameraabstractIn the robotics and computer vision communities, localization and mapping of an unknown environment is a well studied problem. To tackle this problem in real-time using a single camera, state-of-the-art Simultaneous Localization and Mapping (SLAM) or Structure from Motion (SfM) algorithms can be used. To create the model of the unknown environment, the camera moves and adds to the map from point to point, and assumes that these detected points are unique 3D corners. However, the scene usually contains false 3D corners, lying at e.g. occlusion boundaries. Inserting these points into the map may lead to SLAM failure or to less accurate estimations in SfM. In this work, a corner selection scheme is proposed that exploits the amplitude and depth signals of a Time-of- Flight (ToF) camera. The selection scheme detects false 3D corners based on a 3D cornerness measure. We then prove that the rejection of these corners increases the accuracy with a simulated SfM example and show the results of using our selection scheme with the ToF camera sequences. Peter Gemeiner, Peter Jojic, Markus Vincze |
IROS | 3 |
| 2009 | An efficient area-based observation model for monte-carlo robot LocalizationabstractThe problem of mobile robot self-localization is considered as solved since Thrun's et. al pioneering work using monte-carlo filters for robot Localization (MCL). However, MCL is robust and precise under constraints like completely known environments and the sensor data must contain enough ¿true data¿ as contained in the map. In fact these conditions cannot always be guaranteed, which may results in a poor accuracy of the localization. In this paper we present a area-based observation model that is applied to MCL self-localization. The model is based on the idea of tracking the ground area inside the ¿free space¿ (not occupied cells) of a known map. Experimental data shows that the proposed model improves the robustness and accuracy of laser and stereo vision sensors under certain conditions like incomplete map, limited FOV and limited range of sensing. We also present an efficient approximation of our sensor model based on integral images. Sven Olufs, Markus Vincze |
IROS | 2 |
| 2009 | A simple inexpensive interface for robots using the Nintendo Wii controllerabstractTo have a robot at home might be great fun: it could fetch and carry things. However it remains open how to teach the robot the places it should go to in a manner that is cheap and entertaining for the user. This paper presents an easy-to-use interface that takes the robot on a virtual leash: using the Nintendo Wii remote the user can go towards target places while pointing at the robot. Using the inbuilt infrared camera and accelerometers and a couple of LEDs on the robot, the robot will follow the user. We show how a particle filter and an interacting multiple model (IMM) Kalman can be configured such that simple hand gestures with the Wii make the robot follow the user's intention. The concept has been implemented on a mobile robot developed within the robotshome project. The robot leash interface has been tested with 12 volunteers who are interested in new technology but have never controlled a robot. The result is that most users could within a few minutes show the robot the first three places in a home environment. Given the little cost of the interface (about $ 50) the proposed robot leash is a promising human robot interface. Sven Olufs, Markus Vincze |
IROS | 2 |
| 2009 | Integrated vision system for the semantic interpretation of activities where a person handles objects
Markus Vincze, Michael Zillich, Wolfgang Ponweiser, Václav Hlavác, Jiri Matas, Stepán Obdrzálek, Hilary Buxton, A. Jonathan Howell, Kingsley Sage, Antonis A. Argyros, Christof Eberst, Gerald Umgeher |
Comput. Vis. Image Underst. | 1 |
| 2008 | Clustered multiple generalized expected improvement: A novel infill sampling criterion for surrogate modelsabstractSurrogate model-based optimization is a well-known technique for optimizing expensive black-box functions. By applying this function approximation, the number of real problem evaluations can be reduced because the optimization is performed on the model. In this case two contradictory targets have to be achieved: increasing global model accuracy and exploiting potentially optimal areas. The key to these targets is the criterion for selecting the next point, which is then evaluated on the expensive black-box function - the dasiainfill sampling criterionpsila. Therefore, a novel approach - the dasiaClustered Multiple Generalized Expected Improvementpsila (CMGEI) - is introduced and motivated by an empirical study. Furthermore, experiments benchmarking its performance compared to the state of the art are presented. Wolfgang Ponweiser, Tobias Wagner 0001, Markus Vincze |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | Hippocampal-like categorization of object views: A self-organizing learning approach to vision modeling using stochastic grammar inference and associative memoryabstractThe categorization problem in object recognition is the assignment of semantic categories to objects or parts of objects. Here, the best categorization performance provide mammalian brain functions that motivate to partly mimic such a formation for the application to machine learning and automation. Curiosity and experimental willingness support human recognizing and detecting in relation to still unknown objects. In order to learn an understanding of the topology of such objects, objects are observed from a series of different viewpoints. This results in a collection of object views that afterward can be used by cognitive processes. State-of-the-art systems provide learning structures that are subsequently utilized for object recognition and tracking tasks. However, most of these systems aim at very specific goals in restricted domains, and therefore, it remains no room for learning and understanding of object structures. In the work of this paper, we propose a new robot vision model for neural categorization. This is close to our initial idea of mimicking mammalian brain functions for robot vision. We use the combinatorial solution of our cognitive framework with an embedding of a recently presented stochastic n-gram model, supported by a three-dimensional grammar model on a discrete three-dimensional lattice. Furthermore, we use ant colony optimization heuristics for collecting the transition probabilities of the n-gram model. The proposed solution is exemplified by applying the method to several object view series of 69 polytopes generated out of the five platonic solids by truncation; thus, generating images from 32 viewpoints each, yields an object set of 2208 images. These set is further expanded to four sets perturbed by Gaussian-noise with varying sigma = {0, 0.1, 0.5, 1, 2}. Finally, we show results for selected objects and conclude with an outlook on further work. Peter Michael Goebel, Markus Vincze, Bernard Favre-Bulle |
ETFA | 2 |
| 2008 | Towards detection of orthogonal planes in monocular images of indoor environmentsabstractIn this paper, we describe the components of a novel algorithm for the extraction of dominant orthogonal planar structures from monocular images taken in indoor environments. The basic building block of our approach is the use of vanishing points and vanishing lines imposed by the frequently observed dominance of three mutually orthogonal vanishing directions in man-made world. Vanishing points are found by an improved approach, taking no assumptions on known internal or external camera parameters. The problem of detecting planar patches is attacked using a probabilistic framework, searching for the maximum a posteriori probability (MAP) in a Markov Random Field (MRF). For this, we propose a novel formulation fusing geometric information obtained from vanishing points and features, such as rectangles and partial rectangles, together with a color-homogeneity criteria imposed by an image over-segmentation. The method was evaluated on a set of images exhibiting largely varying characteristics concerning image quality and scene complexity. Experiments show that the method, despite the variations, works in a stable manner and that its performance compares favorably to the state-of-the-art. Branislav Micusík, Horst Wildenauer, Markus Vincze |
ICRA | 3 |
| 2008 | Multiobjective Optimization on a Limited Budget of Evaluations Using Model-Assisted -Metric Selection
Wolfgang Ponweiser, Tobias Wagner 0001, Dirk Biermann, Markus Vincze |
PPSN | 4 |
| 2007 | Efficient Texture Representation Using Multi-scale Regions
Horst Wildenauer, Branislav Micusík, Markus Vincze |
ACCV (1) | 3 |
| 2007 | A Cognitive Modeling Approach for the Semantic Aggregation of Object Prototypes from Geometric Primitives: Toward Understanding Implicit Object Topology
Peter Michael Goebel, Markus Vincze |
ACIVS | 2 |
| 2007 | The Multiple Multi Objective Problem - Definition, Solution and Evaluation
Wolfgang Ponweiser, Markus Vincze |
EMO | 2 |
| 2007 | Optical Seam Following for Automated Robot SewingabstractNowadays, robust and light-weight parts used in the automobile and aeronautics industry are made of carbon fibres. To increase the mechanical toughness of the parts the carbon fibres are stitched in the preforming process using a sewing robot. However, current systems miss high flexibility and rely on manual programming of each part. The main target of this work is to develop an automatic system that autonomously sets the structure strengthening seams. Therefore, a rapid and flexible following of the carbon textile edges is required. Due to the black and reflective carbon fibres a laser-stripe sensor is necessary and the processing of the range data is a challenging task. The paper proposes a real time approach where different edge detection methodologies are combined in a voting scheme to increase the edge tracking robustness. The experimental results demonstrate the feasibility of a fully automated, sensor-guided robotic sewing process. The seam can be located to within 0.65mm at a detection rate of 99.3% for individual scans. Georg Biegelbauer, Mario Richtsfeld, Walter Wohlkinger, Markus Vincze, Manuel Herkt |
ICRA | 4 |
| 2007 | Efficient 3D Object Detection by Fitting Superquadrics to Range Image Data for Robot's Object ManipulationabstractFast detection of objects in a home or office environment is relevant for robotic service and assistance applications. In this work we present the automatic localization of a wide variety of differently shaped objects scanned with a laser range sensor from one view in a cluttered setting. The daily-life objects are modeled using approximated superquadrics, which can be obtained from showing the object or another modeling process. Detection is based on a hierarchical RANSAC search to obtain fast detection results and the voting of sorted quality-of-fit criteria. The probabilistic search starts from low resolution and refines hypotheses at increasingly higher resolution levels. Criteria for object shape and the relationship of object parts together with a ranking procedure and a ranked voting process result in a combined ranking of hypothesis using a minimum number of parameters. Experiments from cluttered table top scenes demonstrate the effectiveness and robustness of the approach, feasible for real world object localization and robot grasp planning. Georg Biegelbauer, Markus Vincze |
ICRA | 2 |
| 2007 | MOVEMENT -Modular Versatile Mobility Enhancement SystemabstractAlthough powered wheelchairs provide a well established solution for severely impaired persons they do not cover all needs regarding mobility of people with impairment. In the course of the EC funded research project MOVEMENT a novel approach for a highly adaptable and modular mobility enhancement system is targeted to cover additional user needs. A system consisting of a robotic platform and several dockable application modules is developed that additionally provides assistance for the driving process itself to the user or even takes over the complete driving autonomously. The project also includes development of new solutions for navigation of mobile robot systems including a "low-cost" sensor system as well as adaptable HMI components. This paper describes the concept and the first prototyping results. Peter Mayer 0002, Georg Edelmayer, Gert Jan Gelderblom, Markus Vincze, Peter Einramhof, Marnix Nuttin, Thomas Fuxreiter, Gernot Kronreif |
ICRA | 4 |
| 2006 | Robust Handling of Multiple Multi-Objective OptimisationsabstractTo enable the selection of methods in any domain, a detailed and context dependent performance description of the methods is required. In cases, where this description has to be generated by an evaluation, the introduction of context can be done cleverly without increasing the number of evaluation runs. However the more detailed performance evaluation delivers several objective entries for one test case. This results in a new type of problem, a multiple multiobjective optimisation. This paper presents a novel genetic algorithm and a novel performance metric that deals with the specific of this type of problem in a robust way. Wolfgang Ponweiser, Markus Vincze |
e-Science | 2 |
| 2006 | 3D Vision-Guided Bore Inspection SystemabstractVision systems are more and more used in the fields of industrial automation and quality control. This paper describes a 3D vision system that guides a robot arm to reliably insert an endoscope into a bore hole for a subsequent bore surface inspection. The already industrial deployed system, developed in the project FibreScope [3], performs a quality control of bores with a diameter ranging from 4 to 50mm and a depth of up to 100mm. This requires a bore detection accuracy of the vision system of "}0.3mm and less than 0.5 in 3D space to modify the robot path for safely inserting the endoscope. The challenge in the design of the vision system was that it must work in process real time achieving the required accuracy. This paper introduces the prototype of the robotic system and the 3D vision system. A performance evaluation of the 3D vision system and its bore localization and detection proofs the need of a vision system to compensate positioning uncertainties for a flexible robotic inspection system. Georg Biegelbauer, Markus Vincze |
ICVS | 2 |
| 2006 | Contextual Coordination in a Cognitive Vision System for Symbolic Activity InterpretationabstractIn this paper we present a vision system that gives a natural language interpretation of activities where a person handles objects. The system integrates low-level image components such as hand and object tracking, detection and recognition, posture and gesture recognition, with highlevel processes such as spatio-temporal object relationship generation and activity reasoning. To achieve near realtime operation, a task-oriented approach focuses processing depending on the situation context. Vision components are dynamically coordinated and their interrelation and region of interest continuously adapted to the interpretation task. The task of putting a CD in a CD-player with one or two hands demonstrates the operation of the system. Markus Vincze, Wolfgang Ponweiser, Michael Zillich |
ICVS | 1 |
| 2005 | Motion and Structure Estimation from Vision and Inertial Sensor Data with High Speed CMOS CameraabstractThis paper presents a system developed in the SmartTracking project to simultaneously estimate the egomotion of a mobile platform and the structure of the environment in which the platform moves. This is required in applications such as robot navigation and augmented reality (AR) to overlay virtual information correctly. The egomotion estimation is achieved by integrating visual and inertial sensor data. The structure estimation is based on the detection of corner features in the environment. From a single known starting position, the system can move into an unknown environment. To enable fast robotics applications the system uses specially developed CMOS cameras that operate at a rate of 2000 image windows per second. The vision and inertial data are fused with an extended Kalman filter. The filter is designed to handle asynchronous input from these two sensors, which typically operate at different and possibly varying rates. Additionally, a bank of filters is used to estimate the quality of structure points and to include them into the structure estimation process. The system is demonstrated on a set-up with known ground truth, such that the motion from a known reference template into a new unknown environment can be shown. Peter Gemeiner, Markus Vincze |
ICRA | 2 |
| 2005 | Finding Tables for Home Service Tasks and Safe Mobile Robot NavigationabstractA persistent difficulty of indoor robot navigation is to reliably locate tables for navigation or service tasks. The predominately used laser range finders have difficulties to detect thin metallic table legs and protruding table surfaces. This papers presents a simple yet robust approach to locate table surfaces using stereo vision in home and office environments. Disparity image data is used to first locate the ground plane with an approach building on the results in [3]. Exploiting the expectation of the floor plane, parallel surfaces are detected. Robustness is achieved by using a combination of local grouping processes and probabilistic estimation to cope with the noisy disparity data and to cope with the typically large number of spurious data points. Tables are detected and located independently of clutter on the table or the type of legs. Experiments in a home environment show detection accuracy of less than two centimeters in estimating the table or chair surface height. Robert Vogl, Markus Vincze, Georg Biegelbauer |
ICRA | 2 |
| 2004 | Decomposition of range images using markov random fieldsabstractThis paper describes a computational model for deriving a decomposition of objects from laser rangefinder data. The process aims to produce a set of parts defined by compactness and smoothness of surface connectivity. Relying on a general decomposition rule, any kind of objects made up of free-form surfaces are partitioned. A robust method to partition the object based on Markov random fields (MRF), which allows to incorporate prior knowledge, is presented. Shape index and curvedness descriptors along with discontinuity and concavity distributions are introduced to classify region labels correctly. In addition, a novel way to classify the shape of a surface is proposed resulting in a better distinction of concave, convex and saddle shapes. To achieve a reliable classification a multiscale method provides a stable estimation of the shape index. Andreas Pichler, Robert B. Fisher, Markus Vincze |
ICIP | 3 |
| 2004 | Multi-rate Fusion with Vision and Inertial SensorsabstractThis work presents a multi-rate fusion model, which exploits the complimentary properties of visual and inertial sensors for egomotion estimation in applications such as robot navigation and augmented reality. The sampling of these two sensors is described with size-varying input and output equations without assumed synchronicity and periodicity of measurements. Data fusion is performed with two different multi-rate (MR) filter models, an extended (EKF) and an unscented Kalman filter (UKF). A complete dynamic model for the 6D-tracking task is given together with a method to calculate the dependencies of the covariance matrices. It is further shown that a centripetal acceleration model and the precise description of quaternion prediction for a constant velocity model highly improve the estimation error for rotary motions. The comparison demonstrates that the MR-UKF provides better estimation results at higher computational costs. Leopoldo Armesto, Stefan Chroust, Markus Vincze, Josep Tornero |
ICRA | 3 |
| 2004 | Sensor based Robotics for Fully Automated Inspection of Bores at Low Volume High Variant PartsabstractBores and internal threads play a critical role as integral parts of connections, bearings and engines, hydraulic and pneumatic systems. Defective bore surfaces lead to increased friction and abrasion, leakage, and reduced stability. Therefore, industry demands ever higher quality of components' bore surfaces. Today industrial automation is limited to high part volumes. Current systems miss a high flexibility or are mainly driven by human interaction and control. Therefore the main target of this work is to develop an automatic system for rapid and flexible 100% surface inspection of bores with diameters from 4 to 50 mm. Further a method for the detection and localization of bores on arbitrary metallic objects has to be developed as well as the image processing for the inspection with a vision system has to be done. The main focus is on easy operating of the system achieved with rapid programming for the user. The paper describes the system, in overview and detail, and shows the promising results of the experiments and a public long-term demonstration (4 days) during an industrial fair. Georg Biegelbauer, Markus Vincze, Helmut Nöhmayer, Christof Eberst |
ICRA | 2 |
| 2003 | Detection of Classes of Features for Automated Robot ProgrammingabstractThis paper presents an approach to detect classes of features that are relevant for automating spray painting. Using knowledge about the painting process a set of elementary geometries is defined, where each elementary geometry is related to a specific painting strategy. Hence all parts and part families containing these elementary geometries can be detected. After detection the paint strokes are automatically generated for robot programming. Specifically we show how free-form surfaces, cavities and rib sections are detected in the range image of the parts. Results of detecting these features on a large variety of parts are presented. Markus Vincze, Andreas Pichler, Georg Biegelbauer |
ICRA | 1 |
| 2003 | Measuring Scene Complexity to Adapt Feature Selection of Model-Based Object Tracking
Minu Ayromlou, Michael Zillich, Wolfgang Ponweiser, Markus Vincze |
ICVS | 4 |
| 2003 | A system to navigate a robot into a ship structure
Markus Vincze, Minu Ayromlou, Carlos Beltrán 0002, Antonios Gasteratos, Simon Hoffgaard, Ole Madsen, Wolfgang Ponweiser, Michael Zillich |
Mach. Vis. Appl. | 1 |
| 2002 | A Method for Automatic Spray Painting of Unknown PartsabstractToday's industrial automation of spray painting is limited to high part volumes and robot trajectories that are programmed by off-line programming and manual teach-in. This paper presents an approach that uses range image data to obtain the geometry of an unknown part and to automatically generate the robot spray painting trajectories. Laser strip range sensors are installed in front of the paint booth to acquire a range image of the part. Utilizing process knowledge (a geometric library containing constraints specific for the painting application) geometric primitives are detected in the range data. From the geometric primitives a normal vector field is generated that enables to extract main faces. The main faces are located in a 3D space and the process knowledge related to each geometric primitive is utilized to obtain the trajectory for the paint gun. Results of painting a car mirror and steering column are given. Andreas Pichler, Markus Vincze, Henrik J. Andersen, Ole Madsen, Kurt Häusler |
ICRA | 2 |
| 2001 | Robust tracking of ellipses at frame rate
Markus Vincze |
Pattern Recognit. | 1 |
| 2000 | Fast Tracking of Ellipses Using Edge-Projected Integration of CuesabstractCommercial applications of ellipse tracking require robustness and real-time capability. The method presented tracks ellipses at field rate using a Pentium PC. Robustness is obtained by integrating gradient and intensity values for the detection of contour edges and by using a RANSAC-like method to find the most likely ellipse. The method adapts to the appearance along the ellipse circumference and effectively separates object from background. Experiments document the capabilities of the approach with real-world examples. Markus Vincze, Minu Ayromlou, Michael Zillich |
ICPR | 1 |
| 2000 | Dynamics and System Performance of Visual ServoingabstractThe components in the control loop of a visual servoing system are analysed with respect to the dynamic performance. The performance measure is the velocity of the target that can be tracked and expressed as image pixel error. The best dynamic performance obtains the visual servoing system when the vision system operates at at the highest possible rate of image acquisition. The size of the tracking window results from image acquisition time. The optimal parallel architecture is independent of the computing power and the control method used. Markus Vincze |
ICRA | 1 |
| 1999 | A Modular Vision Guided System for Tracking 3D Objects in Real-Time
Markus Vincze, Minu Ayromlou, Wilfried Kubinger |
CAIP | 1 |
| 1999 | Optimal Image Processing Architecture for Active Vision
Peter Krautgartner, Markus Vincze |
ICVS | 2 |
| 1999 | An Integrated Framework for Robust Real-Time 3D Object Tracking
Markus Vincze, Minu Ayromlou, Wilfried Kubinger |
ICVS | 1 |
| 1998 | Performance Evaluation of Vision-Based Control TasksabstractThe tracking performance of vision-based control systems is evaluated and the optimal system configuration found. Configurations evaluated are serial or parallel image acquisition and processing, and pipeline processing. The basis of the optimization is the design of an optimal controller independent of the system configuration. Using this controller design a relation between system latency and maximum pixel error is derived. This relation is used to find maximum dynamic performance for the system configurations. The performance measure is the maximum velocity of the target in the image that can be tracked. The final comparison shows that processing in a pipeline obtains highest velocity due to high cycle rate of the system. The parameters for the point of maximum velocity are derived i.e., the optimal number of steps in a pipeline. Peter Krautgartner, Markus Vincze |
ICRA | 2 |
| 1998 | Comments on "On the kinematics of robot heads"abstractIn the above paper by Sharkey et al. (ibid. vol.13 (1997)), a standard for the kinematics of a generalized stereo robot head was presented. In this paper the commenters attempt to clarify some statements of the paper, particularly regarding modeling and calibration. The objective is to provide a simple model to researchers in active vision that complies to all needs. Stephan Spiess, Markus Vincze |
IEEE Trans. Robotics Autom. | 2 |
| 1997 | On optimising tracking performance for visual servoingabstractVisual tracking is a fundamental primitive in advanced sensing tasks, such as vision-guided manipulation or surveillance. We investigate the influence of window size and image tessellation on tracking performance and show that optimal performance for uniform image tesselations is obtained when image sampling time equals image processing time. The performance measures cover velocity, acceleration, and jerk by utilizing different types of feature prediction. We then show that one-dimensional windows can improve performance for specific targets. Finally we show that multi-resolution approaches (log-polar and image pyramid) greatly improve tracking performance. Markus Vincze, Carl F. R. Weiman |
ICRA | 1 |
| 1996 | Optimal window size for visual tracking for uniform CCDsabstractIn this paper it is shown that for any form of visual tracking (the tracking of an object with visual feedback) tracking performance can be optimised by optimising the size of the tracking image window. The optimality criterion is the maximal change in predicted feature position, which is proportional to instantaneous acceleration, the biggest problem in tracking. The result is that optimal window size is reached when the time to process the image data equals the time of sampling the image. This result is valid for all algorithms whose calculation time is proportional to the area of the window. Linear windows can improve performance if the feature can be found within a linear window. Markus Vincze |
ICPR | 1 |
| 1995 | Camera System to Detect the Orientation of a Corner Cube in Real TimeabstractThe measurement of the position and orientation of a robot end effector is the most critical issue in calibrating robots. Systems to measure the position are well known but systems to measure the orientation at the same time and with a useful accuracy are not available. A method to do this was developed at the Institute of Flexible Automation (INFA). The new system (LuxWess) bases on a tracking system, a laser beam and image processing techniques. This paper describes the real time image processing system required to record the profile of the laser beam with a sample rate of more than 10000 frames per second and to find the geometry of the corner cube. The detection of the information needed to calculate the orientation of the corner cube with an accuracy of 10 arc seconds is explained. A method to increase the accuracy up to 2 arc seconds is introduced. Karl M. Filz, Markus Vincze, Johann P. Prenninger |
ICRA | 2 |
| 1993 | An external 6D-sensor for industrial robotsabstractThis paper presents a dynamic six-degree-of-freedom measurement system for industrial robots. The measurement system uses the beam of a laser interferometer and a mirror mounted on a cardan joint to follow a retroreflector which is mounted to the robot's end-effector. The position of the end-effector is measured with high accuracy and is computed in real time. The orientation is determined by analyzing the intensity profile of the reflected laser beam. Since the measurement system must be capable of tracking the movements of the end-effector at the maximum speed of the robot, the requirements of the tracking controller regarding dynamic operation are very high. Thus a target tracking controller has been developed and is presented in detail. Helmut Gander, Markus Vincze, Johann P. Prenninger |
IROS | 2 |