Tim Patten

dblp:148/2222 · also Timothy Patten · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
11since 2021 · last 2023
0000-0003-1139-9451ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 6 since 2021Systems, architecture and hardware · 10 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 3D-DAT: 3D-Dataset Annotation Toolkit for Robotic Vision
abstract
Robots operating in the real world are expected to detect, classify, segment, and estimate the pose of objects to accomplish their task. Modern approaches using deep learning not only require large volumes of data but also pixel-accurate annotations in order to evaluate the performance and therefore safety of these algorithms. At present, publicly available tools for annotating data are scarce and those that are available rely on depth sensors, which excludes their use for transparent, metallic, and general non-Lambertian objects. To address this issue, we present a novel method for creating valuable datasets that can be used in these more difficult cases. Our key contribution is a purely RGB-based scene-level annotation approach that uses a neural radiance field-based method to automatically align objects. A set of user studies demonstrates the accuracy and speed of our approach over a purely manual or depth sensor assisted pipeline. We provide an open-source implementation of each component and a ROS-based recorder for capturing data with a eye-in-hand robot system. Code will be made available at https://github.com/markus-suchi/3D-DAT.
Markus Suchi, Bernhard Neuberger, Amanzhol Salykov, Jean-Baptiste Weibel, Tim Patten, Markus Vincze
ICRA5
2023 TrackAgent: 6D Object Tracking via Reinforcement Learning
Konstantin Röhrl, Dominik Bauer, Tim Patten, Markus Vincze
ICVS3
2023 Skirting Line Estimation Using Sparse to Dense Deformation
abstract
Automating the process of fleece contaminant removal has the potential to drastically improve the quality of wool leaving the farm gate. Towards this goal, we present a method to automatically extract skirting lines, i.e., the separations between clean and contaminated wool of a fleece using RGB images. We propose a learning-based sparse-to-dense approach for estimating the non-rigid deformation of fleeces in order to estimate the skirting lines. Our method is bootstrapped from a set of sparse inlier feature correspon-dences, which are heavily filtered through a set of strict criteria. The inlier correspondences are then greedily expanded by adding correspondences from a denser set through a filtering process. This process is based on a learning approach that takes as inputs the pixel similarity and the consistency with their inlier neighbours. Each greedy iteration is initialised with a non-rigid deformation using as-rigid-as-possible as a prior to the filtering process. The proposed method outperforms both a rigid deformation baseline and optic flow deep learning approach, as evidenced by the quantitative evaluation of pixel location error in controlled experiments. To further prove its practicality, we demonstrate qualitative results comparing the predicted skirting line from various methods on images of skirted fleeces collected from several wool sheds.
Daniel Perez Banuelos, Raphael Falque, Tim Patten, Alen Alempijevic
IROS3
2023 COPE: End-to-end trainable Constant Runtime Object Pose Estimation
abstract
State-of-the-art object pose estimation handles multiple instances in a test image by using multi-model formulations: detection as a first stage and then separately trained networks per object for 2D-3D geometric correspondence prediction as a second stage. Poses are subsequently estimated using the Perspective-n-Points algorithm at runtime. Unfortunately, multi-model formulations are slow and do not scale well with the number of object instances involved. Recent approaches show that direct 6D object pose estimation is feasible when derived from the aforementioned geometric correspondences. We present an approach that learns an intermediate geometric representation of multiple objects to directly regress 6D poses of all instances in a test image. The inherent end-to-end trainability overcomes the requirement of separately processing individual object instances. By calculating the mutual Intersection-over-Unions, pose hypotheses are clustered into distinct instances, which achieves negligible runtime overhead with respect to the number of object instances. Results on multiple challenging standard datasets show that the pose estimation performance is superior to single-model state-of-the-art approaches despite being more than ~35 times faster. We additionally provide an analysis showing real-time applicability (> 24 fps) for images where more than 90 object instances are present. Further results show the advantage of supervising geometric correspondence-based object pose estimation with the 6D pose.
Stefan Thalhammer, Tim Patten, Markus Vincze
WACV2
2022 GigaDepth: Learning Depth from Structured Light with Branching Neural Networks
Simon Schreiberhuber, Jean-Baptiste Weibel, Tim Patten, Markus Vincze
ECCV (33)3
2022 A New VR Kitchen Environment for Recording Well Annotated Object Interaction Tasks
abstract
This paper presents the Virtual Annotated Cooking Environment (VACE), a new open-source virtual reality dataset (https://sites.google.com/view/vacedataset) and simulator (https://github.com/michaelkoller/vacesimulator) for object inter-action tasks in a rich kitchen environment. We use the Unity-based VR simulator to create thoroughly annotated video se-quences of a virtual human avatar performing food preparation activities. Based on the MPII Cooking 2 dataset, it enables the recreation of recipes for meals such as sandwiches, pizzas, fruit salads and smaller activity sequences such as cutting vegetables. For complex recipes, multiple samples are present, following different orderings of valid partially ordered plans. The dataset includes an RGB and depth camera view, bounding boxes, object masks segmentation, human joint poses and object poses, as well as ground truth interaction data in the form of temporally labeled semantic predicates (holding, on, in, colliding, moving, cutting). In our effort to make the simulator accessible as an open-source tool, researchers are able to expand the setting and annotation to create additional data samples.
Michael Koller 0002, Tim Patten, Markus Vincze
HRI2
2022 SporeAgent: Reinforced Scene-level Plausibility for Object Pose Refinement
abstract
Observational noise, inaccurate segmentation and ambiguity due to symmetry and occlusion lead to inaccurate object pose estimates. While depth- and RGB-based pose refinement approaches increase the accuracy of the resulting pose estimates, they are susceptible to ambiguity in the observation as they consider visual alignment. We propose to leverage the fact that we often observe static, rigid scenes. Thus, the objects therein need to be under physically plausible poses. We show that considering plausibility reduces ambiguity and, in consequence, allows poses to be more accurately predicted in cluttered environments. To this end, we extend a recent RL-based registration approach towards iterative refinement of object poses. Experiments on the LINEMOD and YCB-VIDEO datasets demonstrate the state-of-the-art performance of our depth-based refinement approach. Code is available at github.com/dornik/sporeagent.
Dominik Bauer, Tim Patten, Markus Vincze
WACV2
2021 ReAgent: Point Cloud Registration Using Imitation and Reinforcement Learning
abstract
Point cloud registration is a common step in many 3D computer vision tasks such as object pose estimation, where a 3D model is aligned to an observation. Classical registration methods generalize well to novel domains but fail when given a noisy observation or a bad initialization. Learning-based methods, in contrast, are more robust but lack in generalization capacity. We propose to consider iterative point cloud registration as a reinforcement learning task and, to this end, present a novel registration agent (ReAgent). We employ imitation learning to initialize its discrete registration policy based on a steady expert policy. Integration with policy optimization, based on our proposed alignment reward, further improves the agent’s registration performance. We compare our approach to classical and learning-based registration methods on both ModelNet40 (synthetic) and ScanObjectNN (real data) and show that our ReAgent achieves state-of-the-art accuracy. The lightweight architecture of the agent, moreover, enables reduced inference time as compared to related approaches. Code is available at github.com/dornik/reagent.
Dominik Bauer, Tim Patten, Markus Vincze
CVPR2
2021 PyraPose: Feature Pyramids for Fast and Accurate Object Pose Estimation under Domain Shift
abstract
Object pose estimation enables robots to understand and interact with their environments. Training with synthetic data is necessary in order to adapt to novel situations. Unfortunately, pose estimation under domain shift, i.e., training on synthetic data and testing in the real world, is challenging. Deep learning-based approaches currently perform best when using encoder-decoder networks but typically do not generalize to new scenarios with different scene characteristics. We argue that patch-based approaches, instead of encoder-decoder networks, are more suited for synthetic-to-real transfer because local to global object information is better represented. To that end, we present a novel approach based on a specialized feature pyramid network to compute multi-scale features for creating pose hypotheses on different feature map resolutions in parallel. Our single-shot pose estimation approach is evaluated on multiple standard datasets and outperforms the state of the art by up to ∼35 %. We also perform grasping experiments in the real world to demonstrate the advantage of using synthetic data to generalize to novel environments.
Stefan Thalhammer, Markus Leitner, Tim Patten, Markus Vincze
ICRA3
2021 Learning Image-Based Contaminant Detection in Wool Fleece from Noisy Annotations
Tim Patten, Alen Alempijevic, Robert Fitch
ICVS1
2021 Object Learning for 6D Pose Estimation and Grasping from RGB-D Videos of In-hand Manipulation
abstract
Object models are highly useful for robots as they enable tasks such as detection, pose estimation and manipulation. However, models are not always easily available, especially in real-world domains of operation such as peoples’ homes. This work presents a pipeline to generate high-quality object reconstructions from human in-hand manipulation to alleviate the necessity of specialised or expensive hardware. Missing data, due to occlusion or unseen sides, is explicitly handled by incorporating shape completion. We demonstrate the usability of the reconstructions by applying a model-based as well as a CNN-based object pose estimator that is trained on synthetic images by employing state-of-the-art texture synthesis. Using our pipeline to cheaply generate object models and synthetic RGB images for training, we achieve competitive performance compared to baselines that require an elaborate set-up to construct models or large amounts of annotated data. Object grasping is also enabled by learning with the reconstructions in simulation, then executing with a real robot. These evaluations show that our reconstructions are comparable to those made under near-perfect conditions and enable 6D object pose estimation as well as real-world grasping.
Tim Patten, Kiru Park, Markus Leitner, Kevin Wolfram, Markus Vincze
IROS1
2020 Neural Object Learning for 6D Pose Estimation Using a Few Cluttered Images
Kiru Park, Tim Patten, Markus Vincze
ECCV (4)2
2020 Robust and Efficient Object Change Detection by Combining Global Semantic Information and Local Geometric Verification
abstract
Identifying new, moved or missing objects is an important capability for robot tasks such as surveillance or maintaining order in homes, offices and industrial settings. However, current approaches do not distinguish between novel objects or simple scene readjustments nor do they sufficiently deal with localization error and sensor noise. To overcome these limitations, we combine the strengths of global and local methods for efficient detection of novel objects in 3D reconstructions of indoor environments. Global structure, determined from 3D semantic information, is exploited to establish object candidates. These are then locally verified by comparing isolated geometry to a reference reconstruction provided by the task. We evaluate our approach on a novel dataset containing different types of rooms with 31 scenes and 260 annotated objects. Experiments show that our proposed approach significantly outperforms baseline methods.
Edith Langer, Tim Patten, Markus Vincze
IROS2
2019 SyDPose: Object Detection and Pose Estimation in Cluttered Real-World Depth Images Trained using Only Synthetic Data
abstract
Object pose estimation is an important problem in robotics because it supports scene understanding and enables subsequent grasping and manipulation. Many methods, including modern deep learning approaches, exploit known object models, however, in industry these are difficult and expensive to obtain. 3D CAD models, on the other hand, are often readily available. Consequently, training a deep architecture for pose estimation exclusively from CAD models leads to a considerable decrease of the data creation effort. While this has been shown to work well for feature-and template-based approaches, real-world data is still required for pose estimation in clutter using deep learning. We use synthetically created depth data with domain-relevant background randomized noise heuristics to train an end-to-end, multi-task network, for pose estimation. We simultaneously detect, classify and estimate the poses of texture-less objects in cluttered real-world depth images of an arbitrary amount of objects. We present the results of our experiments with the LineMOD and the Occlusion dataset.
Stefan Thalhammer, Tim Patten, Markus Vincze
3DV2
2019 Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation
abstract
Estimating the 6D pose of objects using only RGB images remains challenging because of problems such as occlusion and symmetries. It is also difficult to construct 3D models with precise texture without expert knowledge or specialized scanning devices. To address these problems, we propose a novel pose estimation method, Pix2Pose, that predicts the 3D coordinates of each object pixel without textured models. An auto-encoder architecture is designed to estimate the 3D coordinates and expected errors per pixel. These pixel-wise predictions are then used in multiple stages to form 2D-3D correspondences to directly compute poses with the PnP algorithm with RANSAC iterations. Our method is robust to occlusion by leveraging recent achievements in generative adversarial training to precisely recover occluded parts. Furthermore, a novel loss function, the transformer loss, is proposed to handle symmetric objects by guiding predictions to the closest symmetric pose. Evaluations on three different benchmark datasets containing symmetric and occluded objects show our method outperforms the state of the art using only RGB images.
Kiru Park, Tim Patten, Markus Vincze
ICCV2
2019 Multi-Task Template Matching for Object Detection, Segmentation and Pose Estimation Using Depth Images
abstract
Template matching has been shown to accurately estimate the pose of a new object given a limited number of samples. However, pose estimation of occluded objects is still challenging. Furthermore, many robot application domains encounter texture-less objects for which depth images are more suitable than color images. In this paper, we propose a novel framework, Multi-Task Template Matching (MTTM), that finds the nearest template of a target object from a depth image while predicting segmentation masks and a pose transformation between the template and a detected object in the scene using the same feature map of the object region. The proposed feature comparison network computes segmentation masks and pose predictions by comparing feature maps of templates and cropped features of a scene. The segmentation result from this network improves the robustness of the pose estimation by excluding points that do not belong to the object. Experimental results show that MTTM outperforms baseline methods for segmentation and pose estimation of occluded objects despite using only depth images.
Kiru Park, Tim Patten, Johann Prankl, Markus Vincze
ICRA2
2019 ScalableFusion: High-resolution Mesh-based Real-time 3D Reconstruction
abstract
Dense 3D reconstructions generate globally consistent data of the environment suitable for many robot applications. Current RGB-D based reconstructions, however, only maintain the color resolution equal to the depth resolution of the used sensor. This firmly limits the precision and realism of the generated reconstructions. In this paper we present a real-time approach for creating and maintaining a surface reconstruction in as high as possible geometrical fidelity with full sensor resolution for its colorization (or surface texture). A multi-scale memory management process and a Level of Detail scheme enable equally detailed reconstructions to be generated at small scales, such as objects, as well as large scales, such as rooms or buildings. We showcase the benefit of this novel pipeline with a PrimeSense RGB-D camera as well as combining the depth channel of this camera with a high resolution global shutter camera. Further experiments show that our memory management approach allows us to scale up to larger domains that are not achievable with current state-of-the-art methods.
Simon Schreiberhuber, Johann Prankl, Tim Patten, Markus Vincze
ICRA3
2019 EasyLabel: A Semi-Automatic Pixel-wise Object Annotation Tool for Creating Robotic RGB-D Datasets
abstract
Developing robot perception systems for recognizing objects in the real world requires computer vision algorithms to be carefully scrutinized with respect to the expected operating domain. This demands large quantities of ground truth data to rigorously evaluate the performance of algorithms. This paper presents the EasyLabel tool for easily acquiring high-quality ground truth annotation of objects at pixel-level in densely cluttered scenes. In a semi-automatic process, complex scenes are incrementally built and EasyLabel exploits depth changes to extract precise object masks at each step. We use this tool to generate the Object Cluttered Indoor Dataset (OCID) that captures diverse settings of objects, background, context, sensor to scene distance, viewpoint angle and lighting conditions. OCID is used to perform a systematic comparison of existing object segmentation methods. The baseline comparison supports the need for pixel- and object-wise annotation to progress robot vision towards realistic applications. This insight reveals the usefulness of EasyLabel and OCID to better understand the challenges that robots face in the real world.
Markus Suchi, Tim Patten, David Fischinger, Markus Vincze
ICRA2
2019 Robust 3D Object Classification by Combining Point Pair Features and Graph Convolution
abstract
Object classification is an important capability for robots as it provides vital semantic information that underpin most practical high-level tasks. Classic handcrafted features, such as point pair features, have demonstrated their robustness for this task. Combining these features with modern deep learning methods provide discriminative features that are rotation invariant and robust to various sources of noise. In this work, we aim to improve the descriptiveness of point pair features while retaining their robustness. We propose a method to achieve more structured sampling of pairs and combine this information through the use of graph convolutional networks. We introduce a novel attention model based on a repeatable local reference frame. Experiments show that our approach significantly improves the state of the art for object classification on large scale reconstruction such as the Stanford 3D indoor dataset and ScanNet and obtains competitive accuracy on the artificial dataset ModelNet.
Jean-Baptiste Weibel, Tim Patten, Markus Vincze
ICRA2
2019 Leveraging Symmetries to Improve Object Detection and Pose Estimation from Range Data
Sergey V. Alexandrov, Tim Patten, Markus Vincze
ICVS2
2019 Monte Carlo Tree Search on Directed Acyclic Graphs for Object Pose Verification
Dominik Bauer, Tim Patten, Markus Vincze
ICVS2
2018 Action Selection for Interactive Object Segmentation in Clutter
abstract
Robots operating in human environments are often required to recognise, grasp and manipulate objects. Identifying the locations of objects amongst their complex surroundings is therefore an important capability. However, when environments are unstructured and cluttered, as is typical for indoor human environments, reliable and accurate object segmentation is not always possible because the scene representation is often incomplete or ambiguous. We overcome the limitations of static object segmentation by enabling a robot to directly interact with the scene with non-prehensile actions. Our method does not rely on object models to infer object existence. Rather, interaction induces scene motion and this provides an additional clue for associating observed parts to the same object. We use a probabilistic segmentation framework in order to identify segmentation uncertainty. This uncertainty is then used to guide a robot while it manipulates the scene. Our probabilistic segmentation approach recursively updates the segmentation given the motion cues and the segmentation is monitored during interaction, thus providing online feedback. Experiments performed with RGB-D data show that the additional source of information from motion enables more certain object segmentation that was otherwise ambiguous. We then show that our interaction approach based on segmentation uncertainty maintains higher quality segmentation than competing methods with increasing clutter.
Tim Patten, Michael Zillich, Markus Vincze
IROS1
2016 Decentralised Monte Carlo Tree Search for Active Perception
Graeme Best, Oliver M. Cliff, Tim Patten, Ramgopal R. Mettu, Robert Fitch
WAFR3