Dominik Bauer

dblp:00/7183 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Real2Code: Reconstruct Articulated Objects via Code Generation
abstract
We present Real2Code, a novel approach to reconstructing articulated objects via code generation. Given visual observations of an object, we first reconstruct its part geometry using image segmentation and shape completion. We represent these object parts with oriented bounding boxes, from which a fine-tuned large language model (LLM) predicts joint articulation as code. By leveraging pre-trained vision and language models, our approach scales elegantly with the number of articulated parts, and generalizes from synthetic training data to real world objects in unstructured environments. Experimental results demonstrate that Real2Code significantly outperforms the previous state-of-the-art in terms of reconstruction accuracy, and is the first approach to extrapolate beyond objects' structural complexity in the training set, as we show for objects with up to 10 articulated parts. When incorporated with a stereo reconstruction model, Real2Code moreover generalizes to real-world objects, given only a handful of multi-view RGB images and without the need for depth or camera information.
Zhao Mandi, Yijia Weng, Dominik Bauer, Shuran Song
ICLR3
2024 DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects
Dominik Bauer, Zhenjia Xu, Shuran Song
ECCV (41)1
2024 Challenges for Monocular 6-D Object Pose Estimation in Robotics
abstract
Object pose estimation is a core perception task that enables, for example, object manipulation and scene understanding. The widely available, inexpensive, and high-resolution RGB sensors and CNNs that allow for fast inference make monocular approaches especially well-suited for robotics applications. We observe that previous surveys establish the state of the art for varying modalities, single- and multiview settings, and datasets and metrics that consider a multitude of applications. We argue, however, that those works' broad scope hinders the identification of open challenges that are specific to monocular approaches and the derivation of promising future challenges for their application in robotics. By providing a unified view on recent publications from both robotics and computer vision, we find that occlusion handling, pose representations, and formalizing and improving category-level pose estimation are still fundamental challenges that are highly relevant for robotics. Moreover, to further improve robotic performance, large object sets, novel objects, refractive materials, and uncertainty estimates are central and largely unsolved open challenges. In order to address them, ontological reasoning, deformability handling, scene-level reasoning, realistic datasets, and the ecological footprint of algorithms need to be improved.
Stefan Thalhammer, Dominik Bauer, Peter Hönig, Jean-Baptiste Weibel, José García Rodríguez 0001, Markus Vincze
IEEE Trans. Robotics2
2023 TrackAgent: 6D Object Tracking via Reinforcement Learning
Konstantin Röhrl, Dominik Bauer, Tim Patten, Markus Vincze
ICVS2
2022 Learning to Navigate by Pushing
abstract
In this work, we investigate a form of dynamic contact-rich locomotion in which a robot pushes off from obstacles in order to move through its environment. We present a reflex-based approach that switches between optimized hand-crafted reflex controllers and produces smooth and predictable motions. In contrast to previous work, our approach does not rely on periodic movements, complex models of robot and contact dynamics, or extensive hand tuning. We demonstrate the effectiveness of our approach and evaluate its performance compared to a standard model-free RL algorithm. We identify continuous clusters of similar behaviours, which allows us to successfully transfer different push-off motions directly from simulation to a physical robot without further retraining.
Cornelia Bauer, Dominik Bauer, Alisa Allaire, Christopher G. Atkeson, Nancy S. Pollard
ICRA2
2022 Contact Transfer: A Direct, User-Driven Method for Human to Robot Transfer of Grasps and Manipulations
abstract
We present a novel method for the direct transfer of grasps and manipulations between objects and hands through utilization of contact areas. Our method fully preserves contact shapes, and in contrast to existing techniques, is not dependent on grasp families, requires no model training or grasp sampling, makes no assumptions about manipulator morphology or kinematics, and allows user control over both transfer parameters and solution optimization. Despite these accommodations, we show that our method is capable of synthesizing kinematically-feasible whole hand poses in seconds even for poor initializations or hard-to-reach contacts. We additionally highlight the method's benefits in both response to design alterations as well as fast approximation over in-hand manipulation sequences. Finally, we demonstrate a solution generated by our method on a physical, custom-designed prosthetic hand.
Arjun Lakshmipathy, Dominik Bauer, Cornelia Bauer, Nancy S. Pollard
ICRA2
2022 SporeAgent: Reinforced Scene-level Plausibility for Object Pose Refinement
abstract
Observational noise, inaccurate segmentation and ambiguity due to symmetry and occlusion lead to inaccurate object pose estimates. While depth- and RGB-based pose refinement approaches increase the accuracy of the resulting pose estimates, they are susceptible to ambiguity in the observation as they consider visual alignment. We propose to leverage the fact that we often observe static, rigid scenes. Thus, the objects therein need to be under physically plausible poses. We show that considering plausibility reduces ambiguity and, in consequence, allows poses to be more accurately predicted in cluttered environments. To this end, we extend a recent RL-based registration approach towards iterative refinement of object poses. Experiments on the LINEMOD and YCB-VIDEO datasets demonstrate the state-of-the-art performance of our depth-based refinement approach. Code is available at github.com/dornik/sporeagent.
Dominik Bauer, Tim Patten, Markus Vincze
WACV1
2021 ReAgent: Point Cloud Registration Using Imitation and Reinforcement Learning
abstract
Point cloud registration is a common step in many 3D computer vision tasks such as object pose estimation, where a 3D model is aligned to an observation. Classical registration methods generalize well to novel domains but fail when given a noisy observation or a bad initialization. Learning-based methods, in contrast, are more robust but lack in generalization capacity. We propose to consider iterative point cloud registration as a reinforcement learning task and, to this end, present a novel registration agent (ReAgent). We employ imitation learning to initialize its discrete registration policy based on a steady expert policy. Integration with policy optimization, based on our proposed alignment reward, further improves the agent’s registration performance. We compare our approach to classical and learning-based registration methods on both ModelNet40 (synthetic) and ScanObjectNN (real data) and show that our ReAgent achieves state-of-the-art accuracy. The lightweight architecture of the agent, moreover, enables reduced inference time as compared to related approaches. Code is available at github.com/dornik/reagent.
Dominik Bauer, Tim Patten, Markus Vincze
CVPR1
2021 Contact Tracing: A Low Cost Reconstruction Framework for Surface Contact Interpolation
abstract
We present a novel, low cost framework for reconstructing surface contact movements during in-hand manipulations. Unlike many existing methods focused on hand pose tracking, ours models the behavior of contact patches, and by doing so is the first to obtain detailed contact tracking estimates for multi-contact manipulations. Our framework is highly accessible, requiring only low cost, readily available paint materials, a single RGBD camera, and a simple, deterministic interpolation algorithm. Despite its simplicity, we demonstrate the framework’s effectiveness over the course of several manipulations on three common household items. Finally, we demonstrate the use of a generated contact time series in manipulation learning for a simulated robot hand.
Arjun Lakshmipathy, Dominik Bauer, Nancy S. Pollard
IROS2
2019 A Pilot Study on Determining the Relation Between Gaze Aversion and Interaction Experience
abstract
Previous work in HHI and HRI demonstrates the impact of gaze on the human interaction experience (IE). In this paper, we discuss an experimental design that should enable measuring the influence of the gaze aversion ratio (GAR) on the users' IE with a social robot. We assume gaze behavior studied in HHI to be restrictive for autonomous social robots, limiting the time available to robots for perception tasks besides HRI. Our goal is to determine if a deviation from human gaze behavior is accepted by users in HRI. With an in-between experimental design we evaluate the effect of varied GAR on IE and behavioral measures. A pilot study with 9 participants suggests that averting the gaze for longer time spans is favorable.
Michael Koller 0002, Dominik Bauer, Jesse de Pagter, Guglielmo Papagni, Markus Vincze
HRI2
2019 Monte Carlo Tree Search on Directed Acyclic Graphs for Object Pose Verification
Dominik Bauer, Tim Patten, Markus Vincze
ICVS1
2018 A VR-based user study on the effects of vision impairments on recognition distances of escape-route signs in buildings
abstract
In workplaces or publicly accessible buildings, escape routes are signposted according to official norms or international standards that specify distances, angles and areas of interest for the positioning of escape-route signs. In homes for the elderly, in which the residents commonly have degraded mobility and suffer from vision impairments caused by age or eye diseases, the specifications of current norms and standards may be insufficient. Quantifying the effect of symptoms of vision impairments like reduced visual acuity on recognition distances is challenging, as it is cumbersome to find a large number of user study participants who suffer from exactly the same form of vision impairments. Hence, we propose a new methodology for such user studies: By conducting a user study in virtual reality (VR), we are able to use participants with normal or corrected sight and simulate vision impairments graphically. The use of standardized medical eyesight tests in VR allows us to calibrate the visual acuity of all our participants to the same level, taking their respective visual acuity into account. Since we primarily focus on homes for the elderly, we accounted for their often limited mobility by implementing a wheelchair simulation for our VR application.
Katharina Krösl, Dominik Bauer, Michael Schwärzler, Henry Fuchs, Georg Suter, Michael Wimmer 0001
Vis. Comput.2