EDBT 2026 Demo / reviewers in the wild / expert
Stefan Thalhammer
dblp:28/1206
· DBLP profile ↗
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-0008-430XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCOPE: Semantic Cross-Attention Conditioning for Category-Level Object Pose Estimation
Peter Hönig, Jean-Baptiste Weibel, Stefan Thalhammer, Matthias Hirschmanner, Markus Vincze, Andreas Holzinger |
Image Vis. Comput. | 3 |
| 2025 | ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual ServoingabstractVisual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted features and thus are universally applicable without task-specific training, they often struggle with occlusions and environmental variations, whereas learning-based approaches improve robustness but typically require extensive training. We present a visual servoing approach that leverages pretrained vision transformers for semantic feature extraction, combining the advantages of both paradigms while also being able to generalize beyond the provided sample. Our approach achieves full convergence in unperturbed scenarios and surpasses classical image-based visual servoing by up to 31.2% relative improvement in perturbed scenarios. Even the convergence rates of learning-based methods are matched despite requiring no task-or object-specific training. Real-world evaluations confirm robust performance in end-effector positioning, industrial box manipulation, and grasping of unseen objects using only a reference from the same category. Our code and simulation environment are available at: https://alessandroscherl.github.io/ViT-VS/ Alessandro Scherl, Stefan Thalhammer, Bernhard Neuberger, Wilfried Wöber, José García Rodríguez 0001 |
IROS | 2 |
| 2025 | Shape-Biased Texture Agnostic Representations for Improved Textureless and Metallic Object Detection and 6D Pose EstimationabstractRecent advances in machine learning have greatly benefited object detection and 6D pose estimation. However, textureless and metallic objects still pose a significant challenge due to few visual cues and the texture bias of CNNs. To address this issue, we propose a strategy for inducing a shape bias to CNN training. In particular, by randomizing textures applied to object surfaces during data rendering, we create training data without consistent textural cues. This methodology allows for seamless integration into existing data rendering engines, and results in negligible computational overhead for data rendering and network training. Our findings demonstrate that the shape bias we induce via randomized texturing, improves over existing approaches using style transfer. We evaluate with five detectors and two pose estimators. For three object detectors and for pose estimation in general, estimation accuracy improves for textureless and metallic objects. Additionally we show that our approach increases the pose estimation accuracy in the presence of image noise and strong illumination changes. Code available at https://github.com/hoenigpeter/randomized_texturing. Peter Hönig, Stefan Thalhammer, Jean-Baptiste Weibel, Matthias Hirschmanner, Markus Vincze |
WACV | 2 |
| 2024 | ZS6D: Zero-shot 6D Object Pose Estimation using Vision TransformersabstractAs robotic systems increasingly encounter complex and unconstrained real-world scenarios, there is a demand to recognize diverse objects. The state-of-the-art 6D object pose estimation methods rely on object-specific training and therefore do not generalize to unseen objects. Recent novel object pose estimation methods are solving this issue using task-specific fine-tuned CNNs for deep template matching. This adaptation for pose estimation still requires expensive data rendering and training procedures. MegaPose for example is trained on a dataset consisting of two million images showing 20,000 different objects to reach such generalization capabilities. To overcome this shortcoming we introduce ZS6D, for zero-shot novel object 6D pose estimation. Visual descriptors, extracted using pre-trained Vision Transformers (ViT), are used for matching rendered templates against query images of objects and for establishing local correspondences. These local correspondences enable deriving geometric correspondences and are used for estimating the object's 6D pose with RANSAC- based PnP. This approach showcases that the image descriptors extracted by pre-trained ViTs are well-suited to achieve a notable improvement over two state-of-the-art novel object 6D pose estimation methods, without the need for task-specific fine-tuning. Experiments are performed on LMO, YCBV, and TLESS. In comparison to MegaPose, we improve the Average Recall on all three datasets and compared to OSOP we improve on two datasets. The code is available at https://github.com/PhilippAuss/ZS6D. Philipp Ausserlechner, David Haberger, Stefan Thalhammer, Jean-Baptiste Weibel, Markus Vincze |
ICRA | 3 |
| 2024 | EdgeSoil 2.0 - Soil Analyzer Using Convolutional Neural Network and Camera Imaging for Agricultural RoboticsabstractSoil is the most important building element of agriculture and its analysis is crucial for healthy plants and a high crop yield. But apart from its importance, soil analysis is a tedious and time-consuming task. This paper presents EdgeSoil 2.0, a non-invasive, accurate, and real-time robotic system for soil pH prediction, a key parameter of soil status for farmers. The EdgeSoil 2.0 predicts the pH value of the soil in real-time, using a live video stream from a webcam with an average of 7 FPS. The method is suitable to be implemented on edge devices necessary for the application: we are using a mobile robot with the NVIDIA Jetson Nano module which is running a pH-estimator trained with a Convolutional Neural Network (CNN) on a novel dataset we built for this purpose. Predictions are performed while the robot is moving over the plowed field before the planting process starts. In order to achieve the best performance, we train the pH-estimator with different input modalities and validate each result using Mean Squared Error (MSE) and Standard Deviation (SD). We are able to achieve accurate results with the MSE value of 0.08, the SD value of 0.15, and with testing results from the field showing up to ± 0.3 deviation from the GT value during prediction, which is sufficient to comply with agricultural standards. Roni Kasemi, Lara Lammer, Stefan Thalhammer, Markus Vincze |
ICRA | 3 |
| 2024 | Challenges for Monocular 6-D Object Pose Estimation in RoboticsabstractObject pose estimation is a core perception task that enables, for example, object manipulation and scene understanding. The widely available, inexpensive, and high-resolution RGB sensors and CNNs that allow for fast inference make monocular approaches especially well-suited for robotics applications. We observe that previous surveys establish the state of the art for varying modalities, single- and multiview settings, and datasets and metrics that consider a multitude of applications. We argue, however, that those works' broad scope hinders the identification of open challenges that are specific to monocular approaches and the derivation of promising future challenges for their application in robotics. By providing a unified view on recent publications from both robotics and computer vision, we find that occlusion handling, pose representations, and formalizing and improving category-level pose estimation are still fundamental challenges that are highly relevant for robotics. Moreover, to further improve robotic performance, large object sets, novel objects, refractive materials, and uncertainty estimates are central and largely unsolved open challenges. In order to address them, ontological reasoning, deformability handling, scene-level reasoning, realistic datasets, and the ecological footprint of algorithms need to be improved. Stefan Thalhammer, Dominik Bauer, Peter Hönig, Jean-Baptiste Weibel, José García Rodríguez 0001, Markus Vincze |
IEEE Trans. Robotics | 1 |
| 2023 | COPE: End-to-end trainable Constant Runtime Object Pose EstimationabstractState-of-the-art object pose estimation handles multiple instances in a test image by using multi-model formulations: detection as a first stage and then separately trained networks per object for 2D-3D geometric correspondence prediction as a second stage. Poses are subsequently estimated using the Perspective-n-Points algorithm at runtime. Unfortunately, multi-model formulations are slow and do not scale well with the number of object instances involved. Recent approaches show that direct 6D object pose estimation is feasible when derived from the aforementioned geometric correspondences. We present an approach that learns an intermediate geometric representation of multiple objects to directly regress 6D poses of all instances in a test image. The inherent end-to-end trainability overcomes the requirement of separately processing individual object instances. By calculating the mutual Intersection-over-Unions, pose hypotheses are clustered into distinct instances, which achieves negligible runtime overhead with respect to the number of object instances. Results on multiple challenging standard datasets show that the pose estimation performance is superior to single-model state-of-the-art approaches despite being more than ~35 times faster. We additionally provide an analysis showing real-time applicability (> 24 fps) for images where more than 90 object instances are present. Further results show the advantage of supervising geometric correspondence-based object pose estimation with the 6D pose. Stefan Thalhammer, Tim Patten, Markus Vincze |
WACV | 1 |
| 2023 | Self-supervised Vision Transformers for 3D pose estimation of novel objectsabstractObject pose estimation is important for object manipulation and scene understanding. In order to improve the general applicability of pose estimators, recent research focuses on providing estimates for novel objects, that is, objects unseen during training. Such works use deep template matching strategies to retrieve the closest template connected to a query image, which implicitly provides object class and pose. Despite the recent success and improvements of Vision Transformers over CNNs for many vision tasks, the state of the art uses CNN-based approaches for novel object pose estimation. This work evaluates and demonstrates the differences between self-supervised CNNs and Vision Transformers for deep template matching. In detail, both types of approaches are trained using contrastive learning to match training images against rendered templates of isolated objects. At test time such templates are matched against query images of known and novel objects under challenging settings, such as clutter, occlusion and object symmetries, using masked cosine similarity. The presented results not only demonstrate that Vision Transformers improve matching accuracy over CNNs but also that for some cases pre-trained Vision Transformers do not need fine-tuning to achieve the improvement. Furthermore, we highlight the differences in optimization and network architecture when comparing these two types of networks for deep template matching. Stefan Thalhammer, Jean-Baptiste Weibel, Markus Vincze, José García Rodríguez 0001 |
Image Vis. Comput. | 1 |
| 2021 | PyraPose: Feature Pyramids for Fast and Accurate Object Pose Estimation under Domain ShiftabstractObject pose estimation enables robots to understand and interact with their environments. Training with synthetic data is necessary in order to adapt to novel situations. Unfortunately, pose estimation under domain shift, i.e., training on synthetic data and testing in the real world, is challenging. Deep learning-based approaches currently perform best when using encoder-decoder networks but typically do not generalize to new scenarios with different scene characteristics. We argue that patch-based approaches, instead of encoder-decoder networks, are more suited for synthetic-to-real transfer because local to global object information is better represented. To that end, we present a novel approach based on a specialized feature pyramid network to compute multi-scale features for creating pose hypotheses on different feature map resolutions in parallel. Our single-shot pose estimation approach is evaluated on multiple standard datasets and outperforms the state of the art by up to ∼35 %. We also perform grasping experiments in the real world to demonstrate the advantage of using synthetic data to generalize to novel environments. Stefan Thalhammer, Markus Leitner, Tim Patten, Markus Vincze |
ICRA | 1 |
| 2019 | SyDPose: Object Detection and Pose Estimation in Cluttered Real-World Depth Images Trained using Only Synthetic DataabstractObject pose estimation is an important problem in robotics because it supports scene understanding and enables subsequent grasping and manipulation. Many methods, including modern deep learning approaches, exploit known object models, however, in industry these are difficult and expensive to obtain. 3D CAD models, on the other hand, are often readily available. Consequently, training a deep architecture for pose estimation exclusively from CAD models leads to a considerable decrease of the data creation effort. While this has been shown to work well for feature-and template-based approaches, real-world data is still required for pose estimation in clutter using deep learning. We use synthetically created depth data with domain-relevant background randomized noise heuristics to train an end-to-end, multi-task network, for pose estimation. We simultaneously detect, classify and estimate the poses of texture-less objects in cluttered real-world depth images of an arbitrary amount of objects. We present the results of our experiments with the LineMOD and the Occlusion dataset. Stefan Thalhammer, Tim Patten, Markus Vincze |
3DV | 1 |