EDBT 2026 Demo / reviewers in the wild / expert
Stefan Hinterstoißer
dblp:57/261 · also Stefan Hinterstoisser
· DBLP profile ↗
15ranked-venue papers
10as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 10 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
3D vision · 32% Image recognition and object detection · 21% Robot manipulation · 17% | |
| Computer graphics and multimedia
4 papers |
Geometric modeling and processing · 37% Computational photography and imaging · 29% Multimedia analysis and retrieval · 19% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation |
0.3 | 1 | 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation · ICRA 2018 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.3 | 1 | 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation · ICRA 2018 |
Robotics › Robot manipulation
grasping |
0.3 | 1 | 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation · ICRA 2018 |
Robotics › Robot manipulation › grasping
instance grasping |
0.3 | 1 | 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation · ICRA 2018 |
Computer vision › Image recognition and object detection
object detection |
0.3 | 3 | 2010 | Dominant orientation templates for real-time detection of texture-less objects · CVPR 2010 Distance transform templates for object detection and pose estimation · CVPR 2009 Online learning of patch perspective rectification for efficient object detection · CVPR 2008 |
Computer vision › 3D vision
object pose estimation |
0.3 | 3 | 2010 | Dominant orientation templates for real-time detection of texture-less objects · CVPR 2010 Real-time learning of accurate patch rectification · CVPR 2009 Online learning of patch perspective rectification for efficient object detection · CVPR 2008 |
Computer vision › Image recognition and object detection › object detection
texture-less object detection |
0.3 | 2 | 2012 | Gradient Response Maps for Real-Time Detection of Textureless Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2012 Dominant orientation templates for real-time detection of texture-less objects · CVPR 2010 |
Computer vision › 3D vision › geometric estimation
3d registration |
0.2 | 1 | 2016 | Going Further with Point Pair Features · ECCV (3) 2016 |
Geometric modeling and processing › 3d scene understanding
3d object recognition |
0.2 | 1 | 2016 | Going Further with Point Pair Features · ECCV (3) 2016 |
Computer vision › 3D vision › feature matching › 3d correspondence
2d-3d correspondence |
0.2 | 1 | 2014 | Optimal Local Searching for Fast and Robust Textureless 3D Object Tracking in Highly Cluttered Backgrounds · IEEE Trans. Vis. Comput. Graph. 2014 |
Computer vision › Video understanding and tracking
object tracking |
0.2 | 1 | 2014 | Optimal Local Searching for Fast and Robust Textureless 3D Object Tracking in Highly Cluttered Backgrounds · IEEE Trans. Vis. Comput. Graph. 2014 |
Computer vision › Image recognition and object detection › object detection
instance detection |
0.1 | 1 | 2012 | Gradient Response Maps for Real-Time Detection of Textureless Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Computer vision › 3D vision
3d object detection |
0.1 | 2 | 2012 | Online learning of patch perspective rectification for efficient object detection · CVPR 2008 Gradient Response Maps for Real-Time Detection of Textureless Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Multimedia analysis and retrieval
object detection |
0.1 | 1 | 2011 | Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes · ICCV 2011 |
Computer vision › Video understanding and tracking › object tracking › appearance-based tracking
template tracking |
0.1 | 1 | 2010 | Rapid selection of reliable templates for visual tracking · CVPR 2010 |
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation |
0.1 | 1 | 2009 | Distance transform templates for object detection and pose estimation · CVPR 2009 |
Computer vision › Image recognition and object detection › object detection
planar object detection |
0.1 | 1 | 2009 | Distance transform templates for object detection and pose estimation · CVPR 2009 |
Computer vision › 3D vision
pose estimation |
0.1 | 1 | 2009 | Distance transform templates for object detection and pose estimation · CVPR 2009 |
Robotics › Robot navigation and mapping › SLAM
visual SLAM |
0.1 | 1 | 2009 | Real-time learning of accurate patch rectification · CVPR 2009 |
Virtual and augmented reality › tracking
camera pose estimation |
0.1 | 1 | 2007 | N3M: Natural 3D Markers for Real-Time Object Detection and Pose Estimation · ICCV 2007 |
Computer vision › Image recognition and object detection
template matching |
0.0 | 1 | 2012 | Gradient Response Maps for Real-Time Detection of Textureless Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Computer vision › Video understanding and tracking › multi-object tracking
tracking-by-detection |
0.0 | 1 | 2010 | Rapid selection of reliable templates for visual tracking · CVPR 2010 |
Image and video processing › image matching
template matching |
0.0 | 1 | 2009 | Distance transform templates for object detection and pose estimation · CVPR 2009 |
Immersive interaction › augmented reality
industrial augmented reality |
0.0 | 1 | 2007 | An Industrial Augmented Reality Solution For Discrepancy Check · ISMAR 2007 |
Methods — techniques the papers use, named apart from their topics
simulation-to-real transfer · 0.3domain adversarial training · 0.3region appearance modeling · 0.2optimal local searching · 0.2online learning · 0.2surface normal orientations · 0.1spread image gradient orientations · 0.1gradient response maps · 0.1multimodal templates · 0.1learning-based rectification · 0.1depth map fusion · 0.1image feature selection · 0.1template matching · 0.1distance transform · 0.1classifier training · 0.1multi-level validation · 0.1image matching · 0.1feature point selection · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from SimulationabstractLearning-based approaches to robotic manipulation are limited by the scalability of data collection and accessibility of labels. In this paper, we present a multi-task domain adaptation framework for instance grasping in cluttered scenes by utilizing simulated robot experiments. Our neural network takes monocular RGB images and the instance segmentation mask of a specified target object as inputs, and predicts the probability of successfully grasping the specified object for each candidate motor command. The proposed transfer learning framework trains a model for instance grasping in simulation and uses a domain-adversarial loss to transfer the trained model to real robots using indiscriminate grasping data, which is available both in simulation and the real world. We evaluate our model in real-world robot experiments, comparing it with alternative model architectures as well as an indiscriminate grasping baseline. Kuan Fang, Stefan Hinterstoißer, Silvio Savarese, Mrinal Kalakrishnan |
ICRA | 3 |
| 2016 | Going Further with Point Pair Features
Stefan Hinterstoißer, Vincent Lepetit, Naresh Rajkumar, Kurt Konolige |
ECCV (3) | 1 |
| 2014 | Optimal Local Searching for Fast and Robust Textureless 3D Object Tracking in Highly Cluttered BackgroundsabstractEdge-based tracking is a fast and plausible approach for textureless 3D object tracking, but its robustness is still very challenging in highly cluttered backgrounds due to numerous local minima. To overcome this problem, we propose a novel method for fast and robust textureless 3D object tracking in highly cluttered backgrounds. The proposed method is based on optimal local searching of 3D-2D correspondences between a known 3D object model and 2D scene edges in an image with heavy background clutter. In our searching scheme, searching regions are partitioned into three levels (interior, contour, and exterior) with respect to the previous object region, and confident searching directions are determined by evaluating candidates of correspondences on their region levels; thus, the correspondences are searched among likely candidates in only the confident directions instead of searching through all candidates. To ensure the confident searching direction, we also adopt the region appearance, which is efficiently modeled on a newly defined local space (called a searching bundle). Experimental results and performance evaluations demonstrate that our method fully supports fast and robust textureless 3D object tracking even in highly cluttered backgrounds. Byung-Kuk Seo, Hanhoon Park, Jong-Il Park, Stefan Hinterstoißer, Slobodan Ilic |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2012 | Model Based Training, Detection and Pose Estimation of Texture-Less 3D Objects in Heavily Cluttered Scenes
Stefan Hinterstoißer, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary R. Bradski, Kurt Konolige, Nassir Navab |
ACCV (1) | 1 |
| 2012 | Gradient Response Maps for Real-Time Detection of Textureless ObjectsabstractWe present a method for real-time 3D object instance detection that does not require a time-consuming training stage, and can handle untextured objects. At its core, our approach is a novel image representation for template matching designed to be robust to small image transformations. This robustness is based on spread image gradient orientations and allows us to test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. In addition, we demonstrate that if a dense depth sensor is available we can extend our approach for an even better performance also taking 3D surface normal orientations into account. We show how to take advantage of the architecture of modern computers to build an efficient but very discriminant representation of the input images that can be used to consider thousands of templates in real time. We demonstrate in many experiments on real data that our method is much faster and more robust with respect to background clutter than current state-of-the-art methods. Stefan Hinterstoißer, Cedric Cagniart, Slobodan Ilic, Peter F. Sturm, Nassir Navab, Pascal Fua, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenesabstractWe present a method for detecting 3D objects using multi-modalities. While it is generic, we demonstrate it on the combination of an image and a dense depth map which give complementary object information. It works in real-time, under heavy clutter, does not require a time consuming training stage, and can handle untextured objects. It is based on an efficient representation of templates that capture the different modalities, and we show in many experiments on commodity hardware that our approach significantly outperforms state-of-the-art methods on single modalities. Stefan Hinterstoißer, Stefan Holzer, Cedric Cagniart, Slobodan Ilic, Kurt Konolige, Nassir Navab, Vincent Lepetit |
ICCV | 1 |
| 2011 | Learning Real-Time Perspective Patch Rectification
Stefan Hinterstoißer, Vincent Lepetit, Selim Benhimane, Pascal Fua, Nassir Navab |
Int. J. Comput. Vis. | 1 |
| 2010 | Rapid selection of reliable templates for visual trackingabstractWe propose a method that rates the suitability of given templates for template-based tracking in real-time. This is important for applications with online template selection, such as SLAM, where it is essential to track a low number of preferably reliable templates. Our approach is based on simple image features specifically designed to identify texture properties which are problematic for tracking. During a training step, a support vector régresser is learned. It uses a tracking quality measure which considers both convergence rate and speed obtained by simulation of many tracking attempts. Finally, a minimum set of image features is identified to speedup the online selection process. In experiments on real-world video sequences our method improved the detection rate of an existing tracking-by-detection system by 8% on average. Nicolas Alt, Stefan Hinterstoißer, Nassir Navab |
CVPR | 2 |
| 2010 | Dominant orientation templates for real-time detection of texture-less objectsabstractWe present a method for real-time 3D object detection that does not require a time consuming training stage, and can handle untextured objects. At its core, is a novel template representation that is designed to be robust to small image transformations. This robustness based on dominant gradient orientations lets us test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. We show that together with a binary representation that makes evaluation very fast and a branch-and-bound approach to efficiently scan the image, it can detect untextured objects in complex situations and provide their 3D pose in real-time. Stefan Hinterstoißer, Vincent Lepetit, Slobodan Ilic, Pascal Fua, Nassir Navab |
CVPR | 1 |
| 2009 | Real-time learning of accurate patch rectificationabstractRecent work showed that learning-based patch rectification methods are both faster and more reliable than affine region methods. Unfortunately, their performance improvements are founded in a computationally expensive offline learning stage, which is not possible for applications such as SLAM. In this paper we propose an approach whose training stage is fast enough to be performed at run-time without the loss of accuracy or robustness. To this end, we developed a very fast method to compute the mean appearances of the feature points over sets of small variations that span the range of possible camera viewpoints. Then, by simply matching incoming feature points against these mean appearances, we get a coarse estimate of the viewpoint that is refined afterwards. Because there is no need to compute descriptors for the input image, the method is very fast at run-time. We demonstrate our approach on tracking-by-detection for SLAM, real-time object detection and pose estimation applications. Stefan Hinterstoißer, Oliver Kutter, Nassir Navab, Pascal Fua, Vincent Lepetit |
CVPR | 1 |
| 2009 | Distance transform templates for object detection and pose estimationabstractWe propose a new approach for detecting low textured planar objects and estimating their 3D pose. Standard matching and pose estimation techniques often depend on texture and feature points. They fail when there is no or only little texture available. Edge-based approaches mostly can deal with these limitations but are slow in practice when they have to search for six degrees of freedom. We overcome these problems by introducing the distance transform templates, generated by applying the distance transform to standard edge based templates. We obtain robustness against perspective transformations by training a classifier for various template poses. In addition, spatial relations between multiple contours on the template are learnt and later used for outlier removal. At runtime, the classifier provides the identity and a rough 3D pose of the distance transform template, which is further refined by a modified template matching algorithm that is also based on the distance transform. We qualitatively and quantitatively evaluate our approach on synthetic and real-life examples and demonstrate robust real-time performance. Stefan Holzer, Stefan Hinterstoißer, Slobodan Ilic, Nassir Navab |
CVPR | 2 |
| 2008 | Simultaneous Recognition and Homography Extraction of Local Patches with a Simple Linear ClassifierabstractWe show that the simultaneous estimation of keypoint identities and poses is more reliable than the two separate steps undertaken by previous approaches. A simple linear classifier coupled with linear predictors trained during a learning phase appears to be sufficient for this task. The retrieved poses are subpixel accurate due to the linear predictors. We demonstrate the advantages of our approach on real-time 3D object detection and tracking applications. Thanks to the high accuracy, one single keypoint is often enough to precisely estimate the object pose. As a result, we can deal in real-time with objects that are significantly less textured than the ones required by state-of-the-art methods. 1 Stefan Hinterstoißer, Selim Benhimane, Vincent Lepetit, Pascal Fua, Nassir Navab |
BMVC | 1 |
| 2008 | Online learning of patch perspective rectification for efficient object detectionabstractFor a large class of applications, there is time to train the system. In this paper, we propose a learning-based approach to patch perspective rectification, and show that it is both faster and more reliable than state-of-the-art ad hoc affine region detection methods. Our method performs in three steps. First, a classifier provides for every keypoint not only its identity, but also a first estimate of its transformation. This estimate allows carrying out, in the second step, an accurate perspective rectification using linear predictors. We show that both the classifier and the linear predictors can be trained online, which makes the approach convenient. The last step is a fast verification - made possible by the accurate perspective rectification - of the patch identity and its sub-pixel precision position estimation. We test our approach on real-time 3D object detection and tracking applications. We show that we can use the estimated perspective rectifications to determine the object pose and as a result, we need much fewer correspondences to obtain a precise pose estimation. Stefan Hinterstoißer, Selim Benhimane, Nassir Navab, Pascal Fua, Vincent Lepetit |
CVPR | 1 |
| 2007 | N3M: Natural 3D Markers for Real-Time Object Detection and Pose EstimationabstractIn this paper, a new approach for object detection and pose estimation is introduced. The contribution consists in the conception of entities permitting stable detection and reliable pose estimation of a given object. Thanks to a well- defined off-line learning phase, we design local and minimal subsets of feature points that have, at the same time, distinctive photometric and geometric properties. We call these entities Natural 3D Markers (N3Ms). Constraints on the selection and the distribution of the subsets coupled with a multi-level validation approach result in a detection at high frame rates and allow us to determine the precise pose of the object. The method is robust against noise, partial occlusions, background clutter and illumination changes. The experiments show its superiority to existing standard methods. The validation was carried out using simulated ground truth data. Excellent results on real data demonstrated the usefulness of this approach for many computer vision applications. Stefan Hinterstoißer, Selim Benhimane, Nassir Navab |
ICCV | 1 |
| 2007 | An Industrial Augmented Reality Solution For Discrepancy CheckabstractConstruction companies employ CAD software during the planning phase, but what is finally built often does not match the original plan. The procedure of validating the model is called "discrepancy check". The system proposed here allows the user to easily obtain an augmentation in order to find differences between the planned 3D model and the built items. The main difference to previous body of work in this field is the emphasis on usability and acceptance of the solution. While standard image-based solutions use markers or rely on a "perfect" 3D model to find the pose of the camera, our software uses anchor-plates. Anchor-Plates are rectangular structures installed on walls and ceiling in the majority of industrial edifices. We are using them as landmarks because they are the most reliable components often used as reference coordinates by constructors. Furthermore, for real industrial applications, they are the most suitable solutions in terms of general applicability. Unfortunately, they have not been designed with computer vision applications in mind. On the contrary, they are often made or painted in such way that they are not easily popping out. They are therefore difficult targets to segment and to track. This paper proposes a solution to extract and match them to their 3D counterparts. We created a software that uses the detected structures for pose estimation and image augmentation. The software has been successfully employed to find discrepancies in several rooms of two industrial plants. Pierre Fite Georgel, Pierre Schroeder, Selim Benhimane, Stefan Hinterstoißer, Mirko Appel, Nassir Navab |
ISMAR | 4 |