EDBT 2026 Demo / reviewers in the wild / expert
Hayko Riemenschneider
dblp:88/5817
· DBLP profile ↗
34ranked-venue papers
6as first author
2since 2021 · last 2024
0000-0003-2541-7999ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 29 · 6 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
3D vision · 70% Generative modeling · 13% Segmentation and scene understanding · 9% | |
| Computer graphics and multimedia
8 papers |
Geometric modeling and processing · 45% Rendering · 32% Visual content generation and editing · 11% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 96% Algorithms and data structures · 4% |
Topics — the 30 heaviest of 42, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation · NeurIPS 2024 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.8 | 1 | 2024 | BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation · NeurIPS 2024 |
Computer vision › 3D vision › depth estimation › deep depth estimation
zero-shot depth estimation |
0.8 | 1 | 2024 | BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation · NeurIPS 2024 |
Rendering
differentiable rendering |
0.8 | 1 | 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from Images · CVPR 2024 |
Rendering › inverse rendering
material and illumination decomposition |
0.8 | 1 | 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from Images · CVPR 2024 |
Geometric modeling and processing
mesh generation |
0.8 | 1 | 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from Images · CVPR 2024 |
Geometric modeling and processing › mesh processing
remeshing |
0.8 | 1 | 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from Images · CVPR 2024 |
Computer vision › 3D vision
3d reconstruction |
0.6 | 3 | 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from Images · CVPR 2024 Superpixel meshes for fast edge-preserving surface reconstruction · CVPR 2015 Fast, Approximate Piecewise-Planar Modeling Based on Sparse Structure-from-Motion and Superpixels · CVPR 2014 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.4 | 2 | 2015 | Superpixel meshes for fast edge-preserving surface reconstruction · CVPR 2015 Fast, Approximate Piecewise-Planar Modeling Based on Sparse Structure-from-Motion and Superpixels · CVPR 2014 |
Visual content generation and editing
texture synthesis |
0.4 | 2 | 2014 | The Synthesizability of Texture Examples · CVPR 2014 Example-Based Facade Texture Synthesis · ICCV 2013 |
Computer vision › 3D vision › shape matching
shape registration |
0.3 | 1 | 2018 | Consensus Maximization for Semantic Region Correspondences · CVPR 2018 |
Mathematical optimization › global optimization
consensus maximization |
0.3 | 1 | 2018 | Consensus Maximization for Semantic Region Correspondences · CVPR 2018 |
Mathematical optimization
global optimization |
0.3 | 1 | 2018 | Consensus Maximization for Semantic Region Correspondences · CVPR 2018 |
Computer vision › 3D vision
implicit neural representation |
0.2 | 1 | 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from Images · CVPR 2024 |
Computer vision › 3D vision
3d scene understanding |
0.2 | 1 | 2015 | 3D all the way: Semantic segmentation of urban scenes from start to end in 3D · CVPR 2015 |
Computer vision › 3D vision › 3d reconstruction
surface reconstruction |
0.2 | 1 | 2015 | Superpixel meshes for fast edge-preserving surface reconstruction · CVPR 2015 |
Computer vision › Segmentation and scene understanding › scene parsing
facade parsing |
0.2 | 2 | 2015 | Irregular lattices for complex shape grammar facade parsing · CVPR 2012 3D all the way: Semantic segmentation of urban scenes from start to end in 3D · CVPR 2015 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.2 | 1 | 2014 | Learning Where to Classify in Multi-view Semantic Segmentation · ECCV (5) 2014 |
Computer vision › Video understanding and tracking
video summarization |
0.2 | 1 | 2014 | Creating Summaries from User Videos · ECCV (7) 2014 |
Visual content generation and editing › texture synthesis
example-based texture synthesis |
0.2 | 1 | 2014 | The Synthesizability of Texture Examples · CVPR 2014 |
Multimedia analysis and retrieval
image analysis |
0.2 | 1 | 2014 | The Synthesizability of Texture Examples · CVPR 2014 |
Geometric modeling and processing
procedural modeling |
0.2 | 1 | 2013 | Is There a Procedural Logic to Architecture? · CVPR 2013 |
Geometric modeling and processing
shape analysis |
0.2 | 1 | 2013 | Is There a Procedural Logic to Architecture? · CVPR 2013 |
Geometric modeling and processing
urban modeling |
0.2 | 1 | 2013 | Example-Based Facade Texture Synthesis · ICCV 2013 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.1 | 1 | 2012 | Hough Regions for Joining Instance Localization and Segmentation · ECCV (3) 2012 |
Computer vision › 3D vision › 3d scene reconstruction
urban reconstruction |
0.1 | 1 | 2012 | Irregular lattices for complex shape grammar facade parsing · CVPR 2012 |
Geometric modeling and processing › procedural modeling
shape grammar |
0.1 | 1 | 2012 | Irregular lattices for complex shape grammar facade parsing · CVPR 2012 |
Computer vision › 3D vision
structure from motion |
0.1 | 2 | 2015 | Superpixel meshes for fast edge-preserving surface reconstruction · CVPR 2015 Fast, Approximate Piecewise-Planar Modeling Based on Sparse Structure-from-Motion and Superpixels · CVPR 2014 |
Image and video processing
edge detection |
0.1 | 1 | 2010 | Linked edges as stable region boundaries · CVPR 2010 |
Mathematical optimization › continuous optimization
convex optimization |
0.1 | 1 | 2018 | Consensus Maximization for Semantic Region Correspondences · CVPR 2018 |
Methods — techniques the papers use, named apart from their topics
vertex displacement · 1.5neural implicit representation · 1.5catmull-clark subdivision · 1.5pre-alignment · 0.8patch masking · 0.8diffusion model · 0.8linear matrix inequality constraints · 0.7ellipsoid approximation · 0.7branch-and-bound · 0.7image parsing · 0.3genetic algorithm · 0.3example-based synthesis · 0.3depth optimization · 0.2texture features · 0.2predictor training · 0.2grammatical inference · 0.2feature engineering · 0.2distance transformation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from ImagesabstractDespite exciting progress in automatic 3D reconstruction from images, excessive and irregular triangular faces in the resulting meshes still constitute a significant challenge when it comes to adoption in practical artist work-flows. Therefore, we propose a method to extract regular quad-dominant meshes from posed images. More specifically, we generate a high-quality 3D model through de-composition into an easily editable quad-dominant mesh with pixel-level details such as displacement, materials, and lighting. To enable end-to-end learning of shape and quad topology, we QUADify a neural implicit representation using our novel differentiable re-meshing objective. Distinct from previous work, our method exploits artifact-free Catmull-Clark subdivision combined with vertex displacement to extract pixel-level details linked to the base geom-etry. Finally, we apply differentiable rendering techniques for material and lighting decomposition to optimize for image reconstruction. Our experiments show the benefits of end-to-end re-meshing and that our method yields state-of-the-art geometric accuracy while providing lightweight meshes with displacements and textures that are directly compatible with professional renderers and game engines. Maximilian Frühauf, Hayko Riemenschneider, Markus Gross 0001, Christopher Schroers |
CVPR | 2 |
| 2024 | BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth EstimationabstractBy training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhibit a superior ability to extract details, they struggle in geometrically complex scenes that challenge their geometry prior, trained on less diverse 3D data. To leverage the complementary merits of both worlds, we propose BetterDepth to achieve geometrically correct affine-invariant MDE while capturing fine details. Specifically, BetterDepth is a conditional diffusion-based refiner that takes the prediction from pre-trained MDE models as depth conditioning, in which the global depth layout is well-captured, and iteratively refines details based on the input image. For the training of such a refiner, we propose global pre-alignment and local patch masking methods to ensure BetterDepth remains faithful to the depth conditioning while learning to add fine-grained scene details. With efficient training on small-scale synthetic datasets, BetterDepth achieves state-of-the-art zero-shot MDE performance on diverse public datasets and on in-the-wild scenes. Moreover, BetterDepth can improve the performance of other MDE models in a plug-and-play manner without further re-training. Xiang Zhang 0022, Bingxin Ke, Hayko Riemenschneider, Nando Metzger, Anton Obukhov, Markus Gross 0001, Konrad Schindler, Christopher Schroers |
NeurIPS | 3 |
| 2018 | Consensus Maximization for Semantic Region CorrespondencesabstractWe propose a novel method for the geometric registration of semantically labeled regions. We approximate semantic regions by ellipsoids, and leverage their convexity to formulate the correspondence search effectively as a constrained optimization problem that maximizes the number of matched regions, and which we solve globally optimal in a Branch-and-Bound fashion. To this end, we derive suitable linear matrix inequality constraints which describe ellipsoid-to-ellipsoid assignment conditions. Our approach is robust to large percentages of outliers and thus applicable to difficult correspondence search problems. In multiple experiments we demonstrate the flexibility and robustness of our approach on a number of challenging vision problems. Pablo Speciale, Danda Pani Paudel, Martin R. Oswald, Hayko Riemenschneider, Luc Van Gool, Marc Pollefeys |
CVPR | 4 |
| 2017 | Efficient edge-aware surface mesh reconstruction for urban scenes
András Bódis-Szomorú, Hayko Riemenschneider, Luc Van Gool |
Comput. Vis. Image Underst. | 2 |
| 2017 | Efficient architectural structural element decomposition
Nikolay Kobyshev, Hayko Riemenschneider, András Bódis-Szomorú, Luc Van Gool |
Comput. Vis. Image Underst. | 2 |
| 2016 | 3D Saliency for Finding Landmark BuildingsabstractIn urban environments the most interesting and effective factors for localization and navigation are landmark buildings. This paper proposes a novel method to detect such buildings that stand out, i.e. would be given the status of 'landmark'. The method works in a fully unsupervised way, i.e. it can be applied to different cities without requiring annotation. First, salient points are detected, based on the analysis of their features as well as those found in their spatial neighborhood. Second, learning refines the points by finding connected landmark components and training a classifier to distinguish these from common building components. Third, landmark components are aggregated into complete landmark buildings. Experiments on city-scale point clouds show the viability and efficiency of our approach on various tasks. Nikolay Kobyshev, Hayko Riemenschneider, András Bódis-Szomorú, Luc Van Gool |
3DV | 2 |
| 2016 | Efficient volumetric fusion of airborne and street-side data for urban reconstructionabstractAirborne acquisition and on-road mobile mapping provide complementary 3D information of an urban landscape: the former acquires roof structures, ground, and vegetation at a large scale, but lacks the facade and street-side details, while the latter is incomplete for higher floors and often totally misses out on pedestrian-only areas or undriven districts. In this work, we introduce an approach that efficiently unifies a detailed street-side Structure-from-Motion (SfM) or Multi-View Stereo (MVS) point cloud and a coarser but more complete point cloud from airborne acquisition in a joint surface mesh. We propose a point cloud blending and a volumetric fusion based on ray casting across a 3D tetrahedralization (3DT), extended with data reduction techniques to handle large datasets. To the best of our knowledge, we are the first to adopt a 3DT approach for airborne/street-side data fusion. Our pipeline exploits typical characteristics of airborne and ground data, and produces a seamless, watertight mesh that is both complete and detailed. Experiments on 3D urban data from multiple sources and different data densities show the effectiveness and benefits of our approach. András Bódis-Szomorú, Hayko Riemenschneider, Luc Van Gool |
ICPR | 2 |
| 2016 | Dilemma First Search for effortless optimization of NP-hard problemsabstractTo tackle the exponentiality associated with NP-hard problems, two paradigms have been proposed. First, Branch & Bound, like Dynamic Programming, achieve efficient exact inference but requires extensive information and analysis about the problem at hand. Second, meta-heuristics are easier to implement but comparatively inefficient. As a result, a number of problems have been left unoptimized and plain greedy solutions are used. We introduce a theoretical framework and propose a powerful yet simple search method called Dilemma First Search (DFS). DFS exploits the decision heuristic needed for the greedy solution for further optimization. DFS is useful when it is hard to design efficient exact inference. We evaluate DFS on two problems: First, the Knapsack problem, for which efficient algorithms exist, serves as a toy example. Second, Decision Tree inference, where state-of-the-art algorithms rely on the greedy or randomness-based solutions. We further show that decision trees benefit from optimizations that are performed in a fraction of the iterations required by a random-based search. Julien Weissenberg, Hayko Riemenschneider, Ralf Dragon, Luc Van Gool |
ICPR | 2 |
| 2016 | Architectural decomposition for 3D landmark building understandingabstractDecomposing 3D building models into architectural elements is an essential step in understanding their 3D structure. Although we focus on landmark buildings, our approach generalizes to arbitrary 3D objects. We formulate the decomposition as a multi-label optimization that identifies individual elements of a landmark. This allows our system to cope with noisy, incomplete, outlier-contaminated 3D point clouds. We detect three types of structural cues, namely dominant mirror symmetries, rotational symmetries, and polylines capturing free-form shapes of the landmark not explained by symmetry. Combining these cues enables modeling the variability present in complex 3D models, and robustly decomposing them into architectural structural elements. Our architectural decomposition facilitates significant 3D model compression and shape-specific modeling. Nikolay Kobyshev, Hayko Riemenschneider, András Bódis-Szomorú, Luc Van Gool |
WACV | 2 |
| 2016 | Mobile phone and cloud - A dream team for 3D reconstructionabstractRecently, Structure-from-Motion pipelines (SfM) for the 3D reconstruction of scenes from images were pushed from desktop computers onto mobile devices, like phones or tablets. However, mobile devices offer much more than just necessary computational power. A combination of handheld device with camera, display and full connectivity entails possibilities for an on-line 3D reconstruction that would have been difficult to implement otherwise. In this work, we propose a combination of a regular mobile phone as frontend with a centralized server plus annex cloud as backend for collaborative, on-line 3D reconstruction. We illustrate few advantages of this combination of a myriad of new possibilities: First, we automatically balance computational load between the frontend and the backend depending on battery autonomy and available bandwidth. Second, we select the best of algorithms given the available resources to obtain better 3D models. Finally, we allow for collaborative modeling in order to arrive at more complete and more detailed models, especially when the objects or scenes are big. This paper presents an implementation of such a joint mobile-cloud modeling approach and demonstrates its advantages via real-life reconstructions. Alex Locher, Michal Perdoch, Hayko Riemenschneider, Luc Van Gool |
WACV | 3 |
| 2015 | Superpixel meshes for fast edge-preserving surface reconstructionabstractMulti-View-Stereo (MVS) methods aim for the highest detail possible, however, such detail is often not required. In this work, we propose a novel surface reconstruction method based on image edges, superpixels and second-order smoothness constraints, producing meshes comparable to classic MVS surfaces in quality but orders of magnitudes faster. Our method performs per-view dense depth optimization directly over sparse 3D Ground Control Points (GCPs), hence, removing the need for view pairing, image rectification, and stereo depth estimation, and allowing for full per-image parallelization. We use Structure-from-Motion (SfM) points as GCPs, but the method is not specific to these, e.g. LiDAR or RGB-D can also be used. The resulting meshes are compact and inherently edge-aligned with image gradients, enabling good-quality lightweight per-face flat renderings. Our experiments demonstrate on a variety of 3D datasets the superiority in speed and competitive surface quality. András Bódis-Szomorú, Hayko Riemenschneider, Luc Van Gool |
CVPR | 2 |
| 2015 | 3D all the way: Semantic segmentation of urban scenes from start to end in 3DabstractWe propose a new approach for semantic segmentation of 3D city models. Starting from an SfM reconstruction of a street-side scene, we perform classification and facade splitting purely in 3D, obviating the need for slow image-based semantic segmentation methods. We show that a properly trained pure-3D approach produces high quality labelings, with significant speed benefits (20x faster) allowing us to analyze entire streets in a matter of minutes. Additionally, if speed is not of the essence, the 3D labeling can be combined with the results of a state-of-the-art 2D classifier, further boosting the performance. Further, we propose a novel facade separation based on semantic nuances between facades. Finally, inspired by the use of architectural principles for 2D facade labeling, we propose new 3D-specific principles and an efficient optimization scheme based on an integer quadratic programming formulation. Andelo Martinovic, Jan Knopp, Hayko Riemenschneider, Luc Van Gool |
CVPR | 3 |
| 2014 | Matching Features Correctly through Semantic UnderstandingabstractImage-to-image feature matching is the single most restrictive time bottleneck in any matching pipeline. We propose two methods for improving the speed and quality by employing semantic scene segmentation. First, we introduce a way of capturing semantic scene context of a key point into a compact description. Second, we propose to learn correct match ability of descriptors from these semantic contexts. Finally, we further reduce the complexity of matching to only a pre-computed set of semantically close key points. All methods can be used independently and in the evaluation we show combinations for maximum speed benefits. Overall, our proposed methods outperform all baselines and provide significant improvements in accuracy and an order of magnitude faster key point matching. Nikolay Kobyshev, Hayko Riemenschneider, Luc Van Gool |
3DV | 2 |
| 2014 | An Integer Linear Programming Model for View Selection on Overlapping Camera ClustersabstractMulti-View Stereo (MVS) algorithms scale poorly on large image sets, and quickly become unfeasible to run on a single machine with limited memory. Typical solutions to lower the complexity include reducing the redundancy of the image set (view selection), and dividing the image set in groups to be processed independently (view clustering). A novel formulation for view selection is proposed here. We express the problem with an Integer Linear Programming (ILP) model, where cameras are modeled with binary variables, while the linear constraints enforce the completeness of the 3D reconstruction. The solution of the ILP leads to an optimal subset of selected cameras. As a second contribution, we integrate ILP camera selection with a view clustering approach which exploits Leveraged Affinity Propagation (LAP). LAP clustering can efficiently deal with large camera sets. We adapt the original algorithm so that it provides a set of overlapping clusters where the minimum and maximum sizes and the number of overlapping cameras can be specified. Evaluations on four different dataset show our solution provides significant complexity reductions and guarantees near-perfect coverage, making large reconstructions feasible even on a single machine. Massimo Mauro, Hayko Riemenschneider, Alberto Signoroni, Riccardo Leonardi, Luc Van Gool |
3DV | 2 |
| 2014 | Frankenhorse: Automatic Completion of Articulating Objects from Image-based Reconstruction
Alex Mansfield, Nikolay Kobyshev, Hayko Riemenschneider, Will Chang, Luc Van Gool |
BMVC | 3 |
| 2014 | A unified framework for content-aware view selection and planning through view importance
Massimo Mauro, Hayko Riemenschneider, Alberto Signoroni, Riccardo Leonardi, Luc Van Gool |
BMVC | 2 |
| 2014 | Fast, Approximate Piecewise-Planar Modeling Based on Sparse Structure-from-Motion and SuperpixelsabstractState-of-the-art Multi-View Stereo (MVS) algorithms deliver dense depth maps or complex meshes with very high detail, and redundancy over regular surfaces. In turn, our interest lies in an approximate, but light-weight method that is better to consider for large-scale applications, such as urban scene reconstruction from ground-based images. We present a novel approach for producing dense reconstructions from multiple images and from the underlying sparse Structure-from-Motion (SfM) data in an efficient way. To overcome the problem of SfM sparsity and textureless areas, we assume piecewise planarity of man-made scenes and exploit both sparse visibility and a fast over-segmentation of the images. Reconstruction is formulated as an energy-driven, multi-view plane assignment problem, which we solve jointly over superpixels from all views while avoiding expensive photoconsistency computations. The resulting planar primitives -- defined by detailed superpixel boundaries -- are computed in about 10 seconds per image. András Bódis-Szomorú, Hayko Riemenschneider, Luc Van Gool |
CVPR | 2 |
| 2014 | The Synthesizability of Texture ExamplesabstractExample-based texture synthesis (ETS) has been widely used to generate high quality textures of desired sizes from a small example. However, not all textures are equally well reproducible that way. We predict how synthesizable a particular texture is by ETS. We introduce a dataset (21, 302 textures) of which all images have been annotated in terms of their synthesizability. We design a set of texture features, such as 'textureness', homogeneity, repetitiveness, and irregularity, and train a predictor using these features on the data collection. This work is the first attempt to quantify this image property, and we find that texture synthesizability can be learned and predicted. We use this insight to trim images to parts that are more synthesizable. Also we suggest which texture synthesis method is best suited to synthesise a given texture. Our approach can be seen as 'winner-uses-all': picking one method among several alternatives, ending up with an overall superior ETS method. Such strategy could also be considered for other vision tasks: rather than building an even stronger method, choose from existing methods based on some simple preprocessing. Dengxin Dai, Hayko Riemenschneider, Luc Van Gool |
CVPR | 2 |
| 2014 | Creating Summaries from User Videos
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, Luc Van Gool |
ECCV (7) | 3 |
| 2014 | Learning Where to Classify in Multi-view Semantic Segmentation
Hayko Riemenschneider, András Bódis-Szomorú, Julien Weissenberg, Luc Van Gool |
ECCV (5) | 1 |
| 2013 | Overlapping camera clustering through dominant sets for scalable 3D reconstructionabstractIn this work we present a method for clustering large unordered sets of cameras. Our method uses camera view information available from Structure-from-Motion (SfM) for computing a set of overlapping clusters suited for Multi-View Stereo (MVS) reconstruction. Our formulation of the problem uses the game theoretic model of dominant sets to find competing clustering solutions with computational simplicity. The overlapping solutions ensure more robust partial reconstructions. Experimental evaluations show that our method produces more regular cluster and overlap configurations with respect to the state of the art. This allows more scalable and higher quality reconstructions, while speeding up 6 times with respect to a MVS which uses all images at once. c 2013. Massimo Mauro, Hayko Riemenschneider, Luc Van Gool, Riccardo Leonardi |
BMVC | 2 |
| 2013 | Is There a Procedural Logic to Architecture?abstractUrban models are key to navigation, architecture and entertainment. Apart from visualizing facades, a number of tedious tasks remain largely manual (e.g. compression, generating new facade designs and structurally comparing facades for classification, retrieval and clustering). We propose a novel procedural modelling method to automatically learn a grammar from a set of facades, generate new facade instances and compare facades. To deal with the difficulty of grammatical inference, we reformulate the problem. Instead of inferring a compromising, one-size-fits-all, single grammar for all tasks, we infer a model whose successive refinements are production rules tailored for each task. We demonstrate our automatic rule inference on datasets of two different architectural styles. Our method supercedes manual expert work and cuts the time required to build a procedural model of a facade from several days to a few milliseconds. Julien Weissenberg, Hayko Riemenschneider, Mukta Prasad, Luc Van Gool |
CVPR | 2 |
| 2013 | Example-Based Facade Texture SynthesisabstractThere is an increased interest in the efficient creation of city models, be it virtual or as-built. We present a method for synthesizing complex, photo-realistic facade images, from a single example. After parsing the example image into its semantic components, a tiling for it is generated. Novel tilings can then be created, yielding facade textures with different dimensions or with occluded parts in painted. A genetic algorithm guides the novel facades as well as in painted parts to be consistent with the example, both in terms of their overall structure and their detailed textures. Promising results for multiple standard datasets - in particular for the different building styles they contain - demonstrate the potential of the method. Dengxin Dai, Hayko Riemenschneider, Gerhard Schmitt, Luc Van Gool |
ICCV | 2 |
| 2013 | The Interestingness of ImagesabstractWe investigate human interest in photos. Based on our own and others' psychophysical experiments, we identify various cues for "interestingness", namely aesthetics, unusualness and general preferences. For the ranking of retrieved images, interestingness shows to be more appropriate than cues proposed earlier. Interestingness is correlated with what people believe they will remember. This is opposed to actual memorability, which is uncorrelated to both. We introduce a set of features computationally capturing the three main aspects of visual interestingness and build an interestingness predictor from them. Its performance is shown on three datasets with varying context, reflecting the prior knowledge of the viewers. Michael Gygli, Helmut Grabner, Hayko Riemenschneider, Fabian Nater, Luc Van Gool |
ICCV | 3 |
| 2012 | Irregular lattices for complex shape grammar facade parsingabstractHigh-quality urban reconstruction requires more than multi-view reconstruction and local optimization. The structure of facades depends on the general layout, which has to be optimized globally. Shape grammars are an established method to express hierarchical spatial relationships, and are therefore suited as representing constraints for semantic facade interpretation. Usually inference uses numerical approximations, or hard-coded grammar schemes. Existing methods inspired by classical grammar parsing are not applicable on real-world images due to their prohibitively high complexity. This work provides feasible generic facade reconstruction by combining low-level classifiers with mid-level object detectors to infer an irregular lattice. The irregular lattice preserves the logical structure of the facade while reducing the search space to a manageable size. We introduce a novel method for handling symmetry and repetition within the generic grammar. We show competitive results on two datasets, namely the Paris 2010 and the Graz 50. The former includes only Hausmannian, while the latter includes Classicism, Biedermeier, Historicism, Art Nouveau and post-modern architectural styles. Hayko Riemenschneider, Ulrich Krispel, Wolfgang Thaller, Michael Donoser, Sven Havemann, Dieter W. Fellner, Horst Bischof |
CVPR | 1 |
| 2012 | Hough Regions for Joining Instance Localization and Segmentation
Hayko Riemenschneider, Sabine Sternig, Michael Donoser, Peter M. Roth, Horst Bischof |
ECCV (3) | 1 |
| 2011 | Discriminative Learning of Contour Fragments for Object DetectionabstractThe goal of this work is to discriminatively learn contour fragment descriptors for the task of object detection. Unlike previous methods that incorporate learning techniques only for object model generation or for verification after detection, we present a holistic object detection system using solely shape as underlying cue. In the learning phase, we interrelate local shape descriptions (fragments) of the object contour with the corresponding spatial location of the object centroid. We introduce a novel shape fragment descriptor that abstracts spatially connected edge points into a matrix consisting of angular relations between the points. Our proposed descriptor fulfills important properties like distinctiveness, robustness and insensitivity to clutter. During detection, we hypothesize object locations in a generalized Hough voting scheme. The back-projected votes from the fragments allow to approximately delineate the object contour. We evaluate our method e.g. on the well-known ETHZ shape data base, where we achieve an average detection score of 87:5% at 1:0 FPPI only from Hough voting, outperforming the highest scoring Hough voting approaches by almost 8%. Peter Kontschieder, Hayko Riemenschneider, Michael Donoser, Horst Bischof |
BMVC | 2 |
| 2010 | Linked edges as stable region boundariesabstractMany of the recently popular shape based category recognition methods require stable, connected and labeled edges as input. This paper introduces a novel method to find the most stable region boundaries in grayscale images for this purpose. In contrast to common edge detection algorithms as Canny, which only analyze local discontinuities in image brightness, our method integrates mid-level information by analyzing regions that support the local gradient magnitudes. We use a component tree where every node contains a single connected region obtained from thresholding the gradient magnitude image. Edges in the tree are defined by an inclusion relationship between nested regions in different levels of the tree. Region boundaries which are similar in shape (i. e. have a low chamfer distance) across several levels of the tree are included in the final result. Since the component tree can be calculated in quasi-linear time and chamfer matching between nodes in the component tree is reduced to analysis of the distance transformation, results are obtained in an efficient manner. The proposed detection algorithm labels all identified edges during calculation, thus avoiding the cumbersome post-processing of connecting and labeling edge responses. We evaluate our method on two reference data sets and demonstrate improved performance for shape prototype based localization of objects in images. Michael Donoser, Hayko Riemenschneider, Horst Bischof |
CVPR | 2 |
| 2010 | Using Partial Edge Contour Matches for Efficient Object Category Localization
Hayko Riemenschneider, Michael Donoser, Horst Bischof |
ECCV (5) | 1 |
| 2010 | Shape Prototype Signatures for Action RecognitionabstractRecognizing human actions in video sequences is frequently based on analyzing the shape of the human silhouette as the main feature. In this paper we introduce a method for recognizing different actions by comparing signatures of similarities to pre-defined shape prototypes. In training, we build a vocabulary of shape prototypes by clustering a training set of human silhouettes and calculate prototype similarity signatures for all training videos. During testing a prototype signature is calculated for the test video and is aligned to each training signature by dynamic time warping. A simple voting scheme over the similarities to the training videos provides action classification results and temporal alignments to the training videos. Experimental evaluation on a reference data set demonstrates that state-of-the-art results are achieved. Michael Donoser, Hayko Riemenschneider, Horst Bischof |
ICPR | 2 |
| 2010 | Shape Guided Maximally Stable Extremal Region (MSER) TrackingabstractMaximally Stable Extremal Regions (MSERs) are one of the most prominent interest region detectors in computer vision due to their powerful properties and low computational demands. In general MSERs are detected in single images, but given image sequences as input, the repeatability of MSER detection can be improved by exploiting correspondences between subsequent frames by feature based analysis. Such an approach fails during fast movements, in heavily cluttered scenes and in images containing several similar sized regions because of the simple feature based analysis. In this paper we propose an extension of MSER tracking by considering shape similarity as strong cue for defining the frame-to-frame correspondences. Efficient calculation of shape similarity scores ensures that real-time capability is maintained. Experimental evaluation demonstrates improved repeatability and an application for tracking weakly textured, planar objects. Michael Donoser, Hayko Riemenschneider, Horst Bischof |
ICPR | 2 |
| 2009 | Efficient Partial Shape Matching of Outer Contours
Michael Donoser, Hayko Riemenschneider, Horst Bischof |
ACCV (1) | 2 |
| 2009 | Bag of Optical Flow Volumes for Image Sequence RecognitionabstractThis paper introduces a novel 3D interest point detector and feature representation for describing image sequences. The approach considers image sequences as spatio-temporal volumes and detects Maximally Stable Volumes (MSVs) in efficiently calcu-lated optical flow fields. This provides a set of binary optical flow volumes highlighting the dominant motions in the sequences. 3D interest points are sampled on the surface of the volumes which balance well between density and informativeness. The binary opti-cal flow volumes are used as feature representation in a 3D shape context descriptor. A standard bag-of-words approach then allows building discriminant optical flow volume signatures for predicting class labels of previously unseen image sequences by machine learning algorithms. We evaluate the proposed method for the task of action recognition on the well-known Weizmann dataset, and show that we outperform recently proposed state-of-the-art 3D interest point detection and description methods. 1 Hayko Riemenschneider, Michael Donoser, Horst Bischof |
BMVC | 1 |
| 2008 | Online object recognition by MSER trajectoriesabstractThis work presents a robust online learning and recognition system. The basic idea is to exploit information from tracking an object during the recognition and/or learning stage to obtain increased robustness and better recognition results. Object tracking by means of an extended MSER tracker is utilized to detect local features and construct their trajectories. Compact object representations are formed by summarizing the trajectories. All steps are performed online including the MSER detection, tracking, summarization, SIFT description as well as learning and recognition based on a vocabulary tree. The proposed method is evaluated on realistic video sequences which prove the increased performance for robust online recognition. Hayko Riemenschneider, Michael Donoser, Horst Bischof |
ICPR | 1 |