EDBT 2026 Demo / reviewers in the wild / expert
Bastian Goldlücke
dblp:12/4880 · also Bastian Goldluecke
· DBLP profile ↗
54ranked-venue papers
17as first author
12since 2021 · last 2026
0000-0003-3427-4029ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 14 first-author · 9 since 2021Artificial intelligence and machine learning · 39 · 11 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometric 2D Scene Graph GenerationabstractIn production processes for consumer products, assembly instructions are essential not only for planning but also for executing the production process. Likewise in robotics, it is crucial for an assembly robot to understand how components fit together and can be assembled. To facilitate these tasks, we contribute a method for constructing scene graphs to represent and characterize assembly relationships between components. Our approach does not rely on semantic data and is capable of handling a very small dataset. To realize this, the output of a Faster R-CNN model is used to create geometric representations, which are then processed by a transformer architecture to generate an adjacency matrix. This matrix serves as input to a Siamese network that uses message passing based on an attentional graph convolutional network (aGCN) architecture to characterize the connections between the components. We validate our method on a study dataset of toy model components which can be assembled i nto transportation vehicles. Christoph Jahn, Urs Waldmann, Bastian Goldlücke |
ICAART (3) | 3 |
| 2026 | A Deep Network for Object Detection on Inland WatersabstractCollisions on inland waters frequently occur due to the absence of clearly defined navigation guidance. To prevent such accidents, optical sensors such as stereo cameras can be employed to detect obstacles at an early stage, enabling timely warnings to vessel operators or the planning of collision avoidance trajectories. In this context, the paper presents a neural network for object detection on inland waters, designed to address the specific challenges of waterborne vehicles, including strong ego-motion and contextual cues like the shoreline appearing in the background. The proposed network leverages a plane sweep approach to integrate multiple views of a scene and predict object locations in bird’s-eye view (BEV) coordinates. Its ability to incorporate more than two camera perspectives is demonstrated using the KITTI dataset. On real-world inland water data, the network is evaluated against both a traditional maritime stereo-based method and a learning-based stereo detector from autonomous driving, showing consistently higher mean average precision. The code is available at https://github.com/dionysos4/IWOD. Dennis Grießer, Bastian Goldlücke, Matthias O. Franz, Georg Umlauf |
WACV | 2 |
| 2025 | Editor's Note: Special Issue on German Conference on Pattern Recognition (DAGM GCPR)abstractThis special issue consists of two papers selected from the DAGM GCPR conference (German Conference on Pattern Recognition) held in Konstanz, Germany, September 27-30, 2022 jointly with VMV (Vision, Modeling and Visualizaition).It was the 44th edition of an international conference series organized by DAGM (The German Association for Pattern Recognition), which was run under the name DAGM Symposium for Pattern Recognition until 2012 und German Conference on Pattern Recognition afterwards.A selection of four papers ranked highest by an evaluation committee during the conference were invited by us, the program chairs of DAGM GCPR 2022, to submit extended manuscripts to this special issue.We received three submissions, which underwent an additional rigorous peer-review process according to the journal's high standards, and two were finally accepted for this special issue.The first article, by Yinghao Huang, Omid Taheri, Michael J. Black, and Dimitrios Tzionas addresses the challenge of reconstructing 3D interactions between humans and objects from images.They introduce and adress limitations of Inter-Cap, a novel method that uses multi-view RGB-D data and the SMPL-X model to reconstruct whole-body and object interactions.InterCap leverages contact points between the body and objects to improve pose estimation and uses Azure Kinect cameras to reduce occlusions.They further use their method to capture a dataset with 10 subjects interacting with various objects, including contact with hands and feet.Inter-Cap fills a significant gap in the literature and supports diverse research avenues, with data and code available online. Bastian Goldlücke |
Int. J. Comput. Vis. | 1 |
| 2024 | Sparse Views, Near Light: A Practical Paradigm for Uncalibrated Point-Light Photometric StereoabstractNeural approaches have shown a significant progress on camera-based reconstruction. But they require either a fairly dense sampling of the viewing sphere, or pre-training on an existing dataset, thereby limiting their generalizability. In contrast, photometric stereo (PS) approaches have shown great potential for achieving high-quality reconstruction under sparse viewpoints. Yet, they are impractical because they typically require tedious laboratory conditions, are restricted to dark rooms, and often multi-staged, making them subject to accumulated errors. To address these shortcomings, we propose an end-to-end uncalibrated multi-view PS frameworkfor reconstructing high-resolution shapes acquiredfrom sparse viewpoints in a real-world environment. We relax the dark room assumption, and allow a combination of static ambient lighting and dynamic near LED lighting, thereby enabling easy data capture outside the lab. Experimental validation confirms that it outperforms existing baseline approaches in the regime of sparse viewpoints by a large margin. This allows to bring high-accuracy 3D reconstruction from the dark room to the real world, while maintaining a reasonable data capture complexity. Mohammed Brahimi 0002, Bjoern Haefner, Zhenzhang Ye, Bastian Goldlücke, Daniel Cremers |
CVPR | 4 |
| 2024 | SupeRVol: Super-Resolution Shape and Reflectance Estimation in Inverse Volume RenderingabstractWe propose an end-to-end inverse rendering pipeline called SupeRVol that allows us to recover 3D shape and material parameters from a set of color images in a superresolution manner. To this end, we represent both the bidirectional reflectance distribution function’s (BRDF) parameters and the signed distance function (SDF) by multi-layer perceptrons (MLPs). In order to obtain both the surface shape and its reflectance properties, we revert to a differentiable volume renderer with a physically based illumination model that allows us to decouple reflectance and lighting. This physical model takes into account the effect of the camera’s point spread function thereby enabling a reconstruction of shape and material in a super-resolution quality. Experimental validation confirms that SupeRVol achieves state of the art performance in terms of inverse rendering quality. It generates reconstructions that are sharper than the individual input images, making this method ideally suited for 3D modeling from low-resolution imagery. Mohammed Brahimi 0002, Bjoern Haefner, Tarun Yenamandra, Bastian Goldlücke, Daniel Cremers |
WACV | 4 |
| 2024 | 3D-MuPPET: 3D Multi-Pigeon Pose Estimation and TrackingabstractAbstract Markerless methods for animal posture tracking have been rapidly developing recently, but frameworks and benchmarks for tracking large animal groups in 3D are still lacking. To overcome this gap in the literature, we present 3D-MuPPET, a framework to estimate and track 3D poses of up to 10 pigeons at interactive speed using multiple camera views. We train a pose estimator to infer 2D keypoints and bounding boxes of multiple pigeons, then triangulate the keypoints to 3D. For identity matching of individuals in all views, we first dynamically match 2D detections to global identities in the first frame, then use a 2D tracker to maintain IDs across views in subsequent frames. We achieve comparable accuracy to a state of the art 3D pose estimator in terms of median error and Percentage of Correct Keypoints. Additionally, we benchmark the inference speed of 3D-MuPPET, with up to 9.45 fps in 2D and 1.89 fps in 3D, and perform quantitative tracking evaluation, which yields encouraging results. Finally, we showcase two novel applications for 3D-MuPPET. First, we train a model with data of single pigeons and achieve comparable results in 2D and 3D posture estimation for up to 5 pigeons. Second, we show that 3D-MuPPET also works in outdoors without additional annotations from natural environments. Both use cases simplify the domain shift to new species and environments, largely reducing annotation effort needed for 3D posture tracking. To the best of our knowledge we are the first to present a framework for 2D/3D animal posture and trajectory tracking that works in both indoor and outdoor environments for up to 10 individuals. We hope that the framework can open up new opportunities in studying animal collective behaviour and encourages further developments in 3D multi-animal posture tracking. Urs Waldmann, Alex Hoi Hang Chan, Hemal Naik, Nagy Máté, Iain D. Couzin, Oliver Deussen, Bastian Goldlücke, Fumihiro Kano |
Int. J. Comput. Vis. | 7 |
| 2023 | Incremental one-class learning using regularized null-space training for industrial defect detectionabstractOne-class incremental learning is a special case of class-incremental learning, where only a single novel class is incrementally added to an existing classifier instead of multiple classes. This case is relevant in industrial defect detection scenarios, where novel defects usually appear during operation. Existing rolled-out classifiers must be updated incrementally in this scenario with only a few novel examples. In addition, it is often required that the base classifier must not be altered due to approval and warranty restrictions. While simple finetuning often gives the best performance across old and new classes, it comes with the drawback of potentially losing performance on the base classes (catastrophic forgetting [1]). Simple prototype approaches [2] work without changing existing weights and perform very well when the classes are well separated but fail dramatically when not. In theory, null-space training (NSCL) [3] should retain the basis classifier entirely, as parameter updates are restricted to the null space of the network with respect to existing classes. However, as we show, this technique promotes overfitting in the case of one-class incremental learning. In our experiments, we found that unconstrained weight growth in null space is the underlying issue, leading us to propose a regularization term (R-NSCL) that penalizes the magnitude of amplification. The regularization term is added to the standard classification loss and stabilizes null-space training in the one-class scenario by counteracting overfitting. We test the method’s capabilities on two industrial datasets, namely AITEX and MVTec, and compare the performance to state-of-the-art algorithms for class-incremental learning. Matthias Hermann, Georg Umlauf, Bastian Goldlücke, Matthias O. Franz |
ICMV | 3 |
| 2022 | Neural Puppeteer: Keypoint-Based Neural Rendering of Dynamic Shapes
Simon Giebenhain, Urs Waldmann, Ole Johannsen, Bastian Goldlücke |
ACCV (4) | 4 |
| 2022 | GenDR: A Generalized Differentiable RendererabstractIn this work, we present and study a generalized family of differentiable renderers. We discuss from scratch which components are necessary for differentiable rendering and formalize the requirements for each component. We instantiate our general differentiable renderer, which generalizes existing differentiable renderers like SoftRas and DIB-R, with an array of different smoothing distributions to cover a large spectrum of reasonable settings. We evaluate an array of differentiable renderer instantiations on the popular ShapeNet 3D reconstruction benchmark and analyze the implications of our results. Surprisingly, the simple uniform distribution yields the best overall results when averaged over 13 classes; in general, however, the optimal choice of distribution heavily depends on the task. Felix Petersen, Bastian Goldlücke, Christian Borgelt, Oliver Deussen |
CVPR | 2 |
| 2022 | Style Agnostic 3D Reconstruction via Adversarial Style TransferabstractReconstructing the 3D geometry of an object from an image is a major challenge in computer vision. Recently introduced differentiable renderers can be leveraged to learn the 3D geometry of objects from 2D images, but those approaches require additional supervision to enable the renderer to produce an output that can be compared to the input image. This can be scene information or constraints such as object silhouettes, uniform backgrounds, material, texture, and lighting. In this paper, we propose an approach that enables a differentiable rendering-based learning of 3D objects from images with backgrounds without the need for silhouette supervision. Instead of trying to render an image close to the input, we propose an adversarial style-transfer and domain adaptation pipeline that allows to translate the input image domain to the rendered image domain. This allows us to directly compare between a translated image and the differentiable rendering of a 3D object reconstruction in order to train the 3D object reconstruction network. We show that the approach learns 3D geometry from images with backgrounds and provides a better performance than constrained methods for single-view 3D object reconstruction on this task. Felix Petersen, Bastian Goldlücke, Oliver Deussen, Hilde Kuehne |
WACV | 2 |
| 2021 | AIR-Nets: An Attention-Based Framework for Locally Conditioned Implicit RepresentationsabstractThis paper introduces Attentive Implicit Representation Networks (AIR-Nets), a simple, but highly effective architecture for 3D reconstruction from point clouds. Since representing 3D shapes in a local and modular fashion increases generalization and reconstruction quality, AIR-Nets encode an input point cloud into a set of local latent vectors anchored in 3D space, which locally describe the object’s geometry, as well as a global latent description, enforcing global consistency. Our model is the first grid-free, encoder-based approach that locally describes an implicit function. The vector attention mechanism from [62] serves as main point cloud processing module, and allows for permutation invariance and translation equivariance. When queried with a 3D coordinate, our decoder gathers information from the global and nearby local latent vectors in order to predict an occupancy value. Experiments on the ShapeNet dataset [7] show that AIR-Nets significantly outperform previous state-of-the-art encoder-based, implicit shape learning methods and especially dominate in the sparse setting. Furthermore, our model generalizes well to the FAUST dataset [1] in a zero-shot setting. Finally, since AIR-Nets use a sparse latent representation and follow a simple operating scheme, the model offers several exiting avenues for future work. Our code is available at https: //github.com/SimonGiebenhain/AIR-Nets. Simon Giebenhain, Bastian Goldlücke |
3DV | 2 |
| 2021 | Towards Monocular Shape from Refraction
Antonin Sulc, Imari Sato, Bastian Goldlücke, Tali Treibitz |
BMVC | 3 |
| 2020 | L2R GAN: LiDAR-to-Radar Translation
Leichen Wang, Bastian Goldlücke, Carsten Anklam |
ACCV (3) | 2 |
| 2020 | High Dimensional Frustum PointNet for 3D Object Detection from Camera, LiDAR, and RadarabstractFusing the raw data from different automotive sensors for real-world environment perception is still challenging due to their different representations and data formats. In this work, we propose a novel method termed High Dimensional Frustum PointNet for 3D object detection in the context of autonomous driving. Motivated by the goals data diversity and lossless processing of the data, our deep learning approach directly and jointly uses the raw data from the camera, LiDAR, and radar. In more detail, given 2D region proposals and classification from camera images, a high dimensional convolution operator captures local features from a point cloud enhanced with color and temporal information. Radars are used as adaptive plug-in sensors to refine object detection performance. As shown by an extensive evaluation on the nuScenes 3D detection benchmark, our network outperforms most of the previous methods. Leichen Wang, Tianbai Chen, Carsten Anklam, Bastian Goldlücke |
IV | 4 |
| 2020 | ATQAM/MAST'20: Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal TrendsabstractThe Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/ MAST) aims to bring together researchers and professionals working in fields ranging from computer vision, multimedia computing, multimodal signal processing to psychology and social sciences. It is divided into two tracks: ATQAM and MAST. ATQAM track: Visual quality assessment techniques can be divided into image and video technical quality assessment (IQA and VQA, or broadly TQA) and aesthetics quality assessment (AQA). While TQA is a long-standing field, having its roots in media compression, AQA is relatively young. Both have received increased attention with developments in deep learning. The topics have mostly been studied separately, even though they deal with similar aspects of the underlying subjective experience of media. The aim is to bring together individuals in the two fields of TQA and AQA for the sharing of ideas and discussions on current trends, developments, issues, and future directions. MAST track: The research area of media content analytics has been traditionally used to refer to applications involving inference of higher-level semantics from multimedia content. However, multimedia is typically created for human consumption, and we believe it is necessary to adopt a human-centered approach to this analysis, which would not only enable a better understanding of how viewers engage with content but also how they impact each other in the process. Tanaya Guha, Vlad Hosu, Dietmar Saupe, Bastian Goldlücke, Naveen Kumar 0004, Weisi Lin, Victor R. Martinez, Krishna Somandepalli, Shri Narayanan, Wen-Huang Cheng, Kree Cole-McLaughlin, Hartwig Adam, John See, Lai-Kuan Wong |
ACM Multimedia | 4 |
| 2019 | Effective Aesthetics Prediction With Multi-Level Spatially Pooled FeaturesabstractWe propose an effective deep learning approach to aesthetics quality assessment that relies on a new type of pre-trained features, and apply it to the AVA data set, the currently largest aesthetics database. While previous approaches miss some of the information in the original images, due to taking small crops, down-scaling or warping the originals during training, we propose the first method that efficiently supports full resolution images as an input, and can be trained on variable input sizes. This allows us to significantly improve upon the state of the art, increasing the Spearman rank-order correlation coefficient (SRCC) of ground-truth mean opinion scores (MOS) from the existing best reported of 0.612 to 0.756. To achieve this performance, we extract multi-level spatially pooled (MLSP) features from all convolutional blocks of a pre-trained InceptionResNet-v2 network, and train a custom shallow Convolutional Neural Network (CNN) architecture on these new features. Vlad Hosu, Bastian Goldlücke, Dietmar Saupe |
CVPR | 2 |
| 2018 | Light Field Intrinsics With a Deep Encoder-Decoder NetworkabstractWe present a fully convolutional autoencoder for light fields, which jointly encodes stacks of horizontal and vertical epipolar plane images through a deep network of residual layers. The complex structure of the light field is thus reduced to a comparatively low-dimensional representation, which can be decoded in a variety of ways. The different pathways of upconvolution we currently support are for disparity estimation and separation of the lightfield into diffuse and specular intrinsic components. The key idea is that we can jointly perform unsupervised training for the autoencoder path of the network, and supervised training for the other decoders. This way, we find features which are both tailored to the respective tasks and generalize well to datasets for which only example light fields are available. We provide an extensive evaluation on synthetic light field data, and show that the network yields good results on previously unseen real world data captured by a Lytro Illum camera and various gantries. Anna Alperovich, Ole Johannsen, Michael Strecke, Bastian Goldlücke |
CVPR | 4 |
| 2018 | Structure-from-Motion-Aware PatchMatch for Adaptive Optical Flow Estimation
Daniel Maurer 0002, Nico Marniok, Bastian Goldlücke, Andrés Bruhn |
ECCV (8) | 3 |
| 2018 | Real-Time Variational Range Image Fusion and Visualization for Large-Scale Scenes Using GPU Hash TablesabstractWe present a real-time pipeline for large-scale 3D scene reconstruction from a single moving RGB-D camera together with interactive visualization. Our approach combines a time and space efficient data structure capable of representing large scenes, a local variational update algorithm and a visualization system. The environment's structure is reconstructed by integrating the depth image of each camera view into a sparse volume representation using a truncated signed distance function, which is organized via a hash table. Noise from real-world data is efficiently eliminated by immediately performing local variational refinements on newly integrated data. The whole pipeline is able to perform in real-time on consumer-available hardware and allows for simultaneous inspection of the currently reconstructed scene. Nico Marniok, Bastian Goldlücke |
WACV | 2 |
| 2018 | Bring It to the Pitch: Combining Video and Movement Data to Enhance Team Sport AnalysisabstractAnalysts in professional team sport regularly perform analysis to gain strategic and tactical insights into player and team behavior. Goals of team sport analysis regularly include identification of weaknesses of opposing teams, or assessing performance and improvement potential of a coached team. Current analysis workflows are typically based on the analysis of team videos. Also, analysts can rely on techniques from Information Visualization, to depict e.g., player or ball trajectories. However, video analysis is typically a time-consuming process, where the analyst needs to memorize and annotate scenes. In contrast, visualization typically relies on an abstract data model, often using abstract visual mappings, and is not directly linked to the observed movement context anymore. We propose a visual analytics system that tightly integrates team sport video recordings with abstract visualization of underlying trajectory data. We apply appropriate computer vision techniques to extract trajectory data from video input. Furthermore, we apply advanced trajectory and movement analysis techniques to derive relevant team sport analytic measures for region, event and player analysis in the case of soccer analysis. Our system seamlessly integrates video and visualization modalities, enabling analysts to draw on the advantages of both analysis forms. Several expert studies conducted with team sport analysts indicate the effectiveness of our integrated approach. Manuel Stein, Halldór Janetzko, Andreas Lamprecht, Thorsten Breitkreutz, Philipp Zimmermann, Bastian Goldlücke, Tobias Schreck, Gennady L. Andrienko, Michael Grossniklaus, Daniel A. Keim |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | Accurate Depth and Normal Maps from Occlusion-Aware Focal Stack SymmetryabstractWe introduce a novel approach to jointly estimate consistent depth and normal maps from 4D light fields, with two main contributions. First, we build a cost volume from focal stack symmetry. However, in contrast to previous approaches, we introduce partial focal stacks in order to be able to robustly deal with occlusions. This idea already yields significanly better disparity maps. Second, even recent sublabel-accurate methods for multi-label optimization recover only a piecewise flat disparity map from the cost volume, with normals pointing mostly towards the image plane. This renders normal maps recovered from these approaches unsuitable for potential subsequent applications. We therefore propose regularization with a novel prior linking depth to normals, and imposing smoothness of the resulting normal field. We then jointly optimize over depth and normals to achieve estimates for both which surpass previous work in accuracy on a recent benchmark. Michael Strecke, Anna Alperovich, Bastian Goldlücke |
CVPR | 3 |
| 2017 | 4D imaging through spray-on opticsabstractLight fields are a powerful concept in computational imaging and a mainstay in image-based rendering; however, so far their acquisition required either carefully designed and calibrated optical systems (micro-lens arrays), or multi-camera/multi-shot settings. Here, we show that fully calibrated light field data can be obtained from a single ordinary photograph taken through a partially wetted window. Each drop of water produces a distorted view on the scene, and the challenge of recovering the unknown mapping from pixel coordinates to refracted rays in space is a severely underconstrained problem. The key idea behind our solution is to combine ray tracing and low-level image analysis techniques (extraction of 2D drop contours and locations of scene features seen through drops) with state-of-the-art drop shape simulation and an iterative refinement scheme to enforce photo-consistency across features that are seen in multiple views. This novel approach not only recovers a dense pixel-to-ray mapping, but also the refractive geometry through which the scene is observed, to high accuracy. We therefore anticipate that our inherently self-calibrating scheme might also find applications in other fields, for instance in materials science where the wetting properties of liquids on surfaces are investigated. Julian Iseringhausen, Bastian Goldlücke, Nina Pesheva, Stanimir Iliev, Alexander Wender, Martin Fuchs 0001, Matthias B. Hullin |
ACM Trans. Graph. | 2 |
| 2016 | A Variational Model for Intrinsic Light Field Decomposition
Anna Alperovich, Bastian Goldlücke |
ACCV (3) | 2 |
| 2016 | A Dataset and Evaluation Methodology for Depth Estimation on 4D Light Fields
Katrin Honauer, Ole Johannsen, Daniel Kondermann, Bastian Goldlücke |
ACCV (3) | 4 |
| 2016 | Layered Scene Reconstruction from Multiple Light Field Camera Views
Ole Johannsen, Antonin Sulc, Nico Marniok, Bastian Goldlücke |
ACCV (3) | 4 |
| 2016 | What Sparse Light Field Coding Reveals about Scene StructureabstractIn this paper, we propose a novel method for depth estimation in light fields which employs a specifically designed sparse decomposition to leverage the depth-orientation relationship on its epipolar plane images. The proposed method learns the structure of the central view and uses this information to construct a light field dictionary for which groups of atoms correspond to unique disparities. This dictionary is then used to code a sparse representation of the light field. Analyzing the coefficients of this representation with respect to the disparities of their corresponding atoms yields an accurate and robust estimate of depth. In addition, if the light field has multiple depth layers, such as for reflective or transparent surfaces, statistical analysis of the coefficients can be employed to infer the respective depth of the superimposed layers. Ole Johannsen, Antonin Sulc, Bastian Goldlücke |
CVPR | 3 |
| 2016 | Editorial
Jingyi Ju, Bastian Goldlücke, Richard Szeliski, Tomás Pajdla |
Comput. Vis. Image Underst. | 2 |
| 2015 | On Linear Structure from Motion for Light Field CamerasabstractWe present a novel approach to relative pose estimation which is tailored to 4D light field cameras. From the relationships between scene geometry and light field structure and an analysis of the light field projection in terms of Pluecker ray coordinates, we deduce a set of linear constraints on ray space correspondences between a light field camera pair. These can be applied to infer relative pose of the light field cameras and thus obtain a point cloud reconstruction of the scene. While the proposed method has interesting relationships to pose estimation for generalized cameras based on ray-to-ray correspondence, our experiments demonstrate that our approach is both more accurate and computationally more efficient. It also compares favourably to direct linear pose estimation based on aligning the 3D point clouds obtained by reconstructing depth for each individual light field. To further validate the method, we employ the pose estimates to merge light fields captured with hand-held consumer light field cameras into refocusable panoramas. Ole Johannsen, Antonin Sulc, Bastian Goldlücke |
ICCV | 3 |
| 2014 | Spherical Light Fields
Bernd Krolla, Maximilian Diebold, Bastian Goldlücke, Didier Stricker |
BMVC | 3 |
| 2014 | Bayesian View Synthesis and Image-Based Rendering PrinciplesabstractIn this paper, we address the problem of synthesizing novel views from a set of input images. State of the art methods, such as the Unstructured Lumigraph, have been using heuristics to combine information from the original views, often using an explicit or implicit approximation of the scene geometry. While the proposed heuristics have been largely explored and proven to work effectively, a Bayesian formulation was recently introduced, formalizing some of the previously proposed heuristics, pointing out which physical phenomena could lie behind each. However, some important heuristics were still not taken into account and lack proper formalization. We contribute a new physics-based generative model and the corresponding Maximum a Posteriori estimate, providing the desired unification between heuristics-based methods and a Bayesian formulation. The key point is to systematically consider the error induced by the uncertainty in the geometric proxy. We provide an extensive discussion, analyzing how the obtained equations explain the heuristics developed in previous methods. Furthermore, we show that our novel Bayesian model significantly improves the quality of novel views, in particular if the scene geometry estimate is inaccurate. Sergi Pujades, Frederic Devernay, Bastian Goldlücke |
CVPR | 3 |
| 2014 | A Super-Resolution Framework for High-Accuracy Multiview Reconstruction
Bastian Goldlücke, Mathieu Aubry, Kalin Kolev, Daniel Cremers |
Int. J. Comput. Vis. | 1 |
| 2014 | Variational Light Field Analysis for Disparity Estimation and Super-ResolutionabstractWe develop a continuous framework for the analysis of 4D light fields, and describe novel variational methods for disparity reconstruction as well as spatial and angular super-resolution. Disparity maps are estimated locally using epipolar plane image analysis without the need for expensive matching cost minimization. The method works fast and with inherent subpixel accuracy since no discretization of the disparity space is necessary. In a variational framework, we employ the disparity maps to generate super-resolved novel views of a scene, which corresponds to increasing the sampling rate of the 4D light field in spatial as well as angular direction. In contrast to previous work, we formulate the problem of view synthesis as a continuous inverse problem, which allows us to correctly take into account foreshortening effects caused by scene geometry transformations. All optimization problems are solved with state-of-the-art convex relaxation techniques. We test our algorithms on a number of real-world examples as well as our new benchmark data set for light fields, and compare results to a multiview stereo method. The proposed method is both faster as well as more accurate. Data sets and source code are provided online for additional evaluation. Sven Wanner, Bastian Goldlücke |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | The Variational Structure of Disparity and Regularization of 4D Light FieldsabstractUnlike traditional images which do not offer information for different directions of incident light, a light field is defined on ray space, and implicitly encodes scene geometry data in a rich structure which becomes visible on its epipolar plane images. In this work, we analyze regularization of light fields in variational frameworks and show that their variational structure is induced by disparity, which is in this context best understood as a vector field on epipolar plane image space. We derive differential constraints on this vector field to enable consistent disparity map regularization. Furthermore, we show how the disparity field is related to the regularization of more general vector-valued functions on the 4D ray space of the light field. This way, we derive an efficient variational framework with convex priors, which can serve as a fundament for a large class of inverse problems on ray space. Bastian Goldlücke, Sven Wanner |
CVPR | 1 |
| 2013 | Globally Consistent Multi-label Assignment on the Ray Space of 4D Light FieldsabstractWe present the first variational framework for multi-label segmentation on the ray space of 4D light fields. For traditional segmentation of single images, features need to be extracted from the 2D projection of a three-dimensional scene. The associated loss of geometry information can cause severe problems, for example if different objects have a very similar visual appearance. In this work, we show that using a light field instead of an image not only enables to train classifiers which can overcome many of these problems, but also provides an optimal data structure for label optimization by implicitly providing scene geometry information. It is thus possible to consistently optimize label assignment over all views simultaneously. As a further contribution, we make all light fields available online with complete depth and segmentation ground truth data where available, and thus establish the first benchmark data set for light field analysis to facilitate competitive further development of algorithms. Sven Wanner, Christoph N. Straehle, Bastian Goldlücke |
CVPR | 3 |
| 2013 | Tight Convex Relaxations for Vector-Valued LabelingabstractMultilabel problems are of fundamental importance in computer vision and image analysis. Yet, finding global minima of the associated energies is typically a hard computational challenge. Recently, progress has been made by reverting to spatially continuous formulations of respective problems and solving the arising convex relaxation globally. In practice this leads to solutions which are either optimal or within an a posteriori bound of the optimum. Unfortunately, in previous methods, both run time and memory requirements scale linearly in the total number of labels, making these methods very inefficient and often not applicable to problems with higher dimensional label spaces. In this paper, we propose a reduction technique for the case that the label space is a continuous product space and the regularizer is separable, i.e., a sum of regularizers for each dimension of the label space. In typical real-world labeling problems, the resulting convex relaxation requires orders of magnitude less memory and computation time than previous methods. This enables us to apply it to large-scale problems like optic flow, stereo with occlusion detection, segmentation into a very large number of regions, and joint denoising and local noise estimation. Experiments show that despite the drastic gain in performance, we do not arrive at less accurate solutions than the original relaxation. Using the novel method, we can for the first time efficiently compute solutions to the optic flow functional which are within provable bounds (typically 5%) of the global optimum. Bastian Goldlücke, Evgeny Strekalovskiy, Daniel Cremers |
SIAM J. Imaging Sci. | 1 |
| 2012 | Globally consistent depth labeling of 4D light fieldsabstractWe present a novel paradigm to deal with depth reconstruction from 4D light fields in a variational framework. Taking into account the special structure of light field data, we reformulate the problem of stereo matching to a constrained labeling problem on epipolar plane images, which can be thought of as vertical and horizontal 2D cuts through the field. This alternative formulation allows to estimate accurate depth values even for specular surfaces, while simultaneously taking into account global visibility constraints in order to obtain consistent depth maps for all views. The resulting optimization problems are solved with state-of-the-art convex relaxation techniques. We test our algorithm on a number of synthetic and real-world examples captured with a light field gantry and a plenoptic camera, and compare to ground truth where available. All data sets as well as source code are provided online for additional evaluation. Sven Wanner, Bastian Goldlücke |
CVPR | 2 |
| 2012 | Spatial and Angular Variational Super-Resolution of 4D Light Fields
Sven Wanner, Bastian Goldlücke |
ECCV (5) | 2 |
| 2012 | The Natural Vectorial Total Variation Which Arises from Geometric Measure TheoryabstractSeveral ways to generalize scalar total variation to vector-valued functions have been proposed in the past. In this paper, we give a detailed analysis of a variant we denote by $\text{TV}_J$, which has not been previously explored as a regularizer. The contributions of the manuscript are twofold: on the theoretical side, we show that $\text{TV}_J$ can be derived from the generalized Jacobians from geometric measure theory. Thus, within the context of this theory, $\text{TV}_J$ is the most natural form of a vectorial total variation. As an important feature, we derive how $\text{TV}_J$ can be written as the support functional of a convex set in $\mathcal{L}^2$. This property allows us to employ fast and stable minimization algorithms to solve inverse problems. The analysis also shows that in contrast to other total variation regularizers for color images, the proposed one penalizes across a common edge direction for all channels, which is a major theoretical advantage. Our practical contribution consist of an extensive experimental section, where we compare the performance of a number of provable convergent algorithms for inverse problems with our proposed regularizer. In particular, we show in experiments for denoising, deblurring, superresolution, and inpainting that its use leads to a significantly better restoration of color images, both visually and quantitatively. Source code for all algorithms employed in the experiments is provided online. Bastian Goldlücke, Evgeny Strekalovskiy, Daniel Cremers |
SIAM J. Imaging Sci. | 1 |
| 2011 | Decoupling photometry and geometry in dense variational camera calibrationabstractWe introduce a spatially dense variational approach to estimate the calibration of multiple cameras in the context of 3D reconstruction. We propose a relaxation scheme which allows to transform the original photometric error into a geometric one, thereby decoupling the problems of dense matching and camera calibration. In both quantitative and qualitative experiments, we demonstrate that the proposed decoupling scheme allows for robust and accurate estimation of camera parameters. In particular, the presented dense camera calibration formulation leads to substantial improvements both in the reconstructed 3D geometry and in the super-resolution texture estimation. Mathieu Aubry, Kalin Kolev, Bastian Goldlücke, Daniel Cremers |
ICCV | 3 |
| 2011 | Introducing total curvature for image processingabstractWe introduce the novel continuous regularizer total curvature (TC) for images u: Ω → ℝ. It is defined as the Menger-Melnikov curvature of the Radon measure |Du|, which can be understood as a measure theoretic formulation of curvature mathematically related to mean curvature. The functional is not convex, therefore we define a convex relaxation which yields a close approximation. Similar to the total variation, the relaxation can be written as the support functional of a convex set, which means that there are stable and efficient minimization algorithms available when it is used as a regularizer in image processing problems. Our current implementation can handle general inverse problems, inpainting and segmentation. We demonstrate in experiments and comparisons how the regularizer performs in practice. Bastian Goldlücke, Daniel Cremers |
ICCV | 1 |
| 2011 | Tight convex relaxations for vector-valued labeling problemsabstractThe multi-label problem is of fundamental importance to computer vision, yet finding global minima of the associated energies is very hard and usually impossible in practice. Recently, progress has been made using continuous formulations of the multi-label problem and solving a convex relaxation globally, thereby getting a solution with optimality bounds. In this work, we develop a novel framework for continuous convex relaxations, where the label space is a continuous product space. In this setting, we can combine the memory efficient product relaxation of [9] with the much tighter relaxation of [5], which leads to solutions closer to the global optimum. Furthermore, the new setting allows us to formulate more general continuous regularizers, which can be freely combined in the different label dimensions. We also improve upon the relaxation of the products in the data term of [9], which removes the need for artificial smoothing and allows the use of exact solvers. Evgeny Strekalovskiy, Bastian Goldlücke, Daniel Cremers |
ICCV | 2 |
| 2011 | Motion Field Estimation from Alternate Exposure ImagesabstractTraditional optical flow algorithms rely on consecutive short-exposed images. In this work, we make use of an additional long-exposed image for motion field estimation. Long-exposed images integrate motion information directly in the form of motion-blur. With this additional information, more robust and accurate motion fields can be estimated. In addition, the moment of occlusion can be determined. Considering the basic signal-theoretical problem in motion field estimation, we exploit the fact that long-exposed images integrate motion information to prevent temporal aliasing. A suitable image formation model relates the long-exposed image to preceding and succeeding short-exposed images in terms of dense 2D motion and per-pixel occlusion/disocclusion timings. Based on our image formation model, we describe a practical variational algorithm to estimate the motion field not only for visible image regions but also for regions getting occluded. Results for synthetic as well as real-world scenes demonstrate the validity of the approach. Anita Sellent, Martin Eisemann, Bastian Goldlücke, Daniel Cremers, Marcus A. Magnor |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | An approach to vectorial total variation based on geometric measure theoryabstractWe analyze a previously unexplored generalization of the scalar total variation to vector-valued functions, which is motivated by geometric measure theory. A complete mathematical characterization is given, which proves important invariance properties as well as existence of solutions of the vectorial ROF model. As an important feature, there exists a dual formulation for the proposed vectorial total variation, which leads to a fast and stable minimization algorithm. The main difference to previous approaches with similar properties is that we penalize across a common edge direction for all channels, which is a major theoretical advantage. Experiments show that this leads to a significantly better restoration of color edges in practice. Bastian Goldlücke, Daniel Cremers |
CVPR | 1 |
| 2010 | Convex Relaxation for Multilabel Problems with Product Label Spaces
Bastian Goldlücke, Daniel Cremers |
ECCV (5) | 1 |
| 2009 | Superresolution texture maps for multiview reconstructionabstractWe study the scenario of a multiview setting, where several calibrated views of a textured object with known surface geometry are available. The objective is to estimate a diffuse texture map as precisely as possible. A superresolution image formation model based on the camera properties leads to a total variation energy for the desired texture map, which can be recovered as the minimizer of the functional by solving the Euler-Lagrange equation on the surface. The PDE is transformed to planar texture space via an automatically created conformal atlas, where it can be solved using total variation deblurring. The proposed approach allows to recover a high-resolution, high-quality texture map even from lower-resolution photographs, which is of interest for a variety of image-based modeling applications. Bastian Goldlücke, Daniel Cremers |
ICCV | 1 |
| 2007 | Weighted Minimal Hypersurface ReconstructionabstractMany problems in computer vision can be formulated as a minimization problem for an energy functional. If this functional is given as an integral of a scalar-valued weight function over an unknown hypersurface, then the sought-after minimal surface can be determined as a solution of the functional's Euler-Lagrange equation. This paper deals with a general class of weight functions that may depend on surface point coordinates as well as surface orientation. We derive the Euler-Lagrange equation in arbitrary dimensional space without the need for any surface parameterization, generalizing existing proofs. Our work opens up the possibility of solving problems involving minimal hypersurfaces in a dimension higher than three, which were previously impossible to solve in practice. We also introduce two applications of our new framework: We show how to reconstruct temporally coherent geometry from multiple video streams, and we use the same framework for the volumetric reconstruction of refractive and transparent natural phenomena, here bodies of flowing water. Bastian Goldlücke, Ivo Ihrke, Christian Linz, Marcus A. Magnor |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Reconstructing the Geometry of Flowing WaterabstractWe present a recording scheme, image formation model and reconstruction method that enables image-based modeling of flowing bodies of water from multivideo input data. The recorded water is dyed with a fluorescent chemical to measure the thickness of a column of water which leads to an image formation model based on integrated emissivities along a viewing ray. This model allows for a photo-consistency based error measure for a weighted minimal surface, which is recovered using a PDE obtained from the Euler-Lagrangian formulation of the problem. The resulting equation is solved using the level set method. Ivo Ihrke, Bastian Goldlücke, Marcus A. Magnor |
ICCV | 2 |
| 2005 | Spacetime-continuous geometry meshes from multi-view video sequencesabstractWe reconstruct geometry for a time-varying scene given by a number of video sequences. The dynamic geometry is represented by a 3D hypersurface embedded in space-time. The intersection of the hypersurface with planes of constant time then yields the geometry at a single time instant. In this paper, we model the hypersurface with a collection of triangle meshes, one for each time frame. The photo-consistency error is measured by an error functional defined as an integral over the hypersurface. It can be minimized using a PDE driven surface evolution, which simultaneously optimizes space-time continuity as well. Compared to our previous implementation based on level sets, triangle meshes yield more accurate results, while requiring less memory and computation time. Meshes are also directly compatible with triangle-based rendering algorithms, so no additional post-processing is required. Bastian Goldlücke, Marcus A. Magnor |
ICIP (1) | 1 |
| 2004 | Space-Time Isosurface Evolution for Temporally Coherent 3D Reconstruction
Bastian Goldlücke, Marcus A. Magnor |
CVPR (1) | 1 |
| 2004 | Weighted Minimal Hypersurfaces and Their Applications in Computer Vision
Bastian Goldlücke, Marcus A. Magnor |
ECCV (2) | 1 |
| 2003 | Joint 3D-Reconstruction and Background Separation in Multiple Views using Graph CutsabstractThis paper deals with simultaneous depth map estimation and background separation in a multi-view setting with several fixed calibrated cameras, two problems which have previously been addressed separately. We demonstrate that their strong interdependency can be exploited elegantly by minimizing a discrete energy functional, which evaluates both properties at the same time. Our algorithm is derived from the powerful "multi-camera scene reconstruction via graph cuts" algorithm presented by Kolmogorov and Zabih (2002). Experiments with both real-world as well as synthetic scenes demonstrate that the presented combined approach yields even more correct depth estimates. In particular, the additional information gained by taking background into account increases considerably the algorithm's robustness against noise. Bastian Goldlücke, Marcus A. Magnor |
CVPR (1) | 1 |
| 2003 | Real-time microfacet billboarding for free-viewpoint video renderingabstractWe present a hardware-accelerated method for video-based rendering relying on an approximate model of scene geometry. Our goal is to render high-quality views of the scene from arbitrary viewpoints in real-time, using as input synchronized video-streams from only a small number of calibrated cameras distributed around the scene. Despite the fact that only a very coarse geometry reconstruction is possible, our rendering approach based on textured billboards preserves the details present in the source images. By exploiting multi-texturing and blending facilities of modern graphics cards, we achieve real-time frame-rates on current off-the-shelf hardware. One of the possible applications of our algorithm is to use it in an inexpensive system to display 3D-videos, which the user can watch from an interactively chosen viewpoint, or ultimately even live 3D-television. Bastian Goldlücke, Marcus A. Magnor |
ICIP (3) | 1 |
| 2003 | Robust depth estimation from multiple video streams for dynamic light field rendering
Bastian Goldlücke, Marcus A. Magnor |
SIGGRAPH | 1 |
| 2003 | Real-time free-viewpoint video rendering from volumetric geometry
Bastian Goldlücke, Marcus A. Magnor |
VCIP | 1 |