VLDB 2026 Research / reviewers in the wild / expert
Stefan Gumhold
dblp:g/StefanGumhold
· DBLP profile ↗
60ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-2467-5734ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 14 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 2 · 1 first-authorSystems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixed Presence in Mixed Reality: Charting the Challenges and OpportunitiesabstractThis paper investigates the challenges of designing mixed-presence environments for Mixed Reality and suggests future research directions derived from an expert workshop. Developing mixed-presence systems is a complex undertaking that combines the intricacies of both co-located and distributed mixed-reality spaces. Current literature in this field describes various promising design and development approaches but lacks a systematic overview, resulting in fragmented solutions to re-occurring challenges. Therefore, we conducted a comprehensive review of mixed-presence and multi-user remote mixed-reality systems, categorizing the prevalent challenges faced during the development of such systems, but also current trends, common use cases, study tasks and methodologies. Supported by these results, we then conducted an expert ideation workshop to collect and structure promising future research directions. As a result, we provide a detailed resource to orient and prepare developers for probable challenges and support researchers in making informed design decisions for future mixed-presence studies in Mixed Reality. Katja Krug, Wolfgang Büschel, Marc Satkowski, Stefan Gumhold, Raimund Dachselt |
CHI | 4 |
| 2026 | Beyond Links: Exploring Visual Representations of Multi-View Relations in Mixed RealityabstractThis paper investigates associations, explicit representations of relations between multiple views in Mixed Reality (MR). While research on 2D desktop environments offers extensive recommendations for communicating relations between multiple views, MR environments lack such systematic guidance, necessitating adapted solutions that consider their spatial affordances. To address this gap, we systematically explored association techniques in existing research. Building on established 2D multi-view literature and refining insights from prior design principles, we developed a codebook to describe view relations and their representations. Applying it to a corpus of 44 immersive multi-view approaches, we identified recurring design strategies and synthesized them into a design space of visual association techniques adapted for immersive contexts. Based on a lightweight prototyping framework, we validate the utility of the design space through three envisioning scenarios, demonstrating how associations can support exploration, coordination, and sensemaking in MR applications. Our results inform the design of MR multi-view environments. Weizhou Luo, Rufat Rzayev, Benjamin Russig, Sivanon Visutarporn, Marc Satkowski, Stefan Gumhold, Raimund Dachselt |
CHI | 6 |
| 2026 | PointShopVR: Immersive Authoring of Large Point Clouds in Virtual RealityabstractThis paper presents PointShopVR, an immersive point cloud authoring system in Virtual Reality (VR) that enables intuitive creation, manipulation, and refinement of large 3D point clouds. To support real-time interaction with dense data, our system integrates a simple acceleration data structure and a continuous level-of-detail (CLOD) rendering, ensuring high frame rates in VR. PointShopVR provides a handful of authoring operations—including point addition, deletion, labeling, copy-paste, translation and deformation—each offering instant visual feedback for seamless editing. We evaluate PointShopVR through a user study in which users complete diverse editing tasks within minutes, demonstrating both the usability and effectiveness of immersive point cloud authoring. Compared to traditional desktop-based tools, PointShopVR offers enhanced accessibility, natural interaction, and immersive feedback, making it a powerful platform for large-scale 3D data exploration and creative editing in VR. Tianfang Lin, Matthew McGinity, Stefan Gumhold |
VR | 3 |
| 2026 | EASE: Parametric garment design with explicit and local ease control
Kristijan Bartol, Frieda Hentschel, Nataliya Sadretdinova, Benjamin Russig, Melinos Averkiou, Yordan Kyosev, Stefan Gumhold |
Comput. Graph. | 7 |
| 2026 | Tubes or Ribbons? Comparing Texture-space Visualization for Multivariate Line DataabstractAbstract Multivariate line data is critical for analyzing flow fields, agent systems, and dynamic trajectories. Embedding secondary variables along spatial paths using surface‐based primitives such as ribbons and circular tubes introduces challenges related to perspective, scale, and distortion. Perceptual trade‐offs due to these challenges remain unclear. We address this gap through a controlled user study with 10 experts in computational fluid dynamics and visualization, performing four identical analysis tasks involving both spatial (requiring location‐based relationships) and non‐spatial (requiring attribute value comparisons) aspects. The tasks were performed using a prototype with interactivity limited to controlling the camera. While quantitative measures showed no significant performance differences between embedded visualizations on ribbons and tubes, we found a clear subjective preference for tubes among the study participants. Furthermore, their feedback indicates surface‐based embeddings are generally helpful and should be utilized more often. Benjamin Russig, Rufat Rzayev, Raimund Dachselt, Stefan Gumhold |
Comput. Graph. Forum | 4 |
| 2025 | IGUANA: Immersive Guidance, Navigation, and Control for Consumer UAVabstractAs the markets for unmanned aerial vehicles (UAVs) and mixed reality (MR) headsets continue to grow, recent research has increasingly explored their integration, which enables more intuitive, immersive, and situationally aware control systems. We present IGUANA, an MR-based immersive guidance, navigation, and control system for consumer UAVs. IGUANA introduces three key elements beyond conventional control interfaces: (1) a 3D terrain map interface with draggable waypoint markers and live camera preview for high-level control, (2) a novel spatial control metaphor that uses a virtual ball as a physical analogy for low-level control, and (3) a spatial overlay that helps track the UAV when it is not visible with the naked eye or visual line of sight is interrupted. We conducted a user study to evaluate our design, both quantitatively and qualitatively, and found that (1) the 3D map interface is intuitive and easy to use, relieving users from manual control and suggesting improved accuracy and consistency with lower perceived workload relative to conventional dual-stick controller, (2) the virtual ball interface is intuitive but limited by the lack of physical feedback, and (3) the spatial overlay is very useful in enhancing the users’ situational awareness. Victor Victor, Tania Krisanty, Matthew McGinity, Stefan Gumhold, Uwe Aßmann |
VRST | 4 |
| 2024 | Quantile-Based Maximum Likelihood Training for Outlier DetectionabstractDiscriminative learning effectively predicts true object class for image classification. However, it often results in false positives for outliers, posing critical concerns in applications like autonomous driving and video surveillance systems. Previous attempts to address this challenge involved training image classifiers through contrastive learning using actual outlier data or synthesizing outliers for self-supervised learning. Furthermore, unsupervised generative modeling of inliers in pixel space has shown limited success for outlier detection. In this work, we introduce a quantile-based maximum likelihood objective for learning the inlier distribution to improve the outlier separation during inference. Our approach fits a normalizing flow to pre-trained discriminative features and detects the outliers according to the evaluated log-likelihood. The experimental evaluation demonstrates the effectiveness of our method as it surpasses the performance of the state-of-the-art unsupervised methods for outlier detection. The results are also competitive compared with a recent self-supervised approach for outlier detection. Our work allows to reduce dependency on well-sampled negative training data, which is especially important for domains like medical diagnostics or remote sensing. Masoud Taghikhah, Nishant Kumar 0005, Sinisa Segvic, Abouzar Eslami, Stefan Gumhold |
AAAI | 5 |
| 2024 | Enhanced Plant Phenotyping Through Spatio-Temporal Point Cloud Registration
Somnath Dutta, Benjamin Russig, Stefan Gumhold |
CGI (1) | 3 |
| 2024 | Towards Computational Performance Engineering for Unsupervised Concept Drift Detection: Complexities, Benchmarking, Performance Analysis
Elias Werner, Nishant Kumar 0005, Matthias Lieber, Sunna Torge 0001, Stefan Gumhold, Wolfgang E. Nagel |
DATA | 5 |
| 2024 | An immersive labeling method for large point clouds
Tianfang Lin, Zhongyuan Yu, Matthew McGinity, Stefan Gumhold |
Comput. Graph. | 4 |
| 2023 | Pearl: Physical Environment based Augmented Reality Lenses for In-Situ Human Movement AnalysisabstractThis paper presents Pearl, a mixed-reality approach for the analysis of human movement data in situ. As the physical environment shapes human motion and behavior, the analysis of such motion can benefit from the direct inclusion of the environment in the analytical process. We present methods for exploring movement data in relation to surrounding regions of interest, such as objects, furniture, and architectural elements. We introduce concepts for selecting and filtering data through direct interaction with the environment, and a suite of visualizations for revealing aggregated and emergent spatial and temporal relations. More sophisticated analysis is supported through complex queries comprising multiple regions of interest. To illustrate the potential of Pearl, we developed an Augmented Reality-based prototype and conducted expert review sessions and scenario walkthroughs in a simulated exhibition. Our contribution lays the foundation for leveraging the physical environment in the in-situ analysis of movement data. Weizhou Luo, Zhongyuan Yu, Rufat Rzayev, Marc Satkowski, Stefan Gumhold, Matthew McGinity, Raimund Dachselt |
CHI | 5 |
| 2023 | TmoTA: Simple, Highly Responsive Tool for Multiple Object Tracking AnnotationabstractMachine learning is applied in a multitude of sectors with very impressive results. This success is due to the availability of an ever-growing amount of data acquired by omnipresent sensor devices and platforms on the internet. But there is a scarcity of labeled data which is required for most ML methods. However, generation of labeled data requires much time and resources. In this paper, we propose a portable, Open Source, simple and responsive manual Tool for 2D multiple object Tracking Annotation (TmoTA). Besides responsiveness, our tool design provides several features like view centering and looped playback that speed up the annotation process. We evaluate our proposed tool by comparing TmoTA with the widely used manual labeling tools CVAT, Label Studio, and two semi-automated tools Supervisely and VATIC with respect to object labeling time and accuracy. The evaluation includes a user study and pre-case studies showing that the annotation time per object frame can be reduced by 20% to 40% over the first 20 annotated objects compared to the manual labeling tools. Marzan Tasnim Oyshi, Sebastian Vogt 0002, Stefan Gumhold |
CHI | 3 |
| 2023 | Normalizing Flow based Feature Synthesis for Outlier-Aware Object DetectionabstractReal-world deployment of reliable object detectors is crucial for applications such as autonomous driving. However, general-purpose object detectors like Faster R-CNN are prone to providing overconfident predictions for outlier objects. Recent outlier-aware object detection approaches estimate the density of instance-wide features with class-conditional Gaussians and train on synthesized outlier features from their low-likelihood regions. However, this strategy does not guarantee that the synthesized outlier features will have a low likelihood according to the other class-conditional Gaussians. We propose a novel outlier-aware object detection framework that distinguishes outliers from inlier objects by learning the joint data distribution of all inlier classes with an invertible normalizing flow. The appropriate sampling of the flow model ensures that the synthesized outliers have a lower likelihood than inliers of all object classes, thereby modeling a better decision boundary between inlier and outlier objects. Our approach significantly outperforms the state-of-the-art for outlier-aware object detection on both image and video datasets. Nishant Kumar 0005, Sinisa Segvic, Abouzar Eslami, Stefan Gumhold |
CVPR | 4 |
| 2023 | Efficient Raycasting of Volumetric Depth Images for Remote Visualization of Large Volumes at High Frame RatesabstractWe present an efficient raycasting algorithm for rendering Volumetric Depth Images (VDIs), and we show how it can be used in a remote visualization setting with VDIs generated and streamed from a remote server. VDIs are compact view-dependent volume representations that enable interactive visualization of large volumes at high frame rates by decoupling viewpoint changes from expensive rendering calculations. However, current rendering approaches for VDIs struggle with achieving interactive frame rates at high image resolutions. Here, we exploit the properties of perspective projection to simplify intersections of rays with the view-dependent frustums in a VDI and leverage spatial smoothness in the volume data to minimize memory accesses. Benchmarks show that responsive frame rates can be achieved close to the viewpoint of generation for HD display resolutions, providing high-fidelity approximate renderings of Gigabyte-sized volumes. We also propose a method to subsample the VDI for preview rendering, maintaining high frame rates even for large viewpoint deviations. We provide our implementation as an extension of an established open-source visualization library. Aryaman Gupta, Ulrik Günther, Pietro Incardona, Guido Reina, Steffen Frey, Stefan Gumhold, Ivo F. Sbalzarini |
PacificVis | 6 |
| 2023 | On-Tube Attribute Visualization for Multivariate Trajectory DataabstractStylized tubes are an established visualization primitive for line data as encountered in many scientific fields, ranging from characteristic lines in flow fields, fiber tracks reconstructed from diffusion tensor imaging, to trajectories of moving objects as they arise from cyber-physical systems in many engineering disciplines. Typical challenges include large data set sizes demanding for efficient rendering techniques as well as a large number of attributes that cannot be mapped simultaneously to the basic visual attributes provided by a tube-based visualization. In this work, we tackle both challenges with a new on-tube visualization approach. We improve recent work on high-quality GPU ray casting of Hermite spline tubes supporting ambient occlusion and extend it by a new layered procedural texturing technique. In the proposed framework, a large number of data set attributes can be mapped simultaneously to a variety of glyphs and plots that are embedded in texture space and organized in layers. Efficient rendering with minimal data transfer is achieved by generating the glyphs procedurally and drawing them in a deferred shading pass. We integrated these techniques in a prototype visualization tool that facilitates flexible mapping of data set attributes to visual tube and glyph attributes. We studied our approach on a variety of example data from different fields and found it to provide a highly adaptable and extensible toolbox to quickly craft tailor-made tube-based trajectory visualizations. Benjamin Russig, David Groß, Raimund Dachselt, Stefan Gumhold |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Enhancing Fairness of Visual Attribute Predictors
Tobias Hänel, Nishant Kumar 0005, Dmitrij Schlesinger, Mengze Li 0002, Erdem Ünal, Abouzar Eslami, Stefan Gumhold |
ACCV (6) | 7 |
| 2022 | Teaching Distributed and Heterogeneous Robotic CellsabstractTeaching heterogeneous robotic cells is difficult and time-consuming. We present a context-aware robot teaching solution based on an open-source framework that decouples the task instruction from the task execution to simplify the teaching workflow significantly. To demonstrate the advantages of the solution, a human operator uses a VR headset and controllers to teach a virtual robot to perform a task, e.g., pick-and-place in a virtual world. The task is automatically adapted to different settings of operation cells and executed by physical robots located at different geographical locations. Johannes Mey, Sebastian Ebert, Tianfang Lin, Giang T. Nguyen 0002, Stefan Gumhold, Uwe Aßmann |
CCNC | 5 |
| 2022 | Glyph-Based Visual Analysis of Q-Leaning Based Action Policy Ensembles on RacetrackabstractRecently, deep reinforcement learning has become very successful in making complex decisions, achieving super-human performance in Go, chess, and challenging video games. When applied to safety-critical applications, however, like the control of cyber-physical systems with a learned action policy, the need for certification arises. To empower domain experts to decide whether to trust a learned action policy, we propose visualization methods for a detailed assessment of action policies implemented as neural networks trained with Q-learning. We propose a highly responsive visual analysis tool that fosters efficient analysis of Q-learning based action policies over the complete state space of the system, which is essential for verification and gaining detailed insights on policy quality. For efficient visual inspection of the per-action Q-value rating over the state space, we designed three glyphs that provide different levels of detail. In particular, we introduce the two-dimensional Q-Glyph that visually encodes Q-values in a compact manner while preserving directional information of the actions. Placing glyphs in ordered stacks allows for simultaneous inspection of policy ensembles, that for example result from Q-learning meta parameter studies. Further analysis of the policy is supported by enabling inspection of individual traces generated from a chosen start state. A user study was conducted to evaluate the effectiveness of our tool applied to the Racetrack case study, which is a commonly used benchmark in the AI community abstracting driving control. David Groß, Michaela Klauck, Timo P. Gros, Marcel Steinmetz, Jörg Hoffmann 0001, Stefan Gumhold |
IV | 6 |
| 2022 | Real-Time Visualization of Stream-Based Monitoring DataabstractAbstract Stream-based runtime monitors are used in safety-critical applications such as Unmanned Aerial Systems (UAS) to compute comprehensive statistics and logical assessments of system health that provide the human operator with critical information in hand-over situations. In such applications, a visual display of the monitoring data can be much more helpful than the textual alerts provided by a more traditional user interface. This visualization requires extensive real-time data processing, which includes the synchronization of data from different streams, filtering and aggregation, and priorization and management of user attention. We present a visualization approach for theRTLolamonitoring framework. Our approach is based on the principle that the necessary data processing is the responsibility of the monitor itself, rather than the responsibility of some external visualization tool. We show how the various aspects of the data transformation can be described asRTLolastream equations and linked to the visualization component through a bidirectional synchronous interface. In our experience, this approach leads to highly informative visualizations as well as to understandable and easily maintainable monitoring code. Jan Baumeister, Bernd Finkbeiner, Stefan Gumhold, Malte Schledjewski |
RV | 3 |
| 2022 | Understanding multi-modal brain network data: An immersive 3D visualization approach
Britta Pester, Benjamin Russig, Oliver Winke, Carolin Ligges, Raimund Dachselt, Stefan Gumhold |
Comput. Graph. | 6 |
| 2021 | Advanced Rendering of Line Data with Ambient Occlusion and Transparencyabstract3D Lines are a widespread rendering primitive for the visualization of data from research fields like fluid dynamics or fiber tractography. Global illumination effects and transparent rendering improve the perception of three-dimensional features and decrease occlusion within the data set, thus enabling better understanding of complex line data. We present an efficient approach for high quality GPU-based rendering of line data with ambient occlusion and transparency effects. Our approach builds on GPU-based raycasting of rounded cones, which are geometric primitives similar to truncated cones, but with spherical endcaps. Object space ambient occlusion is provided by an efficient voxel cone tracing approach. Our core contribution is a new fragment visibility sorting strategy that allows for interactive visualization of line data sets with millions of line segments. We improve performance further by exploiting hierarchical opacity maps. David Groß, Stefan Gumhold |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Reinforced Feature Points: Optimizing Feature Detection and Description for a High-Level TaskabstractWe address a core problem of computer vision: Detection and description of 2D feature points for image matching. For a long time, hand-crafted designs, like the seminal SIFT algorithm, were unsurpassed in accuracy and efficiency. Recently, learned feature detectors emerged that implement detection and description using neural networks. Training these networks usually resorts to optimizing low-level matching scores, often pre-defining sets of image patches which should or should not match, or which should or should not contain key points. Unfortunately, increased accuracy for these low-level matching scores does not necessarily translate to better performance in high-level vision tasks. We propose a new training methodology which embeds the feature detector in a complete vision pipeline, and where the learnable parameters are trained in an end-to-end fashion. We overcome the discrete nature of key point selection and descriptor matching using principles from reinforcement learning. As an example, we address the task of relative pose estimation between a pair of images. We demonstrate that the accuracy of a state-of-the-art learning-based feature detector can be increased when trained for the task it is supposed to solve at test time. Our training methodology poses little restrictions on the task to learn, and works for any architecture which predicts key point heat maps, and descriptors for key point locations. Aritra Bhowmik, Stefan Gumhold, Carsten Rother, Eric Brachmann |
CVPR | 2 |
| 2020 | TraceVis: Towards Visualization for Deep Statistical Model Checking
Timo P. Gros, David Groß, Stefan Gumhold, Jörg Hoffmann 0001, Michaela Klauck, Marcel Steinmetz |
ISoLA (4) | 3 |
| 2020 | Detection and Visualization of Splat and Antisplat Events in Turbulent FlowsabstractSplat and antisplat events are a widely found phenomenon in three-dimensional turbulent flow fields. Splats are observed when fluid locally impinges on an impermeable surface transferring energy from the normal component to the tangential velocity components, while antisplats relate to the inverted situation. These events affect a variety of flow properties, such as the transfer of kinetic energy between velocity components and the transfer of heat, so that their investigation can provide new insight into these issues. Here, we propose the first Lagrangian method for the detection of splats and antisplats as features of an unsteady flow field. Our method utilizes the concept of strain tensors on flow-embedded flat surfaces to extract disjoint regions in which splat and antisplat events of arbitrary scale occur. We validate the method with artificial flow fields of increasing complexity. Subsequently, the method is used to analyze application data stemming from a direct numerical simulation of the turbulent flow over a backward facing step. Our results show that splat and antisplat events can be identified efficiently and reliably even in such a complex situation, demonstrating that the new method constitutes a well-suited tool for the analysis of turbulent flows. Baldwin Nsonga, Martin Niemann, Jochen Fröhlich, Joachim Staib, Stefan Gumhold, Gerik Scheuermann |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Analysis of the Near-Wall Flow in a Turbine Cascade by Splat VisualizationabstractTurbines are essential components of jet planes and power plants. Therefore, their efficiency and service life are of central engineering interest. In the case of jet planes or thermal power plants, the heating of the turbines due to the hot gas flow is critical. Besides effective cooling, it is a major goal of engineers to minimize heat transfer between gas flow and turbine by design. Since it is known that splat events have a substantial impact on the heat transfer between flow and immersed surfaces, we adapt a splat detection and visualization method to a turbine cascade simulation in this case study. Because splat events are small phenomena, we use a direct numerical simulation resolving the turbulence in the flow as the base of our analysis. The outcome shows promising insights into splat formation and its relation to vortex structures. This may lead to better turbine design in the future. Baldwin Nsonga, Gerik Scheuermann, Stefan Gumhold, Jordi Ventosa-Molina, Denis Koschichow, Jochen Fröhlich |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2019 | Learning to Think Outside the Box: Wide-Baseline Light Field Depth Estimation with EPI-ShiftabstractWe propose a method for depth estimation from light field data, based on a fully convolutional neural network architecture. Our goal is to design a pipeline which achieves highly accurate results for small-and wide-baseline light fields. Since light field training data is scarce, all learning-based approaches use a small receptive field and operate on small disparity ranges. In order to work with wide-baseline light fields, we introduce the idea of EPI-Shift: To virtually shift the light field stack which enables to retain a small receptive field, independent of the disparity range. In this way, our approach "learns to think outside the box of the receptive field"'. Our network performs joint classification of integer disparities and regression of disparity-offsets. A U-Net component provides excellent long-range smoothing. EPI-Shift considerably outperforms the state-of-the-art learning-based approaches and is on par with hand-crafted methods. We demonstrate this on a publicly available, synthetic, small-baseline benchmark and on large-baseline real-world recordings. Titus Leistner, Hendrik Schilling, Radek Mackowiak, Stefan Gumhold, Carsten Rother |
3DV | 4 |
| 2018 | Fast Mapping of the Eloquent Cortex by Learning L2 Penalties
Nico Hoffmann, Uwe Petersohn, Gabriele Schackert, Edmund Koch, Stefan Gumhold, Matthias Kirsch |
MICCAI (3) | 5 |
| 2018 | Generalized motorcycle graphs for imperfect quad-dominant meshesabstractWe introduce a practical pipeline to create UV T-layouts for real-world quad dominant semi-regular meshes. Our algorithm creates large rectangular patches by relaxing the notion of motorcycle graphs and making it insensitive to local irregularities in the mesh structure such as non-quad elements, redundant irregular vertices, T-junctions, and others. Each surface patch, which can contain multiple singularities and/or polygonal elements, is mapped to an axis-aligned rectangle, leading to a simple and efficient UV layout, which is ideal for texture mapping (allowing for mipmapping and artifact-free bilinear interpolation). We demonstrate that our algorithm is an ideal solution for both recent semi-regular, quad-dominant meshing methods, and for the low-poly meshes typically used in games and movies. Nico Schertler, Daniele Panozzo, Stefan Gumhold, Marco Tarini |
ACM Trans. Graph. | 3 |
| 2017 | DSAC - Differentiable RANSAC for Camera LocalizationabstractRANSAC is an important algorithm in robust optimization and a central building block for many computer vision applications. In recent years, traditionally hand-crafted pipelines have been replaced by deep learning pipelines, which can be trained in an end-to-end fashion. However, RANSAC has so far not been used as part of such deep learning pipelines, because its hypothesis selection procedure is non-differentiable. In this work, we present two different ways to overcome this limitation. The most promising approach is inspired by reinforcement learning, namely to replace the deterministic hypothesis selection by a probabilistic selection for which we can derive the expected loss w.r.t. to all learnable parameters. We call this approach DSAC, the differentiable counterpart of RANSAC. We apply DSAC to the problem of camera localization, where deep learning has so far failed to improve on traditional approaches. We demonstrate that by directly minimizing the expected loss of the output camera poses, robustly estimated by RANSAC, we achieve an increase in accuracy. In the future, any deep learning pipeline can use DSAC as a robust optimization component. Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel 0002, Stefan Gumhold, Carsten Rother |
CVPR | 6 |
| 2017 | Global Hypothesis Generation for 6D Object Pose EstimationabstractThis paper addresses the task of estimating the 6D-pose of a known 3D object from a single RGB-D image. Most modern approaches solve this task in three steps: i) compute local features, ii) generate a pool of pose-hypotheses, iii) select and refine a pose from the pool. This work focuses on the second step. While all existing approaches generate the hypotheses pool via local reasoning, e.g. RANSAC or Hough-Voting, we are the first to show that global reasoning is beneficial at this stage. In particular, we formulate a novel fully-connected Conditional Random Field (CRF) that outputs a very small number of pose-hypotheses. Despite the potential functions of the CRF being non-Gaussian, we give a new, efficient two-step optimization procedure, with some guarantees for optimality. We utilize our global hypotheses generation procedure to produce results that exceed state-of-the-art for the challenging "Occluded Object Dataset". Frank Michel 0002, Alexander Kirillov, Eric Brachmann, Alexander Krull, Stefan Gumhold, Bogdan Savchynskyy, Carsten Rother |
CVPR | 5 |
| 2017 | Chamber Recognition in Cave Data SetsabstractQuantitative analysis of cave systems represented as 3D models is becoming more and more important in the field of cave sciences. One open question is the rigorous identification of chambers in a data set, which has a deep impact on subsequent analysis steps such as size calculation. This affects the international recognition of a cave since especially record-holding caves bear significant tourist attraction potential. In the past, chambers have been identified manually, without any clear definition or guidance. While experts agree on core parts of chambers in general, their opinions may differ in more controversial areas. Since this process is heavily subjective, it is not suited for objective quantitative comparison of caves. Therefore, we present a novel fully-automatic curve skeleton-based chamber recognition algorithm that has been derived from requirements from field experts. We state the problem as a binary labeling problem on a curve skeleton and find a solution through energy minimization. A thorough evaluation of our results with the help of expert feedback showed that our algorithm matches real-world requirements very closely and is thus suited as the foundation for any quantitative cave analysis system. Nico Schertler, Manfred F. Buchroithner, Stefan Gumhold |
Comput. Graph. Forum | 3 |
| 2017 | Towards Globally Optimal Normal Orientations for Large Point CloudsabstractAbstract Various processing algorithms on point set surfaces rely on consistently oriented normals (e.g. Poisson surface reconstruction). While several approaches exist for the calculation of normal directions, in most cases, their orientation has to be determined in a subsequent step. This paper generalizes propagation‐based approaches by reformulating the task as a graph‐based energy minimization problem. By applying global solvers, we can achieve more consistent orientations than simple greedy optimizations. Furthermore, we present a streaming‐based framework for orienting large point clouds. This framework orients patches locally and generates a globally consistent patch orientation on a reduced neighbour graph, which achieves similar quality to orienting the full graph. Nico Schertler, Bogdan Savchynskyy, Stefan Gumhold |
Comput. Graph. Forum | 3 |
| 2017 | Field-aligned online surface reconstructionabstractToday's 3D scanning pipelines can be classified into two overarching categories: offline, high accuracy methods that rely on global optimization to reconstruct complex scenes with hundreds of millions of samples, and online methods that produce real-time but low-quality output, usually from structure-from-motion or depth sensors. The method proposed in this paper is the first to combine the benefits of both approaches, supporting online reconstruction of scenes with hundreds of millions of samples from high-resolution sensing modalities such as structured light or laser scanners. The key property of our algorithm is that it sidesteps the signed-distance computation of classical reconstruction techniques in favor of direct filtering, parametrization, and mesh and texture extraction. All of these steps can be realized using only weak notions of spatial neighborhoods, which allows for an implementation that scales approximately linearly with the size of each dataset that is integrated into a partial reconstruction. Combined, these algorithmic differences enable a drastically more efficient output-driven interactive scanning and reconstruction workflow, where the user is able to see the final quality field-aligned textured mesh during the entirety of the scanning procedure. Holes or parts with registration problems are displayed in real-time to the user and can be easily resolved by adding further localized scans, or by adjusting the input point cloud using our interactive editing tools with immediate visual feedback on the output mesh. We demonstrate the effectiveness of our algorithm in conjunction with a state-of-the-art structured light scanner and optical tracking system and test it on a large variety of challenging models. Nico Schertler, Marco Tarini, Wenzel Jakob, Michael M. Kazhdan, Stefan Gumhold, Daniele Panozzo |
ACM Trans. Graph. | 5 |
| 2016 | Uncertainty-Driven 6D Pose Estimation of Objects and Scenes from a Single RGB ImageabstractIn recent years, the task of estimating the 6D pose of object instances and complete scenes, i.e. camera localization, from a single input image has received considerable attention. Consumer RGB-D cameras have made this feasible, even for difficult, texture-less objects and scenes. In this work, we show that a single RGB image is sufficient to achieve visually convincing results. Our key concept is to model and exploit the uncertainty of the system at all stages of the processing pipeline. The uncertainty comes in the form of continuous distributions over 3D object coordinates and discrete distributions over object labels. We give three technical contributions. Firstly, we develop a regularized, auto-context regression framework which iteratively reduces uncertainty in object coordinate and object label predictions. Secondly, we introduce an efficient way to marginalize object coordinate distributions over depth. This is necessary to deal with missing depth information. Thirdly, we utilize the distributions over object labels to detect multiple objects simultaneously with a fixed budget of RANSAC hypotheses. We tested our system for object pose estimation and camera localization on commonly used data sets. We see a major improvement over competing systems. Eric Brachmann, Frank Michel 0002, Alexander Krull, Michael Ying Yang, Stefan Gumhold, Carsten Rother |
CVPR | 5 |
| 2016 | Enhancing Scatterplots with Multi-Dimensional Focal BlurabstractAbstract Scatterplots directly depict two dimensions of multi‐dimensional data points, discarding all other information. To visualize all data, these plots are extended to scatterplot matrices, which distribute the information of each data point over many plots. Problems arising from the resulting visual complexity are nowadays alleviated by concepts like filtering and focus and context. We present a method based on depth of field that contains both aspects and injects information from all dimensions into each scatterplot. Our approach is a natural generalization of the commonly known focus effects from optics. It is based on a multidimensional focus selection body. Points outside of this body are defocused depending on their distance. Our method allows for a continuous transition from data points in focus, over regions of blurry points providing contextual information, to visually filtered data. Our algorithm supports different focus selection bodies, blur kernels, and point shapes. We present an optimized GPU‐based implementation for interactive exploration and show the usefulness of our approach on several data sets. Joachim Staib, Sebastian Grottel, Stefan Gumhold |
Comput. Graph. Forum | 3 |
| 2015 | Pose Estimation of Kinematic Chain Instances via Object Coordinate RegressionabstractAccurate pose estimation of object instances is a key aspect in many applications, including augmented reality or robotics. For example, a task of a domestic robot could be to fetch an item from an open drawer. The poses of both, the drawer and the item have to be known by the robot in order to fulfil the task. 6D pose estimation of rigid objects has been addressed with great success in recent years. In large part, this has been due to the advent of consumer-level RGB-D cameras, which provide rich, robust input data. However, the practical use of state-of-the-art pose estimation approaches is limited by the assumption that objects are rigid. In cluttered, domestic environments this assumption does often not hold. Examples are doors, many types of furniture, certain electronic devices and toys. A robot might encounter these items in any state of articulation. This work considers the task of one-shot pose estimation of articulated object instances from an RGB-D image. In particular, we address objects with the topology of a kinematic chain of any length, i.e. objects are composed of a chain of parts interconnected by joints. We restrict joints to either revolute joints with 1 DOF (degrees of freedom) rotational movement or prismatic joints with 1 DOF translational movement. This topology covers a wide range of common objects (see our dataset for examples). However, our approach can easily be expanded to any topology, and to joints with higher degrees of freedom. Frank Michel 0002, Alexander Krull, Eric Brachmann, Michael Ying Yang, Stefan Gumhold, Carsten Rother |
BMVC | 5 |
| 2015 | Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D ImagesabstractAnalysis-by-synthesis has been a successful approach for many tasks in computer vision, such as 6D pose estimation of an object in an RGB-D image which is the topic of this work. The idea is to compare the observation with the output of a forward process, such as a rendered image of the object of interest in a particular pose. Due to occlusion or complicated sensor noise, it can be difficult to perform this comparison in a meaningful way. We propose an approach that "learns to compare", while taking these difficulties into account. This is done by describing the posterior density of a particular object pose with a convolutional neural network (CNN) that compares observed and rendered images. The network is trained with the maximum likelihood paradigm. We observe empirically that the CNN does not specialize to the geometry or appearance of specific objects. It can be used with objects of vastly different shapes and appearances, and in different backgrounds. Compared to state-of-the-art, we demonstrate a significant improvement on two different datasets which include a total of eleven objects, cluttered background, and heavy occlusion. Alexander Krull, Eric Brachmann, Frank Michel 0002, Michael Ying Yang, Stefan Gumhold, Carsten Rother |
ICCV | 5 |
| 2015 | Wifbs: A Web-Based Image Feature Benchmark System
Marcel Spehr, Sebastian Grottel, Stefan Gumhold |
MMM (2) | 3 |
| 2015 | Visualization of Particle-based Data with Transparency and Ambient OcclusionabstractAbstract Particle‐based simulation techniques, like the discrete element method or molecular dynamics, are widely used in many research fields. In real‐time explorative visualization it is common to render the resulting data using opaque spherical glyphs with local lighting only. Due to massive overlaps, however, inner structures of the data are often occluded rendering visual analysis impossible. Furthermore, local lighting is not sufficient as several important features like complex shapes, holes, rifts or filaments cannot be perceived well. To address both problems we present a new technique that jointly supports transparency and ambient occlusion in a consistent illumination model. Our approach is based on the emission‐absorption model of volume rendering. We provide analytic solutions to the volume rendering integral for several density distributions within a spherical glyph. Compared to constant transparency our approach preserves the three‐dimensional impression of the glyphs much better. We approximate ambient illumination with a fast hierarchical voxel cone‐tracing approach, which builds on a new real‐time voxelization of the particle data. Our implementation achieves interactive frame rates for millions of static or dynamic particles without any preprocessing. We illustrate the merits of our method on real‐world data sets gaining several new insights. Joachim Staib, Sebastian Grottel, Stefan Gumhold |
Comput. Graph. Forum | 3 |
| 2014 | 6-DOF Model Based Tracking via Object Coordinate Regression
Alexander Krull, Frank Michel 0002, Eric Brachmann, Stefan Gumhold, Stephan Ihrke, Carsten Rother |
ACCV (4) | 4 |
| 2014 | Learning 6D Object Pose Estimation Using 3D Object Coordinates
Eric Brachmann, Alexander Krull, Frank Michel 0002, Stefan Gumhold, Jamie Shotton, Carsten Rother |
ECCV (2) | 4 |
| 2014 | Visualizing time-dependent key performance indicator in a graph-based analysisabstractThe usage of visual analytics during the analysis of business warehouse calculated key performance indicators is one emerging challenge in modern business applications. On the one hand, a complex network of key performance indicators has to be supervised. On the other hand, within this network only few key performance indicators change obviously within a short period of time. The sole mapping of the complexity of a network of key performance indicators to a graph-based visualization only covers static information and neglects temporal dependencies. We present a new visualization approach for the enrichment of graph-based visualizations of key performance indicator networks by introducing a multi-encoded visualization of additional functional, contextual and temporal information. The should help the user to understand relationships between KPIs and alert him if something is going wrong. Stefan Hesse, Marcel Spehr, Stefan Gumhold, Rainer Groh 0001 |
ETFA | 3 |
| 2014 | Visual Analysis of Trajectories in Multi-Dimensional State SpacesabstractAbstract Multi‐dimensional data originate from many different sources and are relevant for many applications. One specific sub‐type of such data is continuous trajectory data in multi‐dimensional state spaces of complex systems. We adapt the concept of spatially continuous scatterplots and spatially continuous parallel coordinate plots to such trajectory data, leading to continuous‐time scatterplots and continuous‐time parallel coordinates. Together with a temporal heat map representation, we design coordinated views for visual analysis and interactive exploration. We demonstrate the usefulness of our visualization approach for three case studies that cover examples of complex dynamic systems: cyber‐physical systems consisting of heterogeneous sensors and actuators networks (the collection of time‐dependent sensor network data of an exemplary smart home environment), the dynamics of robot arm movement and motion characteristics of humanoids. Sebastian Grottel, Julian Heinrich, Daniel Weiskopf, Stefan Gumhold |
Comput. Graph. Forum | 4 |
| 2013 | Feature propagation on image webs for enhanced image retrievalabstractThe bag-of-features model is often deployed in content-based image retrieval to measure image similarity. In cases where the visual appearance of semantically similar images differs largely, feature histograms mismatch and the model fails. We increase the robustness of feature histograms by automatically augmenting them with features of related images. We establish image relations by image web construction and adapt a label propagation scheme from the domain of semi-supervised learning for feature augmentation. While the benefit of feature augmentation has been shown before, our approach refrains from the use of semantic labels. Instead we show how to increase the performance of the bag-of-features model substantially on a completely unlabeled image corpus. Eric Brachmann, Marcel Spehr, Stefan Gumhold |
ICMR | 3 |
| 2011 | Directed image search with local parallel feature axesabstractThe browsing of large image data bases has become a standard problem not only on the web but also in private photo collections. Most browsing techniques build on high dimensional feature spaces that are reduced to one or two dimensions when presented to the user. As this approach does not scale well with the size of the data base we propose to use an interface based on the concept of parallel coordinates. Around the currently selected image, we collect images for each feature dimension, which vary only in this feature coordinate, and present them to the user in a row by row fashion. In this way the user can understand the individual feature dimensions independently. Besides directed image search the interface is well suited to explore classes of images and furthermore to evaluate how intuitive individual feature dimensions are for the user. Stefan Gumhold, Marcel Spehr |
ICME | 1 |
| 2011 | Diffusion-Based Snow Cover GenerationabstractAbstract We present a method to generate snow covers on complex scene geometries. Both volumetric snow shapes and photorealistic texturing are computed. We formulate snow accumulation as a diffusive distribution process on a ground scene. Our theoretical framework is motivated by models for granular material deposition. With the framework we can capture the most relevant features of natural snow cover geometries in a concise local computation scheme. Snow bridges and overhangs are also included. Snow surface texture coordinates are computed to create realistic ground–snow interfaces. Several example scenes and a supplementary snow cover growth animation demonstrate the method's efficiency. Niels von Festenberg, Stefan Gumhold |
Comput. Graph. Forum | 2 |
| 2006 | Streaming compression of tetrahedral volume meshes
Martin Isenburg, Peter Lindstrom 0001, Stefan Gumhold, Jonathan Richard Shewchuk |
Graphics Interface | 3 |
| 2005 | Optimizing markov models with applications to triangular connectivity coding
Stefan Gumhold |
SODA | 1 |
| 2005 | Visualization with stylized line primitivesabstractLine primitives are a very powerful visual attribute used for scientific visualization and in particular for 3D vector-field visualization. We extend the basic line primitives with additional visual attributes including color, line width, texture and orientation. To implement the visual attributes we represent the stylized line primitives as generalized cylinders. One important contribution of our work is an efficient rendering algorithm for stylized lines, which is hybrid in the sense that it uses both CPU and GPU based rendering. We improve the depth perception with a shadow algorithm. We present several applications for the visualization with stylized lines among which are the visualizations of 3D vector fields and molecular structures. Carsten Stoll, Stefan Gumhold, Hans-Peter Seidel |
IEEE Visualization | 2 |
| 2005 | Truly selective polygonal mesh hierarchies with error control
Stefan Gumhold |
Comput. Aided Geom. Des. | 1 |
| 2005 | Mesh segmentation driven by Gaussian curvature
Hitoshi Yamauchi, Stefan Gumhold, Rhaleb Zayer, Hans-Peter Seidel |
Vis. Comput. | 2 |
| 2004 | Introduction to situation and task awareness computing
Stefan Gumhold, Stefan Noll |
Comput. Graph. | 1 |
| 2003 | Higher Order Prediction for Geometry CompressionabstractA lot of techniques have been developed for the encoding of triangular meshes as this is a widely used representation for the description of surface models. Although methods for the encoding of the neighbor information, the connectivity, are near optimal, there is still room for better en-codings of vertex locations, the geometry. Our geometry encoding strategy follows the predictive coding paradigm, which is based on a region growing encoding order. Only the delta vectors between original and predicted locations are encoded in a local coordinate system, which splits into two tangential and one normal component. In this paper we introduce so-called higher order prediction for an improved encoding of the normal component. We first encode the tangential components with parallelogram prediction. Then we fit a higher order surface to the so far encoded geometry. As we encode the normal component as a bending angle, it is found by intersecting the higher order surface with the circle defined by the tangential components. Experimental results show that our strategy allows saving one bit per vertex for the normal component independent of the tangential prediction rule used. Stefan Gumhold, Rachida Amjoun |
Shape Modeling International | 1 |
| 2003 | Large Mesh Simplification using Processing SequencesabstractIn this paper we show how out-of-core mesh processing techniques can be adapted to perform their computations based on the new processing sequence paradigm (Isenburg, et al., 2003), using mesh simplification as an example. We believe that this processing concept will also prove useful for other tasks, such a parameterization, remeshing, or smoothing, for which currently only in-core solutions exist. A processing sequence represents a mesh as a particular interleaved ordering of indexed triangles and vertices. This representation allows streaming very large meshes through main memory while maintaining information about the visitation status of edges and vertices. At any time, only a small portion of the mesh is kept in-core, with the bulk of the mesh data residing on disk. Mesh access is restricted to a fixed traversal order, but full connectivity and geometry information is available for the active elements of the traversal. This provides seamless and highly efficient out-of-core access to very large meshes for algorithms that can adapt their computations to this fixed ordering. The two abstractions that are naturally supported by this representation are boundary-based and buffer-based processing. We illustrate both abstractions by adapting two different simplification methods to perform their computation using a prototype of our mesh processing sequence API. Both algorithms benefit from using processing sequences in terms of improved quality, more efficient execution, and smaller memory footprints. Martin Isenburg, Peter Lindstrom 0001, Stefan Gumhold, Jack Snoeyink |
IEEE Visualization | 3 |
| 2003 | Out-of-core compression for gigantic polygon meshesabstractPolygonal models acquired with emerging 3D scanning technology or from large scale CAD applications easily reach sizes of several gigabytes and do not fit in the address space of common 32-bit desktop PCs. In this paper we propose an out-of-core mesh compression technique that converts such gigantic meshes into a streamable, highly compressed representation. During decompression only a small portion of the mesh needs to be kept in memory at any time. As full connectivity information is available along the decompression boundaries, this provides seamless mesh access for incremental in-core processing on gigantic meshes. Decompression speeds are CPU-limited and exceed one million vertices and two million triangles per second on a 1.8 GHz Athlon processor.A novel external memory data structure provides our compression engine with transparent access to arbitrary large meshes. This out-of-core mesh was designed to accommodate the access pattern of our region-growing based compressor, which - in return - performs mesh queries as seldom and as local as possible by remembering previous queries as long as needed and by adapting its traversal slightly. The achieved compression rates are state-of-the-art. Martin Isenburg, Stefan Gumhold |
ACM Trans. Graph. | 2 |
| 2002 | Maximum Entropy Light Source PlacementabstractFinding the "best" viewing parameters for a scene is a difficult but very important problem. Fully automatic procedures seem to be impossible as the notion of "best" strongly depends on human judgment as well as on the application. In this paper a solution to the sub-problem of placing light sources for given camera parameters is proposed. A light position is defined to be optimal, when the resulting illumination reveals more about the scene than illuminations from all other light positions, i.e. the light position maximizes information that is added to the image through the illumination. With the help of an experiment with several subjects we could adapt the information measure to the actually perceived information content. We present fast global optimization procedures and solutions for two and more light sources. Stefan Gumhold |
IEEE Visualization | 1 |
| 2001 | The connectivity shapes videoabstractIn this video we introduce a 3D shape representation that is based sol ely on mesh connectivity -- the {\em connectivity shape}. Given a connectivity, we define its natural geometry as a smooth embedding in space with uniform edge lengths and describe efficient techniques to compute it. Furthermore, we show how to generate connectivity shapes that approximate given shapes. The details will soon be published in form of a full paper. Martin Isenburg, Stefan Gumhold, Craig Gotsman |
SCG | 2 |
| 2001 | Connectivity ShapesabstractWe describe a method to visualize the connectivity graph of a mesh using a natural embedding in 3D space. This uses a 3D shape representation that is based solely on mesh connectivity: the connectivity shape. Given a connectivity, we define its natural geometry as a smooth embedding in space with uniform edge lengths and describe efficient techniques to compute it. Our main contribution is to demonstrate that a surprising amount of geometric information is implicit in the connectivity. We also show how to generate connectivity shapes that approximate given 3D shapes. Potential applications of connectivity shapes to modeling and mesh coding are described. Martin Isenburg, Stefan Gumhold, Craig Gotsman |
IEEE Visualization | 2 |
| 1999 | Tetrahedral Mesh Compression with the Cut-Border MachineabstractIn recent years, substantial progress has been achieved in the area of volume visualization on irregular grids, which is mainly based on tetrahedral meshes. Even moderately fine tetrahedral meshes consume several mega-bytes of storage. For archivation and transmission compression algorithms are essential. In scientific applications lossless compression schemes are of primary interest. This paper introduces a new lossless compression scheme for the connectivity of tetrahedral meshes. Our technique can handle all tetrahedral meshes in three dimensional euclidean space even with non manifold border. We present compression and decompression algorithms which consume for reasonable meshes linear time in the number of tetrahedra. The connectivity is compressed to less than 2.4 bits per tetrahedron for all measured meshes. Thus a tetrahedral mesh can almost be reduced to the vertex coordinates, which consume in a common representation about one quarter of the total storage space. We complete our work with solutions for the compression of vertex coordinates and additional attributes, which might be attached to the mesh. Stefan Gumhold, Stefan Guthe, Wolfgang Straßer |
IEEE Visualization | 1 |
| 1998 | Real Time Compression of Triangle Mesh ConnectivityabstractIn this paper we introduce a new compressed representation for the connectivity of a triangle mesh. We present local compression and decompression algorithms which are fast enough for real time applications. The achieved space compression rates keep pace with the best rates reported for any known global compression algorithm. These nice properties have great benefits for several important applications. Naturally, the technique can be used to compress triangle meshes without significant delay before they are stored on external devices or transmitted over a network. The presented decompression algorithm is very simple allowing a possible hardware realization of the decompression algorithm which could significantly increase the rendering speed of pipelined graphics hardware. CR Categories: I.3.1 [Computer Graphics]: Hardware Architecture; I.3.3 [Computer Graphics]: Picture/Image Generation--- Display algorithms Keywords: Mesh Compression, Algorithms, 3D Graphics Hardware, Graphics 1 In... Stefan Gumhold, Wolfgang Straßer |
SIGGRAPH | 1 |