VLDB 2026 Research / reviewers in the wild / expert
Sebastian Hartwig
dblp:239/4317
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-8642-2789ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance SegmentationabstractTraditionally, algorithms that learn to segment object instances in 2D images have heavily relied on large amounts of human-annotated data. Only recently, novel approaches have emerged tackling this problem in an unsupervised fashion. Generally, these approaches first generate pseudo-masks and then train a class-agnostic detector. While such methods deliver the current state of the art, they often fail to correctly separate instances overlapping in 2D image space since only semantics are considered. To tackle this issue, we instead propose to cut the semantic masks in 3D to obtain the final 2D instances by utilizing a point cloud representation of the scene. Furthermore, we derive a Spatial Importance function, which we use to resharpen the semantics along the 3D borders of instances. Nevertheless, these pseudo-masks are still subject to mask ambiguity. To address this issue, we further propose to augment the training of a class-agnostic detector with three Spatial Confidence components aiming to isolate a clean learning signal. With these contributions, our approach outperforms competing methods across multiple standard benchmarks for unsupervised instance segmentation and object detection. Leon Sick, Dominik Engel 0001, Sebastian Hartwig, Pedro Hermosilla, Timo Ropinski |
ICCV | 3 |
| 2025 | HPSCAN: Human Perception-Based Scattered Data ClusteringabstractAbstract Cluster separation is a task typically tackled by widely used clustering techniques, such as k‐means or DBSCAN. However, these algorithms are based on non‐perceptual metrics, and our experiments demonstrate that their output does not reflect human cluster perception. To bridge the gap between human cluster perception and machine‐computed clusters, we propose HPSCAN, a learning strategy that operates directly on scattered data. To learn perceptual cluster separation on such data, we crowdsourced the labeling of bivariate (scatterplot) datasets to 384 human participants. We train our HPSCAN model on these human‐annotated data. Instead of rendering these data as scatterplot images, we used their x and y point coordinates as input to a modified PointNet++ architecture, enabling direct inference on point clouds. In this work, we provide details on how we collected our dataset, report statistics of the resulting annotations, and investigate the perceptual agreement of cluster separation for real‐world data. We also report the training and evaluation protocol for HPSCAN and introduce a novel metric, that measures the accuracy between a clustering technique and a group of human annotators. We explore predicting point‐wise human agreement to detect ambiguities. Finally, we compare our approach to 10 established clustering techniques and demonstrate that HPSCAN is capable of generalizing to unseen and out‐of‐scope data. Sebastian Hartwig, Christian van Onzenoodt, Dominik Engel 0001, Pedro Hermosilla, Timo Ropinski |
Comput. Graph. Forum | 1 |
| 2025 | A Survey on Quality Metrics for Text-to-Image GenerationabstractAI-based text-to-image models do not only excel at generating realistic images, they also give designers more and more fine-grained control over the image content. Consequently, these approaches have gathered increased attention within the computer graphics research community, which has been historically devoted towards traditional rendering techniques, that offer precise control over scene parameters (e.g., objects, materials, and lighting). While the quality of conventionally rendered images is assessed through well established image quality metrics, such as SSIM or PSNR, the unique challenges of text-to-image generation require other, dedicated quality metrics. These metrics must be able to not only measure overall image quality, but also how well images reflect given text prompts, whereby the control of scene and rendering parameters is interweaved. Within this survey, we provide a comprehensive overview of such text-to-image quality metrics, and propose a taxonomy to categorize these metrics. Our taxonomy is grounded in the assumption, that there are two main quality criteria, namely compositional quality and general quality, that contribute to the overall image quality. Besides the metrics, this survey covers dedicated text-to-image benchmark datasets, over which the metrics are frequently computed. Finally, we identify limitations and open challenges in the field of text-to-image generation, and derive guidelines for practitioners conducting text-to-image evaluation. Sebastian Hartwig, Dominik Engel 0001, Leon Sick, Hannah Kniesel, Tristan Payer, Poonam Poonam, Michael Glöckler, Alex Bäuerle, Timo Ropinski |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | PanoramaViewer - A Framework for Educational Collaborative Virtual Field TripsabstractVirtual field trips (VFTs), which use 360° technologies in particular, have developed into powerful learning tools in recent years. Compared to real field trips, VFTs save time and effort, and compared to other media such as books or websites, VFTs offer a much more authentic learning experience by means of greater visual detail. At present, however, VFTs are typically only based on single-user applications. As a result, the potential positive effects arising from the collaborative aspects of a joint VFTs are not fully exploited. Multi-user environments for VFTs for multiple participants are only occasionally mentioned in the literature. Accordingly, we have implemented a multi-user framework for 360°-based VFTs. In this paper, we describe the basic requirements for such a framework, the technical implementation framework and the results of an evaluation study of a prototypical VFT to a recently built stormwater overflow basin in Germany. As part of the evaluation study, data on presence, motivation, emotion and cognitive load was collected from the study participants (N=7) by using a questionnaire based on standardized instruments. In addition, qualitative feedback and a semi-structured interview with the lecturer were analyzed. Despite its prototypical state, the framework shows its potential by achieving acceptable values for the learning requirements surveyed and leading to positive assessments overall from both lecturers and participants. Accordingly, the study results might be viewed in favor of further development of the framework in order to allow the learning-relevant effects of collaborative VFTs to be investigated further. Mario Wolf, Sebastian Hartwig, Gregor Steinhöfel, Heinrich Söbke, Eckhard Kraft |
ISM | 2 |
| 2024 | Monocular Depth Decomposition of Semi-Transparent Volume RenderingsabstractNeural networks have shown great success in extracting geometric information from color images. Especially, monocular depth estimation networks are increasingly reliable in real-world scenes. In this work we investigate the applicability of such monocular depth estimation networks to semi-transparent volume rendered images. As depth is notoriously difficult to define in a volumetric scene without clearly defined surfaces, we consider different depth computations that have emerged in practice, and compare state-of-the-art monocular depth estimation approaches for these different interpretations during an evaluation considering different degrees of opacity in the renderings. Additionally, we investigate how these networks can be extended to further obtain color and opacity information, in order to create a layered representation of the scene based on a single color image. This layered representation consists of spatially separated semi-transparent intervals that composite to the original input rendering. In our experiments we show that existing approaches to monocular depth estimation can be adapted to perform well on semi-transparent volume renderings, which has several applications in the area of scientific visualization, like re-composition with additional objects and labels or additional shading. Dominik Engel 0001, Sebastian Hartwig, Timo Ropinski |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Learning Human Viewpoint Preferences from Sparsely Annotated ModelsabstractAbstract View quality measures compute scores for given views and are used to determine an optimal view in viewpoint selection tasks. Unfortunately, despite the wide adoption of these measures, they are rather based on computational quantities, such as entropy, than human preferences. To instead tailor viewpoint measures towards humans, view quality measures need to be able to capture human viewpoint preferences. Therefore, we introduce a large‐scale crowdsourced data set, which contains 58k annotated viewpoints for 3220 ModelNet40 models. Based on this data, we derive a neural view quality measure abiding to human preferences. We further demonstrate that this view quality measure not only generalizes to models unseen during training, but also to unseen model categories. We are thus able to predict view qualities for single images, and directly predict human preferred viewpoints for 3D models by exploiting point‐based learning technology, without requiring to generate intermediate images or sampling the view sphere. We will detail our data collection procedure, describe the data analysis and model training and will evaluate the predictive quality of our trained viewpoint measure on unseen models and categories. To our knowledge, this is the first deep learning approach to predict a view quality measure solely based on human preferences. Sebastian Hartwig, Michael Schelling, Christian van Onzenoodt, Pere-Pau Vázquez, Pedro Hermosilla, Timo Ropinski |
Comput. Graph. Forum | 1 |