Volker Rodehorst

dblp:88/3164 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-4815-0118ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2
YearPublicationVenuePosition
2026 Patch Your Matcher: Correspondence-Aware Image-to-Image Translation Unlocks Cross-Modal Matching via Single-Modality Priors
abstract
Matching between image modalities is a high-impact research area. Current state-of-the-art (SOTA) methods rely on extensive multi-million-scale training protocols, which demand significant computational resources. However, the learned cross-modal mapping remains largely opaque and locked within the trained matcher, with limited options for downstream use or transfer to other matchers. To enable such capabilities, we propose Patch Your Matcher (PYM)1, a highly adaptive method for leveraging pre-trained single-modality matchers for cross-modal matching (Fig. 1) by co-learning an explicit two-view geometrically consistent mapping. PYM learns image-to-image (I2I) translations that map new modalities into the original matcher’s modality using a novel adversarial learning approach based on explicit evaluation of 6 DoF two-view correspondence plausibility. Trained with the semi-dense ELoFTR [80], our approach delivers substantially better cross-modal matching than classic I2I techniques, and recovers 97.05% of the matching precision of the extensively trained SOTA multi-modal MINIMA [62] variant. PYM also significantly boosts cross-modal matching performance of uni-modal sparse LightGlue [50] and dense RoMA [23] matchers, demonstrating high transferability of learned mapping.
Anton Frolov 0003, Volker Rodehorst
WACV2
2025 Crackstructures and Crackensembles: The Power of Multi-View for 2.5D Crack Detection
abstract
While research on structural crack segmentation at the image level remains highly active, progress beyond two dimensions has been limited. This stagnation is largely due to the lack of available data for crack detection in higher dimensions. To address this limitation, we introduce Crack-structures, a dataset tailored for real-world 2.5D crack segmentation, encompassing 15 segments from five distinct structures. Additionally, we present Crackensembles, a complementary semi-synthetic dataset that combines real textures with synthetic geometry to enhance the development of learning-based algorithms. Coupled with a baseline for multi-view crack instance segmentation, this work establishes a solid foundation for advancing algorithms that support real-world structural inspection.
Christian Benz, Volker Rodehorst
WACV2
2025 Needles & Haystacks: Dataset and Benchmark for Domain-Agnostic Image-Based Rigid Slice-to-Volume Registration
abstract
We address domain-agnostic slice-to-volume (S2V) registration, the alignment of 2D sliced/tomographic images into 3D volumes without prior knowledge of structure, shape, or orientation. While S2V registration is well-studied in medical imaging, which often relies on auxiliary information (e.g. landmarks, segmentation masks, pre-defined orientations, canonical/atlas volumes), applications such as micro-structure characterization in materials science lack such domain-specific aids. This leaves the task inherently ill-posed due to noise, unstructured regions, repetitive patterns, rotational and translational symmetries. To address this challenge, we present “Needles & Haystacks,”11Project page: https://xaf-cv.github.io/nh-rs2v/ a novel multi-domain algorithm development dataset with 158, 436 unique registration problems and ground-truth solutions, based on diverse and openly licensed real-world volumetric data. Additionally, we provide an online platform with 8, 461 test problems for reproducible evaluation of competing methods. We also propose strong baseline solutions with public implementations and highlight opportunities for further algorithmic advancements.
Anton Frolov 0003, Florian Kleiner, Christiane Rößler, Volker Rodehorst
WACV4
2025 Label Convergence: Defining an Upper Performance Bound in Object Recognition Through Contradictory Annotations
abstract
Annotation errors are a challenge not only during training of machine learning models, but also during their evaluation. Label variations and inaccuracies in datasets often manifest as contradictory examples that deviate from established labeling conventions. Such inconsistencies, when significant, prevent models from achieving optimal performance on metrics such as mean Average Precision (mAP). We introduce the notion of “label convergence” to describe the highest achievable performance under the constraint of contradictory test annotations, essentially defining an upper bound on model accuracy. Recognizing that noise is an inherent characteristic of all data, our study analyzes five real-world datasets, including the LVIS dataset, to investigate the phenomenon of label convergence. We approximate that label convergence is between 62.63-67.52 mAP@ [0.5:0.95:0.05 J for LVIS with 95% confidence, attributing these bounds to the presence of real annotation errors. With current state-of-the-art (SOTA) models at the upper end of the label convergence interval for the well-studied LVIS dataset, we conclude that model capacity is sufficient to solve current object detection problems. Therefore, future efforts should focus on three key aspects: (1) updating the problem specification and adjusting evaluation practices to account for unavoidable label noise, (2) creating cleaner data, especially test data, and (3) including multi-annotated data to investigate annotation variation and make these issues visible from the outset.
David Tschirschwitz, Volker Rodehorst
WACV2
2025 CISOL: An Open and Extensible Dataset for Table Structure Recognition in the Construction Industry
abstract
Reproducibility and replicability are critical pillars of empirical research, particularly in machine learning, where they depend not only on the availability of models, but also on the datasets used to train and evaluate those models. In this paper, we introduce the Construction Industry Steel Or-dering List (CISOL) dataset, which was developed with a focus on transparency to ensure reproducibility, replicability, and extensibility. CISOL provides a valuable new research resource and highlights the importance of having diverse datasets, even in niche application domains such as table extraction in civil engineering. CISOL is unique in that it contains real-world civil engineering documents from industry, making it a distinctive contribution to the field. The dataset contains more than 120,000 annotated instances in over 800 document images, positioning it as a medium-sized dataset that provides a robust foundation for Table Structure Recognition (TSR) and Table Detection (TD) tasks. Benchmarking results show that CISOL achieves 67.22 [email protected]:0.95:0.05 using the YOLOv8 model, outperforming the TSR-specific TATR model. This highlights the effectiveness of CISOL as a benchmark for advancing TSR, especially in specialized domains.
David Tschirschwitz, Volker Rodehorst
WACV2
2024 MVCrackViT: Robust Multi-View Crack Detection For Point Cloud Segmentation Using View Attention
abstract
Adapting computer vision algorithms for inspecting civil structures brings significant societal benefits. Images captured from civil structures often exhibit distinct overlap, typically to perform 3D reconstruction. In this work, the potential of multiple overlapping views is harnessed for robust multi-view crack detection. A transformer approach, named MVCrackViT, is designed to use attention over multiple views, enabling point cloud crack segmentation from the views directly. To address quality issues such as motion blur, defocus, and low exposure commonly found in real-world data, artificial view corruption is applied to accomplish training from image data alone. With reasonable positional tolerance, a performance of approximately 90% clCloudIoU is achieved on a 3D crack dataset, the first of its kind. The powerful clCloudIoU metric is introduced to evaluate crack detection in 3D space.
Christian Benz, Volker Rodehorst
ICIP2
2007 A Benchmarking Dataset for Performance Evaluation of Automatic Surface Reconstruction Algorithms
abstract
Numerous techniques were invented in computer vision and photogrammetry to obtain spatial information from digital images. We intend to describe and improve the performance of these vision techniques by providing test objectives, data, metrics and test protocols. In this paper we propose a comprehensive benchmarking dataset for evaluating a variety of automatic surface reconstruction algorithms (shape-from-X) and a methodology for comparing their results.
Anke Bellmann, Olaf Hellwich, Volker Rodehorst, Ulas Yilmaz
CVPR3
1997 Architectural Image Segmentation Using Digital Watersheds
Volker Rodehorst
CAIP1
1996 Color stereo vision using hierarchical block matching and active color illumination
abstract
Stereo is a well-known technique for obtaining depth information from digital images. Nevertheless, this technique still suffers from a lack in accuracy and/or long computation time needed to match stereo images. A new hierarchical algorithm using an image pyramid for obtaining dense depth maps from color stereo images is presented. We show that matching results of high quality are obtained when using the new hierarchical chromatic block matching algorithm. Most stereo matching algorithms can not compute correct dense depth maps in homogenous image regions. This paper shows that using an active color illumination will considerably improve the quality of the matching results. We present results for synthetic and for real images.
Andreas F. Koschan, Volker Rodehorst, Kathrin Spiller
ICPR2
1993 Algorithms for Shape from Shading, Lighting Direction and Motion
Reinhard Klette, Volker Rodehorst
CAIP2