Sebastian Houben

dblp:47/8484 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-2036-419XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 11 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
abstract
Cross-lingual, cross-task transfer is challenged by task-specific data scarcity, which becomes more severe as language support grows and is further amplified in vision-language models (VLMs). We investigate multilingual generalization in encoder-decoder transformer VLMs to enable zero-shot image captioning in languages encountered only in the translation task. In this setting, the encoder must learn to generate generalizable, task-aware latent vision representations to instruct the decoder via inserted cross-attention layers. To analyze scaling behavior, we train Florence-2 based and Gemma-2 based models (0.4B to 11.2B parameters) on a synthetic dataset using varying compute budgets. While all languages in the dataset have image-aligned translations, only a subset of them include image captions. Notably, we show that captioning can emerge using a language prefix, even when this language only appears in the translation task. We find that indirect learning of unseen task-language pairs adheres to scaling laws that are governed by the multilinguality of the model, model size, and seen training samples. Finally, we demonstrate that the scaling laws extend to downstream tasks, achieving competitive performance through fine-tuning in multimodal machine translation (Multi30K, CoMMuTE), lexical disambiguation (CoMMuTE), and image captioning (Multi30K, XM3600, COCO Karpathy).
Julian Spravil, Sebastian Houben, Sven Behnke
AAAI2
2026 Code-Guided Reasoning in Vision-Language Models for Complex Diagram Understanding
abstract
Understanding complex structured diagrams, such as circuit schematics, molecular structures, musical notation, or business process models, requires precise symbolic, spatial, and relational reasoning.Current vision-language models (VLMs) struggle with such tasks because they lack access to the underlying symbolic structure that governs these diagrams.We introduce a training paradigm in which VLMs explicitly learn to reason through an intermediate symbolic representation of the image that is expressed in code.We generate a large synthetic dataset covering 21 diagram types across 7 domains by prompting large language models to generate code in specific formal representation languages (FRLs) and rendering them into paired code-image samples.During VLM training, the FRL code is provided along with the image, enabling the model to incorporate the symbolic representation during reasoning.Experiments show that models capable of producing valid code benefit from this symbolic intermediate layer, yielding improved accuracy on diagram understanding tasks.Our results demonstrate that integrating symbolic code into VLM training offers a promising direction for VLM design to handle complex visual data by bridging diagram perception with symbolic reasoning. 93
Daniel Steinigen, Lucie Flek, Sebastian Houben
ESANN3
2026 STRUDEL: Unrolling a Benchmark for Evaluating Vision-Language Models on Structured Diagram Understanding across Domains
abstract
Vision-Language Models (VLMs) have achieved impressive progress across diverse multimodal tasks, yet their ability to interpret structured diagrams, such as circuit schematics, molecular structures, musical notation, business process flow charts or class diagrams, which are central to scientific and engineering communication, remains underexplored. We introduce STRUDEL (STRUctured Diagram EvaLuation), a benchmark for evaluating VLMs on structured diagram understanding across 8 domains and 20 image categories. STRUDEL leverages Large-Language Models (LLMs) to synthesize code in domain-specific formal representation languages (FRLs) (e.g. circuit netlists, SMILES, ABC-Notation, BPMN or PlantUML), which are rendered into valid diagrams and paired with generated tasks, functional descriptions, and captions. A multi-stage pipeline filters invalid, cluttered, or redundant samples and employs LLM-as-a-judge scoring to ensure correctness. Through targeted experiments, we evaluate the ability of LLMs to generate valid code in distinct FRLs, demonstrating their capability to successfully perform this task. The resulting benchmark comprises diverse task types covering identification, quantification, structural analysis, image-text association, and image-to-code translation. Evaluating 35 VLMs using STRUDEL reveals that models excel at association tasks, demonstrating strong visual-textual alignment, yet struggle with quantification and identification, where precise structural understanding is required. Performance varies markedly in image-to-code translation, reflecting significant differences in how models connect visual inputs to formal representations. Overall, STRUDEL establishes a scalable foundation for assessing and advancing VLMs torward deeper and more systematic understanding of structured visual information across domains.
Daniel Steinigen, Lucie Flek, Sebastian Houben
LREC3
2025 Robustness Evaluation of the German Extractive Question Answering Task
abstract
To ensure reliable performance of Question Answering (QA) systems, evaluation of robustness is crucial. Common evaluation benchmarks commonly only include performance metrics, such as Exact Match (EM) and the F1 score. However, these benchmarks overlook critical factors for the deployment of QA systems. This oversight can result in systems vulnerable to minor perturbations in the input such as typographical errors. While several methods have been proposed to test the robustness of QA models, there has been minimal exploration of these approaches for languages other than English. This study focuses on the robustness evaluation of German language QA models, extending methodologies previously applied primarily to English. The objective is to nurture the development of robust models by defining an evaluation method specifically tailored to the German language. We assess the applicability of perturbations used in English QA models for German and perform a comprehensive experimental evaluation with eight models. The results show that all models are vulnerable to character-level perturbations. Additionally, the comparison of monolingual and multilingual models suggest that the former are less affected by character and word-level perturbations.
Shalaka Satheesh, Katharina Beckh, Katrin Klug, Héctor Allende-Cid, Sebastian Houben, Teena Hassan
COLING5
2025 Visual Latent Captioning - Towards Verbalizing Vision Transformer Encoders
Sogol Haghighat, Tim Daniel Metzler, Santosh Thoduka, Sebastian Houben
ECIR (2)4
2025 GO-VMP: Global Optimization for View Motion Planning in Fruit Mapping
abstract
Automating labor-intensive tasks such as crop monitoring with robots is essential for enhancing production and conserving resources. However, autonomously monitoring horticulture crops remains challenging due to their complex structures, which often result in fruit occlusions. Existing view planning methods attempt to reduce occlusions but either struggle to achieve adequate coverage or incur high robot motion costs. We introduce a global optimization approach for view motion planning that aims to minimize robot motion costs while maximizing fruit coverage. To this end, we leverage coverage constraints derived from the set covering problem (SCP) within a shortest Hamiltonian path problem (SHPP) formulation. While both SCP and SHPP are well-established, their tailored integration enables a unified framework that computes a global view path with minimized motion while ensuring full coverage of selected targets. Given the NP-hard nature of the problem, we employ a region-prior-based selection of coverage targets and a sparse graph structure to achieve effective optimization outcomes within a limited time. Experiments in simulation demonstrate that our method detects more fruits, enhances surface coverage, and achieves higher volume accuracy than the motion-efficient baseline with a moderate increase in motion cost, while significantly reducing motion costs compared to the coverage-focused baseline. Real-World experiments further confirm the practical applicability of our approach.
Allen Isaac Jose, Sicong Pan, Tobias Zaenker, Rohit U. Menon, Sebastian Houben, Maren Bennewitz
IROS5
2024 HyenaPixel: Global Image Context with Convolutions
abstract
In computer vision, a larger effective receptive field (ERF) is associated with better performance. While attention natively supports global context, its quadratic complexity limits its applicability to tasks that benefit from high-resolution input. In this work, we extend Hyena, a convolution-based attention replacement, from causal sequences to bidirectional data and two-dimensional image space. We scale Hyena’s convolution kernels beyond the feature map size, up to 191×191, to maximize ERF while maintaining sub-quadratic complexity in the number of pixels. We integrate our two-dimensional Hyena, HyenaPixel, and bidirectional Hyena into the MetaFormer framework. For image categorization, HyenaPixel and bidirectional Hyena achieve a competitive ImageNet-1k top-1 accuracy of 84.9% and 85.2%, respectively, with no additional training data, while outperforming other convolutional and large-kernel networks. Combining HyenaPixel with attention further improves accuracy. We attribute the success of bidirectional Hyena to learning the data-dependent geometric arrangement of pixels without a fixed neighborhood definition. Experimental results on downstream tasks suggest that HyenaPixel with large filters and a fixed neighborhood leads to better localization performance.
Julian Spravil, Sebastian Houben, Sven Behnke
ECAI2
2024 Bayesian Inverse Graphics for Few-Shot Concept Learning
Octavio Arriaga, Jichen Guo, Rebecca Adam, Sebastian Houben, Frank Kirchner
NeSy (1)4
2024 Robust Entropy Search for Safe Efficient Bayesian Optimization
abstract
The practical use of Bayesian Optimization (BO) in engineering applications imposes special requirements: high sampling efficiency on the one hand and finding a robust solution on the other hand. We address the case of adversarial robustness, where all parameters are controllable during the optimization process, but a subset of them is uncontrollable or even adversely perturbed at the time of application. To this end, we develop an efficient information-based acquisition function that we call Robust Entropy Search (RES). We empirically demonstrate its benefits in experiments on synthetic and real-life data. The results show that RES reliably finds robust optima, outperforming state-of-the-art algorithms.
Dorina Weichert, Alexander Kister, Sebastian Houben, Patrick Link, Gunar Ernis
UAI3
2022 Object Tracking for the Rotating Table Test
Vincent Scharf, Ibrahim Shakir Syed, Michal Stolarz, Mihir Mehta, Sebastian Houben
RoboCup5
2021 Dynamic Door Modeling for Monocular 3D Vehicle Detection
Thomas Barowski, Andre Brehme, Magdalena Szczot, Sebastian Houben
IV4
2021 Automated Selection of High-Quality Synthetic Images for Data-Driven Machine Learning: A Study on Traffic Signs
abstract
The utilization of automatically generated image training data is a feasible way to enhance existing datasets, e.g., by strengthening underrepresented classes or by adding new lighting or weather conditions for more variety. Synthetic images can also be used to introduce entirely new classes to a given dataset. In order to maximize the positive effects of generated image data on classifier training and reduce the possible downsides of potentially problematic image samples, an automatic quality assessment of each generated image seems sensible for overall quality enhancement of the training set and, thus, of the resulting classifier. In this paper we extend our previous work on synthetic traffic sign images by assessing the quality of a fully generated dataset consisting of 215,000 traffic sign images using four different measures. According to each sample's quality, we successively reduce the size of our training set and evaluate the performance with SVM and CNN classifiers to verify the approach. The comparability of real-world and synthetic training data is investigated by contrasting several classifiers trained on generated data to our baseline w.r.t. actual misclassifications during testing.
Daniela Horn, Lars Janssen, Sebastian Houben
IV3
2020 Fully Automated Traffic Sign Substitution in Real-World Images for Large-Scale Data Augmentation
abstract
Video-based traffic sign recognition is a key ability of autonomous vehicles but a demanding challenge due to the enormous number of classes and natural conditions in the wild. We address this problem with a fully automatic close-to-life image-to-image translation technique for traffic sign substitution in natural images (cf. Fig. 1). The work is intended as data augmentation technique and allows for training rare or unavailable traffic sign classes, or otherwise uncommon cases in visual traffic sign detection and classification. To this end, we extend our previous data generation model [1] and propose a rendering pipeline to create convincing traffic sign images with realistic background and camera recording artifacts. Experiments are conducted by exchanging traffic sign classes on different parts of the German Traffic Sign Recognition Benchmark (GTSRB) [2]. We demonstrate that the pipeline is well-suited for generating representative images of unseen traffic sign classes. A baseline image classification setup trained on real data shows an overall performance similar to being trained with a comparable number of artificial data samples. Our code is made publicly available under an open source license.
Daniela Horn, Sebastian Houben
IV2
2019 Generation of Natural Traffic Sign Images Using Domain Translation with Cycle-Consistent Generative Adversarial Networks
abstract
Video-based traffic sign recognition poses a highly challenging problem due to the significant number of possible classes and large variances of recording conditions in natural environments. Gathering an appropriate amount of data to solve this task with machine learning techniques remains an overall issue. In this study, we assess the suitability of automatically generated traffic sign images for training corresponding image classifiers. To this end, we adapt the recently proposed cycle-consistent generative adversarial networks in order to transfer automatically rendered prototypical traffic sign images for which we control type, pose, and-to a degree-background into their true-to-life counterparts. We test the proposed system by extensive experiments on the German Traffic Sign Recognition Benchmark dataset [1] and learn that both a HOG-feature-based SVM classifier and a state-of-the-art CNN exhibit reasonable performance when solely trained on artificial data. Consequently, it is well suited as data augmentation method and allows for covering uncommon cases and classes.
Dominic Spata, Daniela Horn, Sebastian Houben
IV3
2018 Keyframe-Based Photometric Online Calibration and Color Correction
abstract
Finding the parameters of a vignetting function for a camera currently involves the acquisition of several images in a given scene under very controlled lighting conditions, a cumbersome and error-prone task where the end result can only be confirmed visually. Many computer vision algorithms assume photoconsistency, constant intensity between scene points in different images, and tend to perform poorly if this assumption is violated. We present a real-time online vignetting and response calibration with additional exposure estimation for global-shutter color cameras. Our method does not require uniformly illuminated surfaces, known texture or specific geometry. The only assumptions are that the camera is moving, the illumination is static and reflections are Lambertian. Our method estimates the camera view poses by sparse visual SLAM and models the vignetting function by a small number of thin plate splines (TPS) together with a sixth-order polynomial to provide a dense estimation of attenuation from sparsely sampled scene points. The camera response function (CRF) is jointly modeled by a TPS and a Gamma curve. We evaluate our approach on synthetic datasets and in real-world scenarios with reference data from a Structure-from-Motion (SfM) system. We show clear visual improvement on textured meshes without the need for extensive meshing algorithms. A useful calibration is obtained from a few keyframes which makes an on-the-fly deployment conceivable.
Jan Quenzel, Jannis Horn, Sebastian Houben, Sven Behnke
IROS3
2018 Evaluation of Synthetic Video Data in Machine Learning Approaches for Parking Space Classification
abstract
Most modern computer vision techniques rely on large amounts of meticulously annotated data for training and evaluation. In close-to-market development, this demand is even higher since numerous common and-more important-less common situations have to be tested and must hence be covered datawise. However, gathering the necessary amount of data ready-labeled for the task at hand is a challenge of its own. Depending on the complexity of the objective and the chosen approach, the required amount of data can be vast. At the same time, the effort to capture all possible cases of a given problem grows with their variability. This makes recording new video data unfeasible, even impossible at times. In this work, we regard parking space classification as an exemplary application to target the imbalance of cost and benefit w.r.t. image data creation for machine learning approaches. We rely on a fully-fledged park deck simulation created with Unreal Engine 4 for data creation and replace all conventionally recorded and hand-labeled training data by automatically-annotated synthetic video data. We train several of-the-shelf classifiers with a common choice of feature inputs on synthetic images only and evaluate them on two realworld sequences of different outdoor car parks. We reach a classification performance that matches our previous work on this task in which all classifiers were developed solely with real-life video data.
Daniela Horn, Sebastian Houben
Intelligent Vehicles Symposium2
2017 Online depth calibration for RGB-D cameras using visual SLAM
abstract
Modern consumer RGB-D cameras are affordable and provide dense depth estimates at high frame rates. Hence, they are popular for building dense environment representations. Yet, the sensors often do not provide accurate depth estimates since the factory calibration exhibits a static deformation. We present a novel approach to online depth calibration that uses a visual SLAM system as reference for the measured depth. A sparse map is generated and the visual information is used to correct the static deformation of the measured depth while missing data is extrapolated using a small number of thin plate splines (TPS). The corrected depth can then be used to improve the accuracy of the sparse RGB-D map and the 3D environment reconstruction. As more data becomes available, the depth calibration is updated on the fly. Our method does not rely on a planar geometry like walls or a one-to-one-pixel correspondence between color and depth camera. Our approach is evaluated in real-world scenarios and against ground truth data. Comparison against two popular self-calibration methods is performed. Furthermore, we show clear visual improvement on aggregated point clouds with our method.
Jan Quenzel, Radu Alexandru Rosu, Sebastian Houben, Sven Behnke
IROS3
2016 Efficient multi-camera visual-inertial SLAM for micro aerial vehicles
abstract
Visual SLAM is an area of vivid research and bears countless applications for moving robots. In particular, micro aerial vehicles benefit from visual sensors due to their low weight. Their motion is, however, often faster and more complex than that of ground-based robots which is why systems with multiple cameras are currently evaluated and deployed. This, in turn, drives the computational demand for visual SLAM algorithms.
Sebastian Houben, Jan Quenzel, Nicola Krombach, Sven Behnke
IROS1
2014 Towards the intrinsic self-calibration of a vehicle-mounted omni-directional radially symmetric camera
abstract
Intrinsic calibration, i.e. finding the mapping between a camera's image positions and corresponding view rays, is a cumbersome, yet unavoidable task in order to accurately generate and interpret results from many kinds of image processing algorithms. We address this problem in the context of vehicle-mounted cameras with arbitrary fields of view with applications in advanced driver assistance systems. In particular, we present algorithms to gather the necessary data from unknown scenes and to subsequently estimate the camera parameters. These do rely on vehicle odometry only to resolve the focal scale ambiguity and to recognize when a purely translational motion is performed. We pay special attention to noise handling and circumvention of numerical instabilities. The proposed pipeline is tested by means of simulations to examine its noise sensitivity. Additionally we calibrate a fisheye camera from a natural scene of only 14 seconds length. First results show that the self-calibration in natural scenes is eligible and outperforms the straightforward approach of using all calibration parameters from an identically constructed camera.
Sebastian Houben
Intelligent Vehicles Symposium1
2014 Towards highly automated driving in a parking garage: General object localization and tracking using an environment-embedded camera system
abstract
In this study, we present a new indoor positioning and environment perception system for generic objects based on multiple surveillance cameras. In order to assist highly automated driving, our system detects the vehicle's position and any object along its current path to avoid collisions. A main advantage of the proposed approach is the usage of cameras that are already installed in the majority of parking garages. We generate precise object hypotheses in 3D world coordinates based on a given extrinsic camera calibration. Starting with a background subtraction algorithm for the segmentation of each camera image, we propose a robust view-ray intersection approach that enables the system to match and triangulate segmented hypotheses from all cameras. Comparing with LIDAR-based ground truth, we were able to evaluate the system's mean localization accuracy of 0.37 m for a variety of different sequences.
André Ibisch, Sebastian Houben, Marc Schlipsing, Robert Kesten, Paul Reimche, Florian Schuller, Harald Altinger
Intelligent Vehicles Symposium2
2013 Detection of traffic signs in real-world images: The German traffic sign detection benchmark
abstract
Real-time detection of traffic signs, the task of pinpointing a traffic sign's location in natural images, is a challenging computer vision task of high industrial relevance. Various algorithms have been proposed, and advanced driver assistance systems supporting detection and recognition of traffic signs have reached the market. Despite the many competing approaches, there is no clear consensus on what the state-of-the-art in this field is. This can be accounted to the lack of comprehensive, unbiased comparisons of those methods. We aim at closing this gap by the “German Traffic Sign Detection Benchmark” presented as a competition at IJCNN 2013 (International Joint Conference on Neural Networks). We introduce a real-world benchmark data set for traffic sign detection together with carefully chosen evaluation metrics, baseline results, and a web-interface for comparing approaches. In our evaluation, we separate sign detection from classification, but still measure the performance on relevant categories of signs to allow for benchmarking specialized solutions. The considered baseline algorithms represent some of the most popular detection approaches such as the Viola-Jones detector based on Haar features and a linear classifier relying on HOG descriptors. Further, a recently proposed problem-specific algorithm exploiting shape and color in a model-based Houghlike voting scheme is evaluated. Finally, we present the best-performing algorithms of the IJCNN competition.
Sebastian Houben, Johannes Stallkamp, Jan Salmen, Marc Schlipsing, Christian Igel
IJCNN1
2013 Video-based trailer detection and articulation estimation
abstract
Even for experienced drivers handling a roll trailer with a passenger car is a difficult and often tedious task. Moreover, the driver needs to keep track of the trailer's driving stability on unsteady roads. There are driver assistance systems that can simplify trajectory planning and observe the oscillation amplitude, but they require additional hardware. In this paper, we present a method for trailer detection and articulation angle measurement based on video data from a rear end wide-angle camera. It consists of two stages: to decide whether or not a trailer is coupled to the vehicle and to estimate its articulation angle. These calculations work on single video frames. The vehicle is therefore not required to be in motion. However, we stabilize the single frame estimations by temporal integration. We perform training and parameter optimization and evaluate the accuracy of our approach by comparing the results to those of an articulation measurement unit attached to a test vehicle's hitch. Results show that it can very reliably be determined whether or not a trailer is coupled to the vehicle. Furthermore, its articulation can be estimated with a mean error of less than two degrees.
Lukas Caup, Jan Salmen, Ibro Muharemovic, Sebastian Houben
Intelligent Vehicles Symposium4
2012 Google Street View images support the development of vision-based driver assistance systems
abstract
For the development of vision-based driver assistance systems, large amounts of data are needed, e.g., for training machine learning approaches, tuning parameters, and comparing different methods. There are basically three possible ways to obtain the required data: using freely available benchmark sets, doing own recordings, or falling back to synthesized sequences. In this paper, we show that Google Street View can be incorporated as a valuable source for image data. Street View is the largest publicly available collection of images recorded from a drivers' perspective, covering many different countries and scenarios. We describe how to efficiently access the data and present a framework that allows for virtual driving through a network of images. We assess its performance and show its applicability in practice considering traffic sign recognition as an example. The introduced approach supports an efficient collection of image data relevant to training and evaluating machine vision modules. It is easily adaptable and extendible, whereby Street View becomes a valuable tool for developers of vision-based assistance systems.
Jan Salmen, Sebastian Houben, Marc Schlipsing
Intelligent Vehicles Symposium2
2011 A single target voting scheme for traffic sign detection
abstract
Traffic sign detection and recognition is an important part of advanced driver assistance systems. Many prototype solutions for this task have been developed, and first commercial systems have just become available. Their image processing chain can be divided into three steps, preprocessing, detection, and recognition. Albeit several reliable sign recognition algorithms exist by now sign detection under real-world conditions is still unstable. Therefore, we address the first two steps of the processing chain presenting an analysis of widely used detectors, namely Hough-like methods. We evaluate several preprocessing steps and tweaks to increase their performance. Hence, the detectors are applied to a large, publicly available set of images from real-life traffic scenes. As main result we establish a new probabilistic measure for traffic sign colour detection and, based on the findings in our analysis, propose a novel Hough-like algorithm for detecting circular and triangular shapes. These improvements significantly increased detection performance in our experiments.
Sebastian Houben
Intelligent Vehicles Symposium1
2011 A variational approach to vesicle membrane reconstruction from fluorescence imaging
Kalin Kolev, Norbert Kirchgeßner, Sebastian Houben, Agnes Csiszár, Wolfgang Rubner, Christoph Palm, Björn Eiben, Rudolf Merkel, Daniel Cremers
Pattern Recognit.3