VLDB 2026 Research / reviewers in the wild / expert
Fabian Wolf
dblp:23/3308
· DBLP profile ↗
20ranked-venue papers
11as first author
8since 2021 · last 2025
0000-0001-8842-3718ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 6 since 2021Systems, architecture and hardware · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CM1 - A Dataset for Evaluating Few-Shot Information Extraction with Large Vision Language Models
Fabian Wolf, Oliver Tüselmann, Arthur Matei, Lukas Hennies, Christoph Rass, Gernot A. Fink |
ICDAR (2) | 1 |
| 2024 | Self-training for handwritten word recognition and retrievalabstractAbstract Handwritten text recognition and Word Retrieval, also known as Word Spotting, are traditional problems in the document analysis community. While the use of increasingly large neural network architectures has led to a steady improvement of performances it comes with the drawback of requiring manually annotated training data. This poses a tremendous problem considering their application to new document collections. To overcome this drawback, we propose a self-training approach that allows to train state-of-the-art models for HTR and word spotting. Self-training is a common technique in semi-supervised learning and usually relies on a small labeled dataset and training on pseudo-labels generated by an initial model. In this work, we show that it is feasible to train models on synthetic data that are sufficiently performant to serve as initial models for self-training. Therefore, the proposed training method does not rely on any manually annotated samples. We further investigate visual and language properties of the synthetic datasets. In order to improve performance and robustness of the self-training approach, we propose different confidence measures for both models that allow to identify and remove erroneous pseudo-labels. The presented training approach clearly outperforms other learning-free methods or adaptation strategies under the absence of manually annotated data. Fabian Wolf, Gernot A. Fink |
Int. J. Document Anal. Recognit. | 1 |
| 2022 | Recognition-Free Question Answering on Handwritten Document Collections
Oliver Tüselmann, Friedrich Müller, Fabian Wolf, Gernot A. Fink |
ICFHR | 3 |
| 2022 | Combining Self-training and Minimal Annotations for Handwritten Word Recognition
Fabian Wolf, Gernot A. Fink |
ICFHR | 1 |
| 2022 | Self-Training of Handwritten Word Recognition for Synthetic-to-Real AdaptationabstractPerformances of Handwritten Text Recognition (HTR) models are largely determined by the availability of labeled and representative training samples. However, in many application scenarios, labeled samples are scarce or costly to obtain. In this work, we propose a self-training approach to train a HTR model solely on synthetic samples and unlabeled data. The proposed training scheme uses an initial model trained on synthetic data to make predictions for the unlabeled target dataset. Starting from this initial model with rather poor performance, we show that a considerable adaptation is possible by training against the iteratively predicted pseudo-labels. Therefore, the investigated self-training method does not require any manually annotated training samples. We evaluate the proposed method on four benchmark datasets and show its effectiveness on reducing the gap to a model trained in a fully-supervised manner. Fabian Wolf, Gernot A. Fink |
ICPR | 1 |
| 2021 | Are End-to-End Systems Really Necessary for NER on Handwritten Document Images?
Oliver Tüselmann, Fabian Wolf, Gernot A. Fink |
ICDAR (2) | 2 |
| 2021 | Graph Convolutional Neural Networks for Learning Attribute Representations for Word Spotting
Fabian Wolf, Andreas Fischer 0002, Gernot A. Fink |
ICDAR (1) | 1 |
| 2021 | Annotation-Free Word Spotting with Bag-of-Features HMMsabstractThe annotation-free word spotting method that is proposed in this paper makes document images searchable without requiring any labeled training data. Thus, our method supports the exploration of a document collection directly without demanding any manual efforts from the users for the preparation of a training dataset. Our method works in the query-by-example scenario where the user selects an exemplary occurrence of the query word. Afterwards, the entire collection of document images is searched according to visual similarity to the query. The proposed method requires only minimal assumptions about the visual appearance of text. This is achieved by processing document images as a whole without requiring a given segmentation of the images on word level or on line level. Therefore, the method is also segmentation-free. Word size variabilities can be handled by representing the sequential structure of text with a statistical sequence model. In order to make the computationally costly application of the sequence model feasible in practice, regions are retrieved according to approximate similarity with an efficient model decoding algorithm. Re-ranking these regions according to the visual similarity obtained with the sequence model leads to highly accurate word spotting results. The method is evaluated on five benchmark datasets. In the segmentation-free query-by-example scenario where no annotated training data is available, the method outperforms all other methods that have been evaluated on any of these five benchmarks. Leonard Rothacker, Fabian Wolf, Gernot A. Fink |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2020 | Annotation-Free Learning of Deep Representations for Word Spotting Using Synthetic Data and Self Labeling
Fabian Wolf, Gernot A. Fink |
DAS | 1 |
| 2020 | Identifying and Tackling Key Challenges in Semantic Word SpottingabstractSemantic word spotting is an extension of the traditional word spotting approach that uses not only visual but also semantic information to determine the similarity between a word image and a given query. Current approaches in this area achieve a semantic retrieval by embedding word images into a textually trained semantic space. The related literature presents remarkable results regarding established metrics indicating that the task of semantic word image retrieval is solved. A closer look at the results reveals, however, that this is only partially the case. In this work, we identify and solve current key challenges for semantic word spotting. We analyze the published works in this field towards these challenges and show why they do not solve them. For this purpose, we demonstrate that the used embedding space from current methods contains strong artifacts influencing the retrieval task. Furthermore, we evaluate a more suitable and established embedding approach from Natural Language Processing for semantic word spotting. We also explain the challenges of mapping word images into a semantic embedding space and evaluate different architectures for this task. Thereby, we present a new architecture that outperforms current approaches in this area. In addition, we show that commonly used metrics are not suitable for evaluating a semantic retrieval and present a new evaluation metric for this task. Oliver Tüselmann, Fabian Wolf, Gernot A. Fink |
ICFHR | 2 |
| 2020 | Improving Handwritten Word Synthesis for Annotation-free Word SpottingabstractAnnotation-free word spotting aims at retrieving relevant word images from a document collection without the need of a manually labeled training dataset. As annotated data is usually scarce in the application scenarios of a word spotting system, transfer learning and annotation-free methods became increasingly popular. One possibility to alleviate the annotation problem is to train on synthetically generated word images. Therefore, a common approach is to render word images from electronic fonts and to vary the synthesis parameters randomly. In this work, we show that an annotation-free word spotting method benefits from an adapted synthesis procedure. We investigate the influence of the choice of the underlying vocabulary and the combination of synthesis and data augmentation. Furthermore, we present a method to adapt the style of the synthesized word images to the target dataset. We evaluate the proposed changes to the synthesis procedure on three benchmark datasets and improve performances considerably. Fabian Wolf, Kai Brandenbusch, Gernot A. Fink |
ICFHR | 1 |
| 2019 | Exploring Confidence Measures for Word Spotting in Heterogeneous DatasetsabstractIn recent years, convolutional neural networks (CNNs) took over the field of document analysis and they became the predominant model for word spotting. Especially attribute CNNs, which learn the mapping between a word image and an attribute representation, showed exceptional performances. The drawback of this approach is the overconfidence of neural networks when used out of their training distribution. In this paper, we explore different metrics for quantifying the confidence of a CNN in its predictions, specifically on the retrieval problem of word spotting. With these confidence measures, we limit the inability of a retrieval list to reject certain candidates. We investigate four different approaches that are either based on the network's attribute estimations or make use of a surrogate model. Our approach also aims at answering the question for which part of a dataset the retrieval system gives reliable results. We further show that there exists a direct relation between the proposed confidence measures and the quality of an estimated attribute representation. Fabian Wolf, Philipp Oberdiek, Gernot A. Fink |
ICDAR | 1 |
| 2017 | Agile Procedures of an Automotive OEM - Views from Different Business Areas
Alexander Poth, Fabian Wolf |
EuroSPI | 2 |
| 2008 | U2VAS: A Research Communication Stack for Vehicular NetworksabstractThis paper has shown that U2VAS fulfills all requirements of a truly modular and flexible communication stack for VANETs. Beyond the framework, the current implementation already contains several forms of communication, positioning and security mechanisms. We invite interested developers to join our effort and contribute to the further development of U2VAS. Elmar Schoch, Frank Kargl, Fabian Wolf, Michael Weber 0001 |
VTC Fall | 3 |
| 2005 | Context Sensitive Performance Analysis of Automotive ApplicationsabstractAccurate timing analysis is key to efficient embedded system synthesis and integration. While industrial control software systems are developed using graphical models, such as Matlab/Simulink or ASCET/SD, exhaustive simulation is not suitable for verifying functional and timing behavior. Formal performance analysis is an alternative, but can lead to wide timing intervals because of input data dependency and complex target architectures. Hence, a designer might want to restrict the formal performance analysis to parts of the software system, called context or process modes. We describe how to define and characterize such context information from graphical models. Further, we extend the formal performance analysis to consider contexts. Results front an automotive application demonstrate the applicability of our approach. Jan Staschulat, Rolf Ernst, Andreas Schulze, Fabian Wolf |
DATE | 4 |
| 2003 | Formal Methods for Integration of Automotive SoftwareabstractNovel functionality, configurability and higher efficiency in automotive systems require sophisticated embedded software as well as distributed software development between manufacturers and control unit suppliers. However, at least for engine control units (ECU), there exists today no well-defined software integration process that satisfies all key requirements of automotive manufacturers. We propose a methodology for safe integration of automotive software functions where required performance information is exchanged while each partner's IP is protected. We claim that, in principle, performance requirements and constraints (timing, memory consumption) for each software component and for the complete ECU can be formally validated, and believe that ultimately such formal analysis will be required for legal certification of an ECU. Marek Jersak, Kai Richter 0001, Rolf Ernst, Jörn-Christian Braam, Zheng-Yu Jiang, Fabian Wolf |
DATE | 6 |
| 2003 | Safe Automotive Software Development
Ken Tindell, Hermann Kopetz, Fabian Wolf, Rolf Ernst |
DATE | 3 |
| 2002 | Associative caches in formal software timing analysisabstractPrecise cache analysis is crucial to formally determine program running time. As cache simulation is unsafe with respect to the conservative running time bounds for real-time systems, current cache analysis techniques combine basic block level cache modeling with explicit or implicit program path analysis. We present an approach that extends instruction and data cache modeling from the granularity of basic blocks to program segments thereby increasing the overall running time analysis precision. Data flow analysis and local simulation of program segments are combined to safely predict cache line contents for associative caches in software running time analysis. The experiments show significant improvements in analysis precision over previous approaches on a typical embedded processor. Fabian Wolf, Jan Staschulat, Rolf Ernst |
DAC | 1 |
| 2001 | Execution cost interval refinement in static software analysis
Fabian Wolf, Rolf Ernst |
J. Syst. Archit. | 1 |
| 2001 | Path clustering in software timing analysisabstractVerification of program running time is essential in system design with real-time constraints. Simulation with incomplete test patterns or simple instruction counting are not appropriate for complex architectures. Software running times of embedded systems are process state and input data dependent. Formal analysis of such dependencies leads to software running time intervals rather than single values. These intervals depend on program properties, execution paths, and states of processes, as well as on the target architecture. An approach to analysis of process behavior using running time intervals is presented. It improves our previous work by exploiting program segments with single paths and by taking the execution context into account. The example of an asynchronous transfer mode (ATM) cell handler demonstrates significant improvements in analysis precision. Experimental results show the superiority of the presented approach over well-established approaches. Fabian Wolf, Rolf Ernst, Wei Ye 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |