Lars Vögtlin

dblp:242/9407 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0002-2543-9074ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Are Layout Analysis and OCR Still Useful for Document Information Extraction Using Foundation Models?
Anna Scius-Bertrand, Atefeh Fakhari, Lars Vögtlin, Daniel Ribeiro Cabral, Andreas Fischer 0002
ICDAR (4)3
2024 Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models
Anna Scius-Bertrand, Michael Jungo, Lars Vögtlin, Jean-Marc Spat
ICPR (19)3
2024 Approximate ground truth generation for semantic labeling of historical documents with minimal human effort
abstract
Abstract Deep learning approaches have shown high performance for layout analysis of historical documents, provided that enough labeled data is available. This is not an issue for generic tasks such as image binarization, text graphics separation, or text line and text block detection but can become an impediment for more specialized tasks specific to one or a few books only. This paper addresses layout analysis of medieval books with rich and complex layouts, for which no labeled data is initially available. The proposed strategy consists of training an initial model with artificial data created to reflect the rules a deep neural network should learn. Then, the model is iteratively fine-tuned by mixing the artificial data with real data obtained by previous predictions, post-processed, and manually selected by an expert user. Such a strategy needs less human effort than manual ground truthing. The approach is qualitatively and quantitatively assessed and shows that the system converges to an accurate model that finally produces approximate ground truth stable and good enough to train a final model to solve the targeted task with high accuracy.
Najoua Rahal, Lars Vögtlin, Rolf Ingold
Int. J. Document Anal. Recognit.2
2023 Layout Analysis of Historical Document Images Using a Light Fully Convolutional Network
Najoua Rahal, Lars Vögtlin, Rolf Ingold
ICDAR (5)2
2023 Historical document image analysis using controlled data for pre-training
abstract
Abstract Using neural networks for semantic labeling has become a dominant technique for layout analysis of historical document images. However, to train or fine-tune appropriate models, large labeled datasets are needed. This paper addresses the case when only limited labeled data are available and promotes a novel approach using so-called controlled data to pre-train the networks. Two different strategies are proposed: The first addresses the real labeling task by using artificial data; the second uses real data to pre-train the networks with a pretext task. To assess these strategies, a large set of experiments has been carried out on a text line detection and classification task using different variants of U-Net. The observations, obtained from two different datasets, show that globally the approach reduces the training time while offering similar or better performance. Furthermore, the effect is bigger on lightweight network architectures.
Najoua Rahal, Lars Vögtlin, Rolf Ingold
Int. J. Document Anal. Recognit.2
2021 Generating Synthetic Handwritten Historical Documents with OCR Constrained GANs
Lars Vögtlin, Manuel Drazyk, Vinaychandran Pondenkandath, Michele Alberti, Rolf Ingold
ICDAR (3)1
2019 Labeling, Cutting, Grouping: An Efficient Text Line Segmentation Method for Medieval Manuscripts
abstract
This paper introduces a new way for text-line extraction by integrating deep-learning based pre-classification and state-of-the-art segmentation methods. Text-line extraction in complex handwritten documents poses a significant challenge, even to the most modern computer vision algorithms. Historical manuscripts are a particularly hard class of documents as they present several forms of noise, such as degradation, bleed-through, interlinear glosses, and elaborated scripts. In this work, we propose a novel method which uses semantic segmentation at pixel level as intermediate task, followed by a text-line extraction step. We measured the performance of our method on a recent dataset of challenging medieval manuscripts and surpassed state-of-the-art results by reducing the error by 80.7%. Furthermore, we demonstrate the effectiveness of our approach on various other datasets written in different scripts. Hence, our contribution is two-fold. First, we demonstrate that semantic pixel segmentation can be used as strong denoising pre-processing step before performing text line extraction. Second, we introduce a novel, simple and robust algorithm that leverages the high-quality semantic segmentation to achieve a text-line extraction performance of 99.42% line IU on a challenging dataset.
Michele Alberti, Lars Vögtlin, Vinaychandran Pondenkandath, Mathias Seuret, Rolf Ingold, Marcus Liwicki
ICDAR2