Lars Vögtlin

dblp:242/9407 · DBLP profile ↗
← Back
4ranked-venue papers in the field
1as first author
3since 2021 · last 2024
0000-0002-2543-9074ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 4 (1 first)
YearPublicationVenuePosition
2024 Are Layout Analysis and OCR Still Useful for Document Information Extraction Using Foundation Models?
Anna Scius-Bertrand, Atefeh Fakhari, Lars Vögtlin, Daniel Ribeiro Cabral, Andreas Fischer 0002
ICDAR (4)3
2023 Layout Analysis of Historical Document Images Using a Light Fully Convolutional Network
Najoua Rahal, Lars Vögtlin, Rolf Ingold
ICDAR (5)2
2021 Generating Synthetic Handwritten Historical Documents with OCR Constrained GANs
Lars Vögtlin, Manuel Drazyk, Vinaychandran Pondenkandath, Michele Alberti, Rolf Ingold
ICDAR (3)1
2019 Labeling, Cutting, Grouping: An Efficient Text Line Segmentation Method for Medieval Manuscripts
abstract
This paper introduces a new way for text-line extraction by integrating deep-learning based pre-classification and state-of-the-art segmentation methods. Text-line extraction in complex handwritten documents poses a significant challenge, even to the most modern computer vision algorithms. Historical manuscripts are a particularly hard class of documents as they present several forms of noise, such as degradation, bleed-through, interlinear glosses, and elaborated scripts. In this work, we propose a novel method which uses semantic segmentation at pixel level as intermediate task, followed by a text-line extraction step. We measured the performance of our method on a recent dataset of challenging medieval manuscripts and surpassed state-of-the-art results by reducing the error by 80.7%. Furthermore, we demonstrate the effectiveness of our approach on various other datasets written in different scripts. Hence, our contribution is two-fold. First, we demonstrate that semantic pixel segmentation can be used as strong denoising pre-processing step before performing text line extraction. Second, we introduce a novel, simple and robust algorithm that leverages the high-quality semantic segmentation to achieve a text-line extraction performance of 99.42% line IU on a challenging dataset.
Michele Alberti, Lars Vögtlin, Vinaychandran Pondenkandath, Mathias Seuret, Rolf Ingold, Marcus Liwicki
ICDAR2