Tobias Grüning

dblp:124/0668 · also Tobias Gruning · DBLP profile ↗
← Back
7ranked-venue papers in the field
2as first author
2since 2021 · last 2022
0000-0003-0031-4942ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 7 (2 first)
YearPublicationVenuePosition
2022 Rescoring Sequence-to-Sequence Models for Text Line Recognition with CTC-Prefixes
Christoph Wick, Jochen Zöllner, Tobias Grüning
DAS3
2021 Transformer for Handwritten Text Recognition Using Bidirectional Post-decoding
Christoph Wick, Jochen Zöllner, Tobias Grüning
ICDAR (3)3
2019 End-to-End Measure for Text Recognition
abstract
Measuring the performance of text recognition and text line detection engines is an important step to objectively compare systems and their configuration. There exist well-established measures for both tasks separately. However, there is no sophisticated evaluation scheme to measure the quality of a combined text line detection and text recognition system. The F-measure on word level is a well-known methodology, which is sometimes used in this context. Nevertheless, it does not take into account the alignment of hypothesis and ground truth text and can lead to deceptive results. Since users of automatic information retrieval pipelines in the context of text recognition are mainly interested in the end-to-end performance of a given system, there is a strong need for such a measure. Hence, we present a measure to evaluate the quality of an end-to-end text recognition system. The basis for this measure is the well established and widely used character error rate, which is limited - in its original form - to aligned hypothesis and ground truth texts. The proposed measure is flexible in a way that it can be configured to penalize different reading orders between the hypothesis and ground truth and can take into account the geometric position of the text lines. Additionally, it can ignore over-and under-segmentation of text lines. With these parameters it is possible to get a measure fitting best to its own needs.
Gundram Leifert, Roger Labahn, Tobias Grüning, Svenja Leifert
ICDAR3
2019 Evaluating Sequence-to-Sequence Models for Handwritten Text Recognition
abstract
Encoder-decoder models have become an effective approach for sequence learning tasks like machine translation, image captioning and speech recognition, but have yet to show competitive results for handwritten text recognition. To this end, we propose an attention-based sequence-to-sequence model. It combines a convolutional neural network as a generic feature extractor with a recurrent neural network to encode both the visual information, as well as the temporal context between characters in the input image, and uses a separate recurrent neural network to decode the actual character sequence. We make experimental comparisons between various attention mechanisms and positional encodings, in order to find an appropriate alignment between the input and output sequence. The model can be trained end-to-end and the optional integration of a hybrid loss allows the encoder to retain an interpretable and usable output, if desired. We achieve competitive results on the IAM and ICFHR2016 READ data sets compared to the state-of-the-art without the use of a language model, and we significantly improve over any recent sequence-to-sequence approaches.
Johannes Michael, Roger Labahn, Tobias Grüning, Jochen Zöllner
ICDAR3
2018 READ-BAD: A New Dataset and Evaluation Scheme for Baseline Detection in Archival Documents
abstract
Text line detection is crucial for any application associated with Automatic Text Recognition or Keyword Spotting. Modern algorithms perform good on well-established datasets since they either comprise clean data or simple/homogeneous page layouts. We have collected and annotated 2036 archival document images from different locations and time periods. The dataset contains varying page layouts and degradations that challenge text line segmentation methods. Well established text line segmentation evaluation schemes such as the Detection Rate or Recognition Accuracy demand for binarized data that is annotated on a pixel level. Producing ground truth by these means is laborious and not needed to determine a method's quality. In this paper we propose a new evaluation scheme that is based on baselines. The proposed scheme has no need for binarization and it can handle skewed as well as rotated text lines. The ICDAR 2017 Competition on Baseline Detection and the ICDAR 2017 Competition on Layout Analysis for Challenging Medieval Manuscripts used this evaluation scheme. Finally, we present results achieved by a recently published text line detection algorithm.
Tobias Grüning, Roger Labahn, Markus Diem, Florian Kleber, Stefan Fiel
DAS1
2017 cBAD: ICDAR2017 Competition on Baseline Detection
abstract
The cBAD competition aims at benchmarking state-of-the-art baseline detection algorithms. It is in line with previous competitions such as the ICDAR 2013 Handwriting Segmentation Contest. A new, challenging, dataset was created to test the behavior of state-of-the-art systems on real world data. Since traditional evaluation schemes are not applicable to the size and modality of this dataset, we present a new one that introduces baselines to measure performance. We received submissions from five different teams for both tracks.
Markus Diem, Florian Kleber, Stefan Fiel, Tobias Grüning, Basilios Gatos
ICDAR4
2017 A Robust and Binarization-Free Approach for Text Line Detection in Historical Documents
abstract
Text line extraction from complex handwritten documents, especially for historical collections, is still an unsolved problem. There is a strong demand for reliable and robust approaches since text line extraction is a crucial pre-processing step for modern text recognition and keyword spotting systems. We propose a binarization-free system which employs a newly developed clustering approach based on so-called 'superpixels'. Although multiple ways of generating superpixels were developed in the past, we demonstrate that even a standard method yields impressive results. Our clustering approach is applicable to various scenarios by making use of general characteristics of text lines (e.g., curvilinearity, interline spacings, local homogeneity), and without adapting its parametrization. State-of-the-art results are achieved by the same parametrization for 8 different well-established benchmarking datasets. These datasets cover historical and modern texts as well as images with diverse resolutions and fonts. The system is developed for detecting text lines in complex scenarios. It is not tuned to assign foreground pixels to detected text lines. Thus, superior performance is achieved for the historical datasets for which no pixel hit accuracy of 95% is required. Remarkably, for the dataset of the ICDAR 2015 Competition on Text Line Detection in Historical Documents, the average cost per text line was reduced from 9.77 (winning team) to 8.19.
Tobias Grüning, Gundram Leifert, Tobias Strauß, Roger Labahn
ICDAR1