Mathias Seuret

dblp:156/2985 · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0001-9153-1031ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 26 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 23 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 An Analysis of Lightweight Models for Document Image Machine Translation
Abantika Bose, Thomas Gorges, Lukas Hüttner, Linda-Sophie Schneider, Mathias Seuret, Fei Wu 0025, Vincent Christlein
ICDAR (2)5
2025 Zero-Shot Paragraph-level Handwriting Imitation with Latent Diffusion Models
abstract
Abstract The imitation of cursive handwriting is mainly limited to generating handwritten words or lines. Multiple synthetic outputs must be stitched together to create paragraphs or whole pages, whereby consistency and layout information are lost. To close this gap, we propose a method for imitating handwriting at the paragraph level that also works for unseen writing styles. Therefore, we introduce a modified latent diffusion model that enriches the encoder-decoder mechanism with specialized loss functions that explicitly preserve the style and content. We enhance the attention mechanism of the diffusion model with adaptive 2D positional encoding and the conditioning mechanism to work with two modalities simultaneously: a style image and the target text. This significantly improves the realism of the generated handwriting. We set a new benchmark in our comprehensive evaluation, achieving 61 % mAP and 56 % top-1 accuracy in style preservation, significantly outperforming the previous best method (37 % mAP, 30 % top-1). We are making our code publicly available for reproducibility, supporting research in this area and research into potential countermeasures: https://github.com/M4rt1nM4yr/paragraph_handwriting_imitation_ldm
Martin Mayr, Marcel Dreier, Florian Kordon, Mathias Seuret, Jochen Zöllner, Fei Wu 0025, Andreas K. Maier, Vincent Christlein
Int. J. Comput. Vis.4
2025 Lightweight cross-attention-based HookNet for historical handwritten document layout analysis
Fei Wu 0025, Mathias Seuret, Martin Mayr, Florian Kordon, Jochen Zöllner, Sebastian Wind, Andreas K. Maier, Vincent Christlein
Int. J. Document Anal. Recognit.2
2025 Data-efficient handwritten text recognition of diplomatic historical text
abstract
Abstract Traditional methods in handwritten text recognition primarily focus on generating basic transcriptions, which often fall short for in-depth humanities research. Our study enhances this by providing diplomatic transcriptions for German studies, meticulously reproducing the original manuscripts, including layout and expanded abbreviations. State-of-the-art sequence-to-sequence approaches for handwritten text recognition predominantly use Connectionist Temporal Classification (CTC) as an auxiliary loss of the encoder output to improve robustness and accuracy. This is not possible in this task due to the great differences in the length of diplomatic transcriptions. We propose using the basic transcription instead of the diplomatic one as an additional target for the CTC feedback. Additionally, we introduce positional encoding at the intersection between the encoder and decoder to resolve the conflict of competing encoder objectives, balancing CTC loss reduction with the maintenance of implicit positional encoding for the decoder. Our empirical tests on the newly created dataset “Nuremberg Letterbooks” demonstrate significant data efficiency improvements. With only 4000 training lines (about 130 transcribed pages), we achieve a Character Error Rate (CER) of 9.39% without expanded abbreviations and 12.07% with expanded abbreviations, outperforming the baseline errors of 14.26% and 68.21%, respectively.
Martin Mayr, Katharina Neumeier, Julian Krenz, Simon Bürcky, Florian Kordon, Mathias Seuret, Jochen Zöllner, Fei Wu 0025, Andreas K. Maier, Vincent Christlein
Multim. Tools Appl.6
2024 fang: Fast Annotation of Glyphs in Historical Printed Documents
Florian Kordon, Nikolaus Weichselbaumer, Randall Herz, Janne van der Loop, Stephen Mossman, Edward Potten, Mathias Seuret, Martin Mayr, Fei Wu 0025, Vincent Christlein
DAS7
2024 ICDAR 2024 Competition on Multi Font Group Recognition and OCR
Janne van der Loop, Florian Kordon, Martin Mayr, Vincent Christlein, Fei Wu 0025, Dalia Rodríguez-Salas, Nikolaus Weichselbaumer, Mathias Seuret
ICDAR (6)8
2024 Evaluating learned feature aggregators for writer retrieval
abstract
Abstract Transformers have emerged as the leading methods in natural language processing, computer vision, and multi-modal applications due to their ability to capture complex relationships and dependencies in data. In this study, we explore the potential of transformers as feature aggregators in the context of patch-based writer retrieval, with the objective of improving the quality of writer retrieval by effectively summarizing the relevant features from image patches. Our investigation underscores the complexity of leveraging transformers as feature aggregators in patch-based writer retrieval. While we have experimented with various model configurations, augmentations, and learning objectives, the performance of transformers in this task has room for improvement. This observation highlights the challenges in this domain and emphasizes the need for further research to enhance their effectiveness. By shedding light on the limitations of transformers in this context, our study contributes to the growing body of knowledge in the field of writer retrieval and provides valuable insights for future research and development in this area.
Alexander Mattick, Martin Mayr, Mathias Seuret, Florian Kordon, Fei Wu 0025, Vincent Christlein
Int. J. Document Anal. Recognit.3
2023 WordStylist: Styled Verbatim Handwritten Text Generation with Latent Diffusion Models
Konstantina Nikolaidou, George Retsinas, Vincent Christlein, Mathias Seuret, Giorgos Sfikas, Elisa H. Barney Smith, Hamam Mokayed, Marcus Liwicki
ICDAR (2)4
2023 Multi-stage Fine-Tuning Deep Learning Models Improves Automatic Assessment of the Rey-Osterrieth Complex Figure Test
Benjamin Schuster, Florian Kordon, Martin Mayr, Mathias Seuret, Stefanie Jost, Josef Kessler, Vincent Christlein
ICDAR (1)4
2023 Combining OCR Models for Reading Early Modern Books
Mathias Seuret, Janne van der Loop, Nikolaus Weichselbaumer, Martin Mayr, Janina Molnar, Tatjana Hass, Vincent Christlein
ICDAR (5)1
2023 ICDAR 2023 Competition on Detection and Recognition of Greek Letters on Papyri
Mathias Seuret, Isabelle Marthot-Santaniello, Stephen A. White, Olga Serbaeva Saraogi, Selaudin Agolli, Guillaume Carrière, Dalia Rodríguez-Salas, Vincent Christlein
ICDAR (2)1
2023 Classification of incunable glyphs and out-of-distribution detection with joint energy-based models
abstract
Abstract Optical character recognition (OCR) has proved a powerful tool for the digital analysis of printed historical documents. However, its ability to localize and identify individual glyphs is challenged by the tremendous variety in historical type design, the physicality of the printing process, and the state of conservation. We propose to mitigate these problems by a downstream fine-tuning step that corrects for pathological and undesirable extraction results. We implement this idea by using a joint energy-based model which classifies individual glyphs and simultaneously prunes potential out-of-distribution (OOD) samples like rubrications, initials, or ligatures. During model training, we introduce specific margins in the energy spectrum that aid this separation and explore the glyph distribution’s typical set to stabilize the optimization procedure. We observe strong classification at 0.972 AUPRC across 42 lower- and uppercase glyph types on a challenging digital reproduction of Johannes Balbus’ Catholicon, matching the performance of purely discriminative methods. At the same time, we achieve OOD detection rates of 0.989 AUPRC and 0.946 AUPRC for OOD ‘clutter’ and ‘ligatures’ which substantially improves upon recently proposed OOD detection techniques. The proposed approach can be easily integrated into the postprocessing phase of current OCR to aid reproduction and shape analysis research.
Florian Kordon, Nikolaus Weichselbaumer, Randall Herz, Stephen Mossman, Edward Potten, Mathias Seuret, Martin Mayr, Vincent Christlein
Int. J. Document Anal. Recognit.6
2022 Investigating the Effect of Using Synthetic and Semi-synthetic Images for Historical Document Font Classification
Konstantina Nikolaidou, Richa Upadhyay, Mathias Seuret, Marcus Liwicki
DAS3
2022 A Fair Evaluation of Various Deep Learning-Based Document Image Binarization Approaches
Richin Sukesh, Mathias Seuret, Anguelos Nicolaou, Martin Mayr, Vincent Christlein
DAS2
2022 A survey of historical document image datasets
abstract
Abstract This paper presents a systematic literature review of image datasets for document image analysis, focusing on historical documents, such as handwritten manuscripts and early prints. Finding appropriate datasets for historical document analysis is a crucial prerequisite to facilitate research using different machine learning algorithms. However, because of the very large variety of the actual data (e.g., scripts, tasks, dates, support systems, and amount of deterioration), the different formats for data and label representation, and the different evaluation processes and benchmarks, finding appropriate datasets is a difficult task. This work fills this gap, presenting a meta-study on existing datasets. After a systematic selection process (according to PRISMA guidelines), we select 65 studies that are chosen based on different factors, such as the year of publication, number of methods implemented in the article, reliability of the chosen algorithms, dataset size, and journal outlet. We summarize each study by assigning it to one of three pre-defined tasks: document classification, layout structure, or content analysis. We present the statistics, document type, language, tasks, input visual aspects, and ground truth information for every dataset. In addition, we provide the benchmark tasks and results from these papers or recent competitions. We further discuss gaps and challenges in this domain. We advocate for providing conversion tools to common formats (e.g., COCO format for computer vision tasks) and always providing a set of evaluation metrics, instead of just one, to make results comparable across studies.
Konstantina Nikolaidou, Mathias Seuret, Hamam Mokayed, Marcus Liwicki
Int. J. Document Anal. Recognit.2
2021 SmartPatch: Improving Handwritten Word Imitation with Patch Discriminators
Alexander Mattick, Martin Mayr, Mathias Seuret, Andreas K. Maier, Vincent Christlein
ICDAR (1)3
2021 ICDAR 2021 Competition on Historical Document Classification
Mathias Seuret, Anguelos Nicolaou, Dalia Rodríguez-Salas, Nikolaus Weichselbaumer, Dominique Stutzmann, Martin Mayr, Andreas K. Maier, Vincent Christlein
ICDAR (4)1
2020 Re-Ranking for Writer Identification and Writer Retrieval
Simon Jordan, Mathias Seuret, Pavel Král, Ladislav Lenc, Jirí Martínek, Barbara Wiermann, Tobias Schwinger, Andreas K. Maier, Vincent Christlein
DAS2
2020 The Notary in the Haystack - Countering Class Imbalance in Document Processing with CNNs
Martin Leipert, Georg Vogeler, Mathias Seuret, Andreas K. Maier, Vincent Christlein
DAS3
2020 ICFHR 2020 Competition on Image Retrieval for Historical Handwritten Fragments
abstract
This competition succeeds upon a line of competitions for writer and style analysis of historical document images. In particular, we investigate the performance of large-scale retrieval of historical document fragments in terms of style and writer identification. The analysis of historic fragments is a difficult challenge commonly solved by trained humanists. In comparison to previous competitions, we make the results more meaningful by addressing the issue of sample granularity and moving from writer to page fragment retrieval. The two approaches, style and author identification, provide information on what kind of information each method makes better use of and indirectly contribute to the interpretability of the participating method. Therefore, we created a large dataset consisting of more than 120 000 fragments. Although the most teams submitted methods based on convolutional neural networks, the winning entry achieves an mAP below 40 %.
Mathias Seuret, Anguelos Nicolaou, Andreas K. Maier, Vincent Christlein, Dominique Stutzmann
ICFHR1
2020 Trainable Spectrally Initializable Matrix Transformations in Convolutional Neural Networks
abstract
In this work, we introduce a new architectural component to Neural Network (NN), i.e., trainable and spectrally initializable matrix transformations on feature maps. While previous literature has already demonstrated the possibility of adding static spectral transformations as feature processors, our focus is on more general trainable transforms. We study the transforms in various architectural configurations on four datasets of different nature: from medical (ColorectalHist, HAM10000) and natural (Flowers) images to historical documents (CB55). With rigorous experiments that control for the number of parameters and randomness, we show that networks utilizing the introduced matrix transformations outperform vanilla neural networks. The observed accuracy increases appreciably across all datasets. In addition, we show that the benefit of spectral initialization leads to significantly faster convergence, as opposed to randomly initialized matrix transformations. The transformations are implemented as auto-differentiable PyTorch modules that can be incorporated into any neural network architecture. The entire code base is open-source.
Michele Alberti, Angela Botros, Narayan Schütz, Rolf Ingold, Marcus Liwicki, Mathias Seuret
ICPR6
2019 Labeling, Cutting, Grouping: An Efficient Text Line Segmentation Method for Medieval Manuscripts
abstract
This paper introduces a new way for text-line extraction by integrating deep-learning based pre-classification and state-of-the-art segmentation methods. Text-line extraction in complex handwritten documents poses a significant challenge, even to the most modern computer vision algorithms. Historical manuscripts are a particularly hard class of documents as they present several forms of noise, such as degradation, bleed-through, interlinear glosses, and elaborated scripts. In this work, we propose a novel method which uses semantic segmentation at pixel level as intermediate task, followed by a text-line extraction step. We measured the performance of our method on a recent dataset of challenging medieval manuscripts and surpassed state-of-the-art results by reducing the error by 80.7%. Furthermore, we demonstrate the effectiveness of our approach on various other datasets written in different scripts. Hence, our contribution is two-fold. First, we demonstrate that semantic pixel segmentation can be used as strong denoising pre-processing step before performing text line extraction. Second, we introduce a novel, simple and robust algorithm that leverages the high-quality semantic segmentation to achieve a text-line extraction performance of 99.42% line IU on a challenging dataset.
Michele Alberti, Lars Vögtlin, Vinaychandran Pondenkandath, Mathias Seuret, Rolf Ingold, Marcus Liwicki
ICDAR4
2019 ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents
abstract
This competition investigates the performance of large-scale retrieval of historical document images based on writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing a total of 20 000 document images representing about 10 000 writers, divided in three types: writers of (i) manuscript books, (ii) letters, (iii) charters and legal documents. We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as writer retrieval. The most teams submitted traditional methods not using deep learning techniques. The competition results show that a combination of methods is outperforming single methods. Furthermore, letters are much more difficult to retrieve than manuscripts.
Vincent Christlein, Anguelos Nicolaou, Mathias Seuret, Dominique Stutzmann, Andreas K. Maier
ICDAR3
2019 Deep Generalized Max Pooling
abstract
Global pooling layers are an essential part of Convolutional Neural Networks (CNN). They are used to aggregate activations of spatial locations to produce a fixed-size vector in several state-of-the-art CNNs. Global average pooling or global max pooling are commonly used for converting convolutional features of variable size images to a fix-sized embedding. However, both pooling layer types are computed spatially independent: each individual activation map is pooled and thus activations of different locations are pooled together. In contrast, we propose Deep Generalized Max Pooling that balances the contribution of all activations of a spatially coherent region by re-weighting all descriptors so that the impact of frequent and rare ones is equalized. We show that this layer is superior to both average and max pooling on the classification of Latin medieval manuscripts (CLAMM'16, CLAMM'17), as well as writer identification (Historical-WI'17).
Vincent Christlein, Lukas Spranger, Mathias Seuret, Anguelos Nicolaou, Pavel Král, Andreas K. Maier
ICDAR3
2018 A Semi-automatized Modular Annotation Tool for Ancient Manuscript Annotation
abstract
In this paper, we present DIVAnnotation, an ancient document annotation tool which is freely available as open source. This software is easily modular thanks to the splitting of the different annotation steps through the use of a tabbed graphical user interface. State-of-the-art document image analysis methods are included through web services, thus allowing users to generate automatically annotations and correct them manually when needed. The annotations are stored into a highly structured TEI file which makes data access and manipulation simple. A Java library for managing TEI files generated by DIVAnnotation is also provided as open source.
Mathias Seuret, Manuel Bouillon, Foteini Liwicki, Marcel Gygli, Marcus Liwicki, Rolf Ingold
DAS1
2017 Convolutional Neural Networks for Page Segmentation of Historical Document Images
abstract
This paper presents a page segmentation method for handwritten historical document images based on a Convolutional Neural Network (CNN). We consider page segmentation as a pixel labeling problem, i.e., each pixel is classified as one of the predefined classes. Traditional methods in this area rely on hand-crafted features carefully tuned considering prior knowledge. In contrast, we propose to learn features from raw image pixels using a CNN. While many researchers focus on developing deep CNN architectures to solve different problems, we train a simple CNN with only one convolution layer. We show that the simple architecture achieves competitive results against other deep architectures on different public datasets. Experiments also demonstrate the effectiveness and superiority of the proposed method compared to previous methods.
Kai Chen 0011, Mathias Seuret, Jean Hennebert, Rolf Ingold
ICDAR2
2017 PCA-Initialized Deep Neural Networks Applied to Document Image Analysis
abstract
In this paper, we present a novel approach for initializing deep neural networks, i.e., by using Principal Component Analysis (PCA) to initialize neural layers. Usually, the initialization of the weights of a deep neural network is done in one of the three following ways: 1) with random values, 2) layer-wise, usually as Deep Belief Network or as auto-encoder, and 3) re-use of layers from another network (transfer learning). Therefore, typically, many training epochs are needed before meaningful weights are learned, or a rather similar dataset is required for seeding a fine-tuning of transfer learning. In this paper, we describe how to turn a PCA into an auto-encoder, by generating an encoder layer of the PCA parameters and furthermore adding a decoding layer. We analyze the initialization technique on real documents. First, we show that a PCA-based initialization is quick and leads to a very stable initialization. Furthermore, for the task of layout analysis we investigate the effectiveness of PCA-based initialization and show that it outperforms state-of-the-art random weight initialization methods.
Mathias Seuret, Michele Alberti, Marcus Liwicki, Rolf Ingold
ICDAR1
2017 ICDAR2017 Competition on Layout Analysis for Challenging Medieval Manuscripts
abstract
This paper reports on the ICDAR2017 Competition on Layout Analysis for Challenging Medieval Manuscripts (HisDoc-Layout-Comp) and provides further details and discussions. In this competition we introduce a new challenging dataset and state-of-the-art benchmark results for pixel-labelling and text line segmentation. The DIVA-HisDB comprises medieval manuscripts with complex layout in contrast to previous datasets, where rectangular text blocks and only a few decorative elements exist. In particular, the images of this competition contain many interlinear and marginal glosses as well as texts in various sizes and decorated letters. This makes the distinction of the four target labels (text, comment, decoration, and background) more difficult. In addition, to reflect the needs of scholars in the humanities, we request multi-labeling of certain regions (decorated text as text and decoration). Furthermore, we measure not just the accuracy, but the Intersection over Union (IU) of pixel sets, which better reflects the real performance. Indeed, in our results we observe that the accuracy appears to be rather high, but the IU reveals, that there is still room for improvement. For the task of line segmentation, the recognition results are rather low (overall error higher than 5%). Noteworthy, a combination of the best layout analysis method with an adapted seam-carving based method achieves better results than the best contestant.
Foteini Liwicki, Manuel Bouillon, Mathias Seuret, Marcel Gygli, Michele Alberti, Rolf Ingold, Marcus Liwicki
ICDAR3
2017 Selecting Fine-Tuned Features for Layout Analysis of Historical Documents
abstract
In this paper, we investigate fine-tuned features learned by deep neural networks in the context of layout analysis. Pre-training and fine-tuning are techniques used in deep neural networks to learn representations (features) of input. However, it is not clear if the fine-tuned features are all useful for a following classification task. We investigate this problem using feature selection. Firstly, features are learned by a deep neural network, where stacked autoencoders are used for pre-training and then the whole network is fine-tuned. Then, a feature selection method is used to select relevant features for classification. We observe that despite fine-tuning, a significant number of the features are still redundant or irrelevant for layout classification. Furthermore, features from the top layer of the stacked autoencoders are generally more relevant for classification than those from lower layers.
Hao Wei 0001, Mathias Seuret, Marcus Liwicki, Rolf Ingold, Pei Fu
ICDAR2
2017 A User-Centered Segmentation Method for Complex Historical Manuscripts Based on Document Graphs
abstract
In historical manuscripts, humans can detect handwritten words, lines, and decorations with lightness even if they do not know the language or the script. Yet for automatic processing this task has proven elusive, especially in the case of handwritten documents with complex layouts, which is why semiautomatic methods that integrate the human user into the process are needed. In this paper, we introduce a user-centered segmentation method based on document graphs and scribbling interaction. The graphs capture a sparse representation of the document's structure that can then be edited by the user with a stylus on a touch-sensitive screen. We evaluate the proposed method on a newly introduced database of historical manuscripts with complex layout and demonstrate, first, that the document graphs are already close to the desired segmentation and, second, that scribbling allows a natural and efficient interaction.
Angelika Garz, Mathias Seuret, Andreas Fischer 0002, Rolf Ingold
IEEE Trans. Hum. Mach. Syst.2
2016 Page Segmentation for Historical Document Images Based on Superpixel Classification with Unsupervised Feature Learning
abstract
In this paper, we present an efficient page segmentation method for historical document images. Many existing methods either rely on hand-crafted features or perform rather slow as they treat the problem as a pixel-level assignment problem. In order to create a feasible method for real applications, we propose to use superpixels as basic units of segmentation, and features are learned directly from pixels. An image is first oversegmented into superpixels with the simple linear iterative clustering (SLIC) algorithm. Then, each superpixel is represented by the features of its central pixel. The features are learned from pixel intensity values with stacked convolutional autoencoders in an unsupervised manner. A support vector machine (SVM) classifier is used to classify superpixels into four classes: periphery, background, text block, and decoration. Finally, the segmentation results are refined by a connected component based smoothing procedure. Experiments on three public datasets demonstrate that compared to our previous method, the proposed method is much faster and achieves comparable segmentation results. Additionally, much fewer pixels are used for classifier training.
Kai Chen 0011, Mathias Seuret, Marcus Liwicki, Jean Hennebert, Rolf Ingold
DAS3
2016 Creating Ground Truth for Historical Manuscripts with Document Graphs and Scribbling Interaction
abstract
Ground truth is both - indispensable for training and evaluating document analysis methods, and yet very tedious to create manually. This especially holds true for complex historical manuscripts that exhibit challenging layouts with interfering and overlapping handwriting. In this paper, we propose a novel semi-automatic system to support layout annotations in such a scenario based on document graphs and a pen-based scribbling interaction. On the one hand, document graphs provide a sparse page representation that is already close to the desired ground truth and on the other hand, scribbling facilitates an efficient and convenient pen-based interaction with the graph. The performance of the system is demonstrated in the context of a newly introduced database of historical manuscripts with complex layouts.
Angelika Garz, Mathias Seuret, Foteini Liwicki, Andreas Fischer 0002, Rolf Ingold
DAS2
2016 Text Detection in Arabic News Video Based on SWT Operator and Convolutional Auto-Encoders
abstract
Text detection in videos is a challenging problem due to variety of text specificities, presence of complex background and anti-aliasing/compression artifacts. In this paper, we present an approach for horizontally aligned artificial text detection in Arabic news video. The novelty of this method revolves around the combination of two techniques: an adapted version of the Stroke Width Transform (SWT) algorithm and a convolutional auto-encoder (CAE). First, the SWT extracts text candidates' components. They are then filtered and grouped using geometric constraints and Stroke Width information. Second, the CAE is used as an unsupervised feature learning method to discriminate the obtained textline candidates as text or non-text. We assess the proposed approach on the public Arabic-Text-in-Video database (AcTiV-DB) using different evaluation protocols including data from several TV channels. Experiments indicate that the use of learned features significantly improves the text detection results.
Oussama Zayene, Mathias Seuret, Sameh Masmoudi Touj, Jean Hennebert, Rolf Ingold, Najoua Essoukri Ben Amara
DAS2
2016 Page Segmentation for Historical Handwritten Document Images Using Conditional Random Fields
abstract
In this paper, we present a Conditional Random Field (CRF) model to deal with the problem of segmenting handwritten historical document images into different regions. We consider page segmentation as a pixel-labeling problem, i.e., each pixel is assigned to one of a set of labels. Features are learned from pixel intensity values with stacked convolutional autoencoders in an unsupervised manner. The features are used for the purpose of initial classification with a multilayer perceptron. Then a CRF model is introduced for modeling the local and contextual information jointly in order to improve the segmentation. For the purpose of decreasing the time complexity, we perform labeling at superpixel level. In the CRF model, graph nodes are represented by superpixels. The label of each pixel is determined by the label of the superpixel to which it belongs. Experiments on three public datasets demonstrate that, compared to previous methods, the proposed method achieves more accurate segmentation results and is much faster.
Kai Chen 0011, Mathias Seuret, Marcus Liwicki, Jean Hennebert, Rolf Ingold
ICFHR2
2016 N-Light-N: A Highly-Adaptable Java Library for Document Analysis with Convolutional Auto-Encoders and Related Architectures
abstract
This paper presents a novel, highly-adaptable Java framework N-light-N, for the work with deep neural networks, especially with CAEs. While the most popular deep learning libraries focus on fast processing and high performance, they only implement the main-stream network architectures and network units. In recent research in the document domain, however, we have shown that modified networks, units, and training processes significantly improve the performance in various tasks. To enable the document research community with such capabilities, in this paper we introduce a novel, publicly available Deep Learning framework which is easy to use, adapt, and extend. Furthermore, we present successful applications for three tasks, including two in the domain of handwritten historical documents, and show how the framework can be used for adaptation, optimization, and deeper analysis.
Mathias Seuret, Rolf Ingold, Marcus Liwicki
ICFHR1
2016 DIVA-HisDB: A Precisely Annotated Large Dataset of Challenging Medieval Manuscripts
abstract
This paper introduces a publicly available historical manuscript database DIVA-HisDB for the evaluation of several Document Image Analysis (DIA) tasks. The database consists of 150 annotated pages of three different medieval manuscripts with challenging layouts. Furthermore, we provide a layout analysis ground-truth which has been iterated on, reviewed, and refined by an expert in medieval studies. DIVA-HisDB and the ground truth can be used for training and evaluating DIA tasks, such as layout analysis, text line segmentation, binarization and writer identification. Layout analysis results of several representative baseline technologies are also presented in order to help researchers evaluate their methods and advance the frontiers of complex historical manuscripts analysis. An optimized state-of-the-art Convolutional Auto-Encoder (CAE) performs with around 95% accuracy, demonstrating that for this challenging layout there is much room for improvement. Finally, we show that existing text line segmentation methods fail due to interlinear and marginal text elements.
Foteini Liwicki, Mathias Seuret, Nicole Eichenberger, Angelika Garz, Marcus Liwicki, Rolf Ingold
ICFHR2
2015 Page segmentation of historical document images with convolutional autoencoders
abstract
In this paper, we present an unsupervised feature learning method for page segmentation of historical handwritten documents available as color images. We consider page segmentation as a pixel labeling problem, i.e., each pixel is classified as either periphery, background, text block, or decoration. Traditional methods in this area rely on carefully hand-crafted features or large amounts of prior knowledge. In contrast, we apply convolutional autoencoders to learn features directly from pixel intensity values. Then, using these features to train an SVM, we achieve high quality segmentation without any assumption of specific topologies and shapes. Experiments on three public datasets demonstrate the effectiveness and superiority of the proposed approach.
Kai Chen 0011, Mathias Seuret, Marcus Liwicki, Jean Hennebert, Rolf Ingold
ICDAR2
2015 Gradient-domain degradations for improving historical documents images layout analysis
abstract
We present a novel method for adding realistic degradations to historical document images in order to generate more training data. Degradation patches are extracted from other documents and applied to the target document in the gradient domain. Working in the gradient domain has not been done for this purpose in document images analysis so far. It has the advantage to prevent color inconsistencies and allows to efficiently avoid border effects. This paper contains the detailed description of our novel method, with a focus on the mathematical aspect of the transition to and from the gradient domain. Furthermore, we perform quantitative experiments where we investigate the effects of using synthetically generated training data on historical documents with different kind of degradations.
Mathias Seuret, Kai Chen 0011, Nicole Eichenberger, Marcus Liwicki, Rolf Ingold
ICDAR1
2014 Pixel Level Handwritten and Printed Content Discrimination in Scanned Documents
abstract
Classification of the content of a scanned document as either printed or handwritten is typically tackled as a segmentation problem of pages into text lines or words. However these methods are not applicable on documents where handwritten annotations overlay printed text. In this paper we propose to treat the task as a pixel classification task, i.e., To classify individual foreground pixels into either printed or handwritten pixels. Our method uses various features of diverse nature taking the surrounding window into account. The influence of the features and their parameters are investigated and optimized on a validation set. Each foreground pixel is then classified by a multilayer perceptron using feature vectors based on a pixel neighborhood. Finally, a post-processing step corrects typical misclassifications, i.e., It removes outliers based on several heuristics. We evaluated our method on printed documents with real handwritten annotations and reached an accuracy of 96.10% on the test set. This is significantly higher than a previously published methods based on local features.
Mathias Seuret, Marcus Liwicki, Rolf Ingold
ICFHR1