Martin Mayr

dblp:05/6334 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-3706-285XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Zero-Shot Paragraph-level Handwriting Imitation with Latent Diffusion Models
abstract
Abstract The imitation of cursive handwriting is mainly limited to generating handwritten words or lines. Multiple synthetic outputs must be stitched together to create paragraphs or whole pages, whereby consistency and layout information are lost. To close this gap, we propose a method for imitating handwriting at the paragraph level that also works for unseen writing styles. Therefore, we introduce a modified latent diffusion model that enriches the encoder-decoder mechanism with specialized loss functions that explicitly preserve the style and content. We enhance the attention mechanism of the diffusion model with adaptive 2D positional encoding and the conditioning mechanism to work with two modalities simultaneously: a style image and the target text. This significantly improves the realism of the generated handwriting. We set a new benchmark in our comprehensive evaluation, achieving 61 % mAP and 56 % top-1 accuracy in style preservation, significantly outperforming the previous best method (37 % mAP, 30 % top-1). We are making our code publicly available for reproducibility, supporting research in this area and research into potential countermeasures: https://github.com/M4rt1nM4yr/paragraph_handwriting_imitation_ldm
Martin Mayr, Marcel Dreier, Florian Kordon, Mathias Seuret, Jochen Zöllner, Fei Wu 0025, Andreas K. Maier, Vincent Christlein
Int. J. Comput. Vis.1
2025 Lightweight cross-attention-based HookNet for historical handwritten document layout analysis
Fei Wu 0025, Mathias Seuret, Martin Mayr, Florian Kordon, Jochen Zöllner, Sebastian Wind, Andreas K. Maier, Vincent Christlein
Int. J. Document Anal. Recognit.3
2025 Data-efficient handwritten text recognition of diplomatic historical text
abstract
Abstract Traditional methods in handwritten text recognition primarily focus on generating basic transcriptions, which often fall short for in-depth humanities research. Our study enhances this by providing diplomatic transcriptions for German studies, meticulously reproducing the original manuscripts, including layout and expanded abbreviations. State-of-the-art sequence-to-sequence approaches for handwritten text recognition predominantly use Connectionist Temporal Classification (CTC) as an auxiliary loss of the encoder output to improve robustness and accuracy. This is not possible in this task due to the great differences in the length of diplomatic transcriptions. We propose using the basic transcription instead of the diplomatic one as an additional target for the CTC feedback. Additionally, we introduce positional encoding at the intersection between the encoder and decoder to resolve the conflict of competing encoder objectives, balancing CTC loss reduction with the maintenance of implicit positional encoding for the decoder. Our empirical tests on the newly created dataset “Nuremberg Letterbooks” demonstrate significant data efficiency improvements. With only 4000 training lines (about 130 transcribed pages), we achieve a Character Error Rate (CER) of 9.39% without expanded abbreviations and 12.07% with expanded abbreviations, outperforming the baseline errors of 14.26% and 68.21%, respectively.
Martin Mayr, Katharina Neumeier, Julian Krenz, Simon Bürcky, Florian Kordon, Mathias Seuret, Jochen Zöllner, Fei Wu 0025, Andreas K. Maier, Vincent Christlein
Multim. Tools Appl.1
2025 SSL4SAR: Self-Supervised Learning for Glacier Calving Front Extraction From SAR Imagery
abstract
Glaciers are losing ice mass at unprecedented rates, increasing the need for accurate, year-round monitoring to understand frontal ablation, particularly the factors driving the calving process. Deep learning models can extract calving front positions from Synthetic Aperture Radar imagery to track seasonal ice losses at the calving fronts of marine- and lake-terminating glaciers. The current state-of-the-art model relies on ImageNet-pretrained weights. However, they are suboptimal due to the domain shift between the natural images in ImageNet and the specialized characteristics of remote sensing imagery, in particular for Synthetic Aperture Radar imagery. To address this challenge, we propose two novel self-supervised multimodal pretraining techniques that leverage SSL4SAR, a new unlabeled dataset comprising 9,563 Sentinel-1 and 14 Sentinel-2 images of Arctic glaciers, with one optical image per glacier in the dataset. Additionally, we introduce a novel hybrid model architecture that combines a Swin Transformer encoder with a residual Convolutional Neural Network (CNN) decoder. When pretrained on SSL4SAR, this model achieves a mean distance error of 293m on the “CAlving Fronts and where to Find thEm” (CaFFe) benchmark dataset, outperforming the prior best model by 67 m. Evaluating an ensemble of the proposed model on a multi-annotator study of the benchmark dataset reveals a mean distance error of 75 m, approaching the human performance of 38 m. This advancement enables precise monitoring of seasonal changes in glacier calving fronts.
Nora Gourmelon, Marcel Dreier, Martin Mayr, Thorsten Seehaus, Dakota Pyles, Matthias H. Braun, Andreas K. Maier, Vincent Christlein
IEEE Trans. Geosci. Remote. Sens.3
2024 fang: Fast Annotation of Glyphs in Historical Printed Documents
Florian Kordon, Nikolaus Weichselbaumer, Randall Herz, Janne van der Loop, Stephen Mossman, Edward Potten, Mathias Seuret, Martin Mayr, Fei Wu 0025, Vincent Christlein
DAS8
2024 ICDAR 2024 Competition on Multi Font Group Recognition and OCR
Janne van der Loop, Florian Kordon, Martin Mayr, Vincent Christlein, Fei Wu 0025, Dalia Rodríguez-Salas, Nikolaus Weichselbaumer, Mathias Seuret
ICDAR (6)3
2024 Evaluating learned feature aggregators for writer retrieval
abstract
Abstract Transformers have emerged as the leading methods in natural language processing, computer vision, and multi-modal applications due to their ability to capture complex relationships and dependencies in data. In this study, we explore the potential of transformers as feature aggregators in the context of patch-based writer retrieval, with the objective of improving the quality of writer retrieval by effectively summarizing the relevant features from image patches. Our investigation underscores the complexity of leveraging transformers as feature aggregators in patch-based writer retrieval. While we have experimented with various model configurations, augmentations, and learning objectives, the performance of transformers in this task has room for improvement. This observation highlights the challenges in this domain and emphasizes the need for further research to enhance their effectiveness. By shedding light on the limitations of transformers in this context, our study contributes to the growing body of knowledge in the field of writer retrieval and provides valuable insights for future research and development in this area.
Alexander Mattick, Martin Mayr, Mathias Seuret, Florian Kordon, Fei Wu 0025, Vincent Christlein
Int. J. Document Anal. Recognit.2
2023 Multi-stage Fine-Tuning Deep Learning Models Improves Automatic Assessment of the Rey-Osterrieth Complex Figure Test
Benjamin Schuster, Florian Kordon, Martin Mayr, Mathias Seuret, Stefanie Jost, Josef Kessler, Vincent Christlein
ICDAR (1)3
2023 Combining OCR Models for Reading Early Modern Books
Mathias Seuret, Janne van der Loop, Nikolaus Weichselbaumer, Martin Mayr, Janina Molnar, Tatjana Hass, Vincent Christlein
ICDAR (5)4
2023 Classification of incunable glyphs and out-of-distribution detection with joint energy-based models
abstract
Abstract Optical character recognition (OCR) has proved a powerful tool for the digital analysis of printed historical documents. However, its ability to localize and identify individual glyphs is challenged by the tremendous variety in historical type design, the physicality of the printing process, and the state of conservation. We propose to mitigate these problems by a downstream fine-tuning step that corrects for pathological and undesirable extraction results. We implement this idea by using a joint energy-based model which classifies individual glyphs and simultaneously prunes potential out-of-distribution (OOD) samples like rubrications, initials, or ligatures. During model training, we introduce specific margins in the energy spectrum that aid this separation and explore the glyph distribution’s typical set to stabilize the optimization procedure. We observe strong classification at 0.972 AUPRC across 42 lower- and uppercase glyph types on a challenging digital reproduction of Johannes Balbus’ Catholicon, matching the performance of purely discriminative methods. At the same time, we achieve OOD detection rates of 0.989 AUPRC and 0.946 AUPRC for OOD ‘clutter’ and ‘ligatures’ which substantially improves upon recently proposed OOD detection techniques. The proposed approach can be easily integrated into the postprocessing phase of current OCR to aid reproduction and shape analysis research.
Florian Kordon, Nikolaus Weichselbaumer, Randall Herz, Stephen Mossman, Edward Potten, Mathias Seuret, Martin Mayr, Vincent Christlein
Int. J. Document Anal. Recognit.7
2022 Is Multitask Learning Always Better?
Alexander Mattick, Martin Mayr, Andreas K. Maier, Vincent Christlein
DAS2
2022 Combining Visual and Linguistic Models for a Robust Recipient Line Recognition in Historical Documents
Martin Mayr, Alex Felker, Andreas K. Maier, Vincent Christlein
DAS1
2022 A Fair Evaluation of Various Deep Learning-Based Document Image Binarization Approaches
Richin Sukesh, Mathias Seuret, Anguelos Nicolaou, Martin Mayr, Vincent Christlein
DAS4
2021 SmartPatch: Improving Handwritten Word Imitation with Patch Discriminators
Alexander Mattick, Martin Mayr, Mathias Seuret, Andreas K. Maier, Vincent Christlein
ICDAR (1)2
2021 ICDAR 2021 Competition on Historical Document Classification
Mathias Seuret, Anguelos Nicolaou, Dalia Rodríguez-Salas, Nikolaus Weichselbaumer, Dominique Stutzmann, Martin Mayr, Andreas K. Maier, Vincent Christlein
ICDAR (4)6
2019 Weakly Supervised Segmentation of Cracks on Solar Cells Using Normalized Lp Norm
abstract
Photovoltaic is one of the most important renewable energy sources for dealing with world-wide steadily increasing energy consumption. This raises the demand for fast and scalable automatic quality management during production and operation. However, the detection and segmentation of cracks on electroluminescence (EL) images of mono- or polycrystalline solar modules is a challenging task. In this work, we propose a weakly supervised learning strategy that only uses image-level annotations to obtain a method that is capable of segmenting cracks on EL images of solar cells. We use a modified ResNet-50 to derive a segmentation from network activation maps. We use defect classification as a surrogate task to train the network. To this end, we apply normalized Lpnormalization to aggregate the activation maps into single scores for classification. In addition, we provide a study how different parameterizations of the normalized Lplayer affect the segmentation performance. This approach shows promising results for the given task. However, we think that the method has the potential to solve other weakly supervised segmentation problems as well.
Martin Mayr, Mathis Hoffmann, Andreas K. Maier, Vincent Christlein
ICIP1