VLDB 2026 Research / reviewers in the wild / expert
Marcus Liwicki
dblp:28/1247 · also Marcus Eichenberger-Liwicki
· DBLP profile ↗
162ranked-venue papers
26as first author
22since 2021 · last 2026
0000-0003-4029-6574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 107 · 16 first-author · 16 since 2021Databases, data management, data science and information retrieval · 68 · 13 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tiny Data, Big Policy? Generative AI Agents for Context-Sensitive Climate Governance
Aparup Khatua, Hamam Mokayed, Foteini Liwicki, Marcus Liwicki |
COMPSAC | 4 |
| 2025 | ASTrA: Adversarial Self-supervised Training with Adaptive-AttacksabstractExisting self-supervised adversarial training (self-AT) methods rely on hand-crafted adversarial attack strategies for PGD attacks, which fail to adapt to the evolving learning dynamics of the model and do not account for instance-specific characteristics of images. This results in sub-optimal adversarial robustness and limits the alignment between clean and adversarial data distributions. To address this, we propose $\textit{ASTrA}$ ($\textbf{A}$dversarial $\textbf{S}$elf-supervised $\textbf{Tr}$aining with $\textbf{A}$daptive-Attacks), a novel framework introducing a learnable, self-supervised attack strategy network that autonomously discovers optimal attack parameters through exploration-exploitation in a single training episode. ASTrA leverages a reward mechanism based on contrastive loss, optimized with REINFORCE, enabling adaptive attack strategies without labeled data or additional hyperparameters. We further introduce a mixed contrastive objective to align the distribution of clean and adversarial examples in representation space. ASTrA achieves state-of-the-art results on CIFAR10, CIFAR100, and STL10 while integrating seamlessly as a plug-and-play module for other self-AT methods. ASTrA shows scalability to larger datasets, demonstrates strong semi-supervised performance, and is resilient to robust overfitting, backed by explainability analysis on optimal attack strategies. Project page for source code and other details at https://prakashchhipa.github.io/projects/ASTrA. Prakash Chandra Chhipa, Gautam Vashishtha, Settur Jithamanyu, Rajkumar Saini, Mubarak Shah, Marcus Liwicki |
ICLR | 6 |
| 2024 | LCM: Log Conformal Maps for Robust Representation Learning to Mitigate Perspective Distortion
Meenakshi Subhash Chippa, Prakash Chandra Chhipa, Kanjar De, Marcus Liwicki, Rajkumar Saini |
ACCV (8) | 4 |
| 2024 | Möbius Transform for Mitigating Perspective Distortions in Representation Learning
Prakash Chandra Chhipa, Meenakshi Subhash Chippa, Kanjar De, Rajkumar Saini, Marcus Liwicki, Mubarak Shah |
ECCV (73) | 5 |
| 2024 | DiffusionPen: Towards Controlling the Style of Handwritten Text Generation
Konstantina Nikolaidou, George Retsinas, Giorgos Sfikas, Marcus Liwicki |
ECCV (85) | 4 |
| 2024 | A Historical Handwritten Dataset for Ethiopic OCR with Baseline Models and Human-Level Performance
Birhanu Belay, Isabelle Guyon, Tadele Mengiste, Bezawork Tilahun, Marcus Liwicki, Tesfa Tegegne, Romain Egele |
ICDAR (3) | 5 |
| 2024 | Editorial for special issue on "advanced topics in document analysis and recognition"
Elisa H. Barney Smith, Marcus Liwicki, Liangrui Peng, Simone Marinai |
Int. J. Document Anal. Recognit. | 2 |
| 2023 | WordStylist: Styled Verbatim Handwritten Text Generation with Latent Diffusion Models
Konstantina Nikolaidou, George Retsinas, Vincent Christlein, Mathias Seuret, Giorgos Sfikas, Elisa H. Barney Smith, Hamam Mokayed, Marcus Liwicki |
ICDAR (2) | 8 |
| 2023 | Towards End-to-End Semi-Supervised Table Detection with Deformable Transformer
Tahira Shehzadi, Khurram Azeem Hashmi, Didier Stricker, Marcus Liwicki, Muhammad Zeshan Afzal |
ICDAR (2) | 4 |
| 2023 | Functional Knowledge Transfer with Self-supervised Representation LearningabstractThis work investigates the unexplored usability of self-supervised representation learning in the direction of functional knowledge transfer. In this work, functional knowledge transfer is achieved by joint optimization of self-supervised learning pseudo task and supervised learning task, improving supervised learning task performance. Recent progress in self-supervised learning uses a large volume of data, which becomes a constraint for its applications on small-scale datasets. This work shares a simple yet effective joint training framework that reinforces human-supervised task learning by learning self-supervised representations just-in-time and vice versa. Experiments on three public datasets from different visual domains, Intel Image, CIFAR, and APTOS, reveal a consistent track of performance improvements on classification tasks during joint optimization. Qualitative analysis also supports the robustness of learnt representations. Source code and trained models are available on GitHub1. Prakash Chandra Chhipa, Muskaan Chopra, Gopal Mengi, Varun Gupta 0005, Richa Upadhyay, Meenakshi Subhash Chippa, Kanjar De, Rajkumar Saini, Seiichi Uchida, Marcus Liwicki |
ICIP | 10 |
| 2023 | AfriWOZ: Corpus for Exploiting Cross-Lingual Transfer for Dialogue Generation in Low-Resource, African LanguagesabstractDialogue generation is an important NLP task fraught with many challenges. The challenges become more daunting for low-resource African languages. To enable the creation of dialogue agents for African languages, we contribute the first high-quality dialogue datasets for 6 African languages: Swahili, Wolof, Hausa, Nigerian Pidgin English, Kinyarwanda & Yorùbá. There are a total of 9,000 turns, each language having 1,500 turns, which we translate from a portion of the English multi-domain MultiWOZ dataset. Subsequently, we benchmark by investigating & analyzing the effectiveness of modelling through transfer learning by utilziing state-of-the-art (SoTA) deep monolingual models: DialoGPT and BlenderBot. We compare the models with a simple seq2seq baseline using perplexity. Besides this, we conduct human evaluation of single-turn conversations by using majority votes and measure inter-annotator agreement (IAA). We find that the hypothesis that deep monolingual models learn some abstractions that generalize across languages holds. We observe human-like conversations, to different degrees, in 5 out of the 6 languages. The language with the most transferable properties is the Nigerian Pidgin English, with a human-likeness score of 78.1%, of which 34.4% are unanimous. We freely provide the datasets and host the model checkpoints/demos on the HuggingFace hub for public access. Tosin P. Adewumi, Mofe Adeyemi, Aremu Anuoluwapo, Bukola Peters, Happy Buzaaba, Oyerinde Samuel, Amina Mardiyyah Rufai, Benjamin Ajibade, Tajudeen Gwadabe, Mory Moussou Koulibaly Traore, Tunde Ajayi, Shamsuddeen Hassan Muhammad, Ahmed Baruwa, Paul Owoicho, Tolúlopé Ògúnrèmí, Phylis Ngigi, Orevaoghene Ahia, Ruqayya Nasir, Foteini Liwicki, Marcus Liwicki |
IJCNN | 20 |
| 2023 | Domain Adaptable Self-supervised Representation Learning on Remote Sensing Satellite ImageryabstractThis work presents a novel domain adaption paradigm for studying contrastive self-supervised representation learning and knowledge transfer using remote sensing satellite data. Major state-of-the-art remote sensing visual domain ef-forts primarily focus on fully supervised learning approaches that rely entirely on human annotations. On the other hand, human annotations in remote sensing satellite imagery are always subject to limited quantity due to high costs and domain expertise, making transfer learning a viable alternative. The proposed approach investigates the knowledge transfer of self-supervised representations across the distinct source and target data distributions in depth in the remote sensing data domain. In this arrangement, self-supervised contrastive learning- based pretraining is performed on the source dataset, and downstream tasks are performed on the target datasets in a round-robin fashion. Experiments are conducted on three publicly avail-able datasets, UC Merced Landuse (UCMD), SIRI-WHU, and MLRSNet, for different downstream classification tasks versus label efficiency. In self-supervised knowledge transfer, the pro-posed approach achieves state-of-the-art performance with label efficiency labels and outperforms a fully supervised setting. A more in-depth qualitative examination reveals consistent evidence for explainable representation learning. The source code and trained models are published on GitHub1. Muskaan Chopra, Prakash Chandra Chhipa, Gopal Mengi, Varun Gupta 0005, Marcus Liwicki |
IJCNN | 5 |
| 2023 | Learning Self-Supervised Representations for Label Efficient Cross-Domain Knowledge Transfer on Diabetic Retinopathy Fundus ImagesabstractThis work presents a novel label-efficient self-supervised representation learning-based approach for classifying diabetic retinopathy (DR) images in cross-domain settings. Most of the existing DR image classification methods are based on supervised learning which requires a lot of time-consuming and expensive medical domain experts-annotated data for training. The proposed approach uses the prior learning from the source DR image dataset to classify images drawn from the target datasets. The image representations learned from the unlabeled source domain dataset through contrastive learning are used to classify DR images from the target domain dataset. Moreover, the proposed approach requires a few labeled images to perform successfully on DR image classification tasks in cross-domain settings. The proposed work experiments with four publicly available datasets: EyePACS, APTOS 2019, MESSIDOR-I, and Fundus Images for self-supervised representation learning-based DR image classification in cross-domain settings. The proposed method achieves state-of-the-art results on binary and multi-classification of DR images, even in cross-domain settings. The proposed method outperforms the existing DR image binary and multi-class classification methods proposed in the literature. The proposed method is also validated qualitatively using class activation maps, revealing that the method can learn explainable image representations. The source code and trained models are published on GitHub11https://github.com/prakashchhipa/Learning-Self-Supervised-Representations-for-Label-Efficient-Cross-Domain-Knowledge-Transfer-on-DRF. Ekta Gupta, Varun Gupta 0005, Muskaan Chopra, Prakash Chandra Chhipa, Marcus Liwicki |
IJCNN | 5 |
| 2023 | Multi-Task Meta Learning: learn how to adapt to unseen tasksabstractThis work proposes Multi-task Meta Learning (MTML), integrating two learning paradigms Multi-Task Learning (MTL) and meta learning, to bring together the best of both worlds. In particular, it focuses simultaneous learning of multiple tasks, an element of MTL and promptly adapting to new tasks, a quality of meta learning. It is important to highlight that we focus on heterogeneous tasks, which are of distinct kind, in contrast to typically considered homogeneous tasks (e.g., if all tasks are classification or if all tasks are regression tasks). The fundamental idea is to train a multi-task model, such that when an unseen task is introduced, it can learn in fewer steps whilst offering a performance at least as good as conventional single task learning on the new task or inclusion within the MTL. By conducting various experiments, we demonstrate this paradigm on two datasets and four tasks: NYU-v2 and the taskonomy dataset for which we perform semantic segmentation, depth estimation, surface normal estimation, and edge detection. MTML achieves state-of-the-art results for three out of four tasks for the NYU-v2 dataset and two out of four for the taskonomy dataset. In the taskonomy dataset, it was discovered that many pseudo-labeled segmentation masks lacked classes that were expected to be present in the ground truth; however, our MTML approach was found to be effective in detecting these missing classes, delivering good qualitative results. While, quantitatively its performance was affected due to the presence of incorrect ground truth labels. The the source code for reproducibility can be found at https://github.com/ricupa/MTML-learn-how-to-adapt-to-unseen-tasks. Richa Upadhyay, Prakash Chandra Chhipa, Ronald Phlypo, Rajkumar Saini, Marcus Liwicki |
IJCNN | 5 |
| 2023 | Magnification Prior: A Self-Supervised Method for Learning Representations on Breast Cancer Histopathological ImagesabstractThis work presents a novel self-supervised pre-training method to learn efficient representations without labels on histopathology medical images utilizing magnification factors. Other state-of-the-art works mainly focus on fully supervised learning approaches that rely heavily on human annotations. However, the scarcity of labeled and unlabeled data is a long-standing challenge in histopathology. Currently, representation learning without labels remains unexplored in the histopathology domain. The proposed method, Magnification Prior Contrastive Similarity (MPCS), enables self-supervised learning of representations without labels on small-scale breast cancer dataset BreakHis by exploiting magnification factor, inductive transfer, and reducing human prior. The proposed method matches fully supervised learning state-of-the-art performance in malignancy classification when only 20% of labels are used in fine-tuning and outperform previous works in fully supervised learning settings for three public breast cancer datasets, including BreakHis. Further, It provides initial support for a hypothesis that reducing human-prior leads to efficient representation learning in self-supervision, which will need further investigation. The implementation of this work is available online on GitHub1. Prakash Chandra Chhipa, Richa Upadhyay, Gustav Pihlgren, Rajkumar Saini, Seiichi Uchida, Marcus Liwicki |
WACV | 6 |
| 2022 | Investigating the Effect of Using Synthetic and Semi-synthetic Images for Historical Document Font Classification
Konstantina Nikolaidou, Richa Upadhyay, Mathias Seuret, Marcus Liwicki |
DAS | 4 |
| 2022 | HaT5: Hate Language Identification using Text-to-Text Transfer TransformerabstractWe investigate the performance of a state-of-the-art (SoTA) architecture T5 (available on the SuperGLUE) and compare it with 3 other previous SoTA architectures across 5 different tasks from 2 relatively diverse datasets. The datasets are diverse in terms of the number and types of tasks they have. To improve performance, we augment the training data by using a new autoregressive conversational AI model checkpoint. We achieve near-SoTA results on a couple of the tasks - macro F1 scores of 81.66% for task A of the OLID 2019 dataset and 82.54% for task A of the hate speech and offensive content (HASOC) 2021 dataset, where SoTA are 82.9% and 83.05%, respectively. We perform error analysis and explain why one of the models (Bi-LSTM) makes the predictions it does by using a publicly available algorithm: Integrated Gradient (IG). This is because explainable artificial intelligence (XAI) is essential for earning the trust of users. The main contributions of this work are the implementation method of T5, which is discussed; the data augmentation, which brought performance improvements; and the revelation on the shortcomings of the HASOC 2021 dataset. The revelation shows the difficulties of poor data annotation by using a small set of examples where the T5 model made the correct predictions, even when the ground truth of the test set were incorrect (in our opinion). We also provide our model checkpoints on the HuggingFace hub11https://huggingface.co/sana-ngu/HaT5_augmentation https://huggingface.co/sana-ngu/HaT5. Sana Sabah Sabry, Tosin P. Adewumi, Nosheen Abid, György Kovács 0001, Foteini Liwicki, Marcus Liwicki |
IJCNN | 6 |
| 2022 | Potential Idiomatic Expression (PIE)-English: Corpus for Classes of IdiomsabstractWe present a fairly large, Potential Idiomatic Expression (PIE) dataset for Natural Language Processing (NLP) in English. The challenges with NLP systems with regards to tasks such as Machine Translation (MT), word sense disambiguation (WSD) and information retrieval make it imperative to have a labelled idioms dataset with classes such as it is in this work. To the best of the authors’ knowledge, this is the first idioms corpus with classes of idioms beyond the literal and the general idioms classification. In particular, the following classes are labelled in the dataset: metaphor, simile, euphemism, parallelism, personification, oxymoron, paradox, hyperbole, irony and literal. We obtain an overall inter-annotator agreement (IAA) score, between two independent annotators, of 88.89%. Many past efforts have been limited in the corpus size and classes of samples but this dataset contains over 20,100 samples with almost 1,200 cases of idioms (with their meanings) from 10 classes (or senses). The corpus may also be extended by researchers to meet specific needs. The corpus has part of speech (PoS) tagging from the NLTK library. Classification experiments performed on the corpus to obtain a baseline and comparison among three common models, including the BERT model, give good results. We also make publicly available the corpus and the relevant codes for working with it for NLP tasks. Tosin P. Adewumi, Roshanak Vadoodi, Aparajita Tripathy, Konstantina Nikolaidou, Foteini Liwicki, Marcus Liwicki |
LREC | 6 |
| 2022 | A survey of historical document image datasetsabstractAbstract This paper presents a systematic literature review of image datasets for document image analysis, focusing on historical documents, such as handwritten manuscripts and early prints. Finding appropriate datasets for historical document analysis is a crucial prerequisite to facilitate research using different machine learning algorithms. However, because of the very large variety of the actual data (e.g., scripts, tasks, dates, support systems, and amount of deterioration), the different formats for data and label representation, and the different evaluation processes and benchmarks, finding appropriate datasets is a difficult task. This work fills this gap, presenting a meta-study on existing datasets. After a systematic selection process (according to PRISMA guidelines), we select 65 studies that are chosen based on different factors, such as the year of publication, number of methods implemented in the article, reliability of the chosen algorithms, dataset size, and journal outlet. We summarize each study by assigning it to one of three pre-defined tasks: document classification, layout structure, or content analysis. We present the statistics, document type, language, tasks, input visual aspects, and ground truth information for every dataset. In addition, we provide the benchmark tasks and results from these papers or recent competitions. We further discuss gaps and challenges in this domain. We advocate for providing conversion tools to common formats (e.g., COCO format for computer vision tasks) and always providing a set of evaluation metrics, instead of just one, to make results comparable across studies. Konstantina Nikolaidou, Mathias Seuret, Hamam Mokayed, Marcus Liwicki |
Int. J. Document Anal. Recognit. | 4 |
| 2022 | Robust Scene Text Detection for Partially Annotated Training DataabstractThis article analyzed the impact of training data containing un-annotated text instances, i.e., partial annotation in scene text detection, and proposed a text region refinement approach to address it. Scene text detection is a problem that has attracted the attention of the research community for decades. Impressive results have been obtained for fully supervised scene text detection with recent deep learning approaches. These approaches, however, need a vast amount of completely labeled datasets, and the creation of such datasets is a challenging and time-consuming task. Research literature lacks the analysis of the partial annotation of training data for scene text detection. We have found that the performance of the generic scene text detection method drops significantly due to the partial annotation of training data. We have proposed a text region refinement method that provides robustness against the partially annotated training data in scene text detection. The proposed method works as a two-tier scheme. Text-probable regions are obtained in the first tier by applying hybrid loss that generates pseudo-labels to refine text regions in the second-tier during training. Extensive experiments have been conducted on a dataset generated from ICDAR 2015 by dropping the annotations with various drop rates and on a publicly available SVT dataset. The proposed method exhibits a significant improvement over the baseline and existing approaches for the partially annotated training data. Prateek Keserwani, Rajkumar Saini, Marcus Liwicki, Partha Pratim Roy 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | A Blended Attention-CTC Network Architecture for Amharic Text-image Recognition
Birhanu Belay, Tewodros Habtegebrial, Marcus Liwicki, Gebeyehu Belay, Didier Stricker |
ICPRAM | 3 |
| 2021 | HyperEmbed: Tradeoffs Between Resources and Performance in NLP Tasks with Hyperdimensional Computing Enabled Embedding of n-gram StatisticsabstractRecent advances in Deep Learning have led to a significant performance increase on several NLP tasks, however, the models become more and more computationally demanding. Therefore, this paper tackles the domain of computationally efficient algorithms for NLP tasks. In particular, it investigates distributed representations of$n$-gram statistics of texts. The representations are formed using hyperdimensional computing enabled embedding. These representations then serve as features, which are used as input to standard classifiers. We investigate the applicability of the embedding on one large and three small standard datasets for classification tasks using nine classifiers. The embedding achieved on par$F_{1}$scores while decreasing the time and memory requirements by several times compared to the conventional$n$-gram statistics, e.g., for one of the classifiers on a small dataset, the memory reduction was 6.18 times; while train and test speed-ups were 4.62 and 3.84 times, respectively. For many classifiers on the large dataset, memory reduction was ca. 100 times and train and test speed-ups were over 100 times. Importantly, the usage of distributed representations formed via hyperdimensional computing allows dissecting strict dependency between the dimensionality of the representation and n-gram size, thus, opening a room for tradeoffs. Kumar Shridhar, Denis Kleyko, Evgeny Osipov, Marcus Liwicki |
IJCNN | 5 |
| 2020 | Cross-Encoded Meta Embedding towards Transfer Learning
György Kovács 0001, Rickard Brännvall, Johan Öhman, Marcus Liwicki |
ESANN | 4 |
| 2020 | Data Fusion and Artificial Neural Networks for Modelling Crop Disease SeverityabstractThis paper analyzes the possibility of applying data fusion combined with artificial neural networks (ANN) on a dataset combining hard and soft data for prediction of one of the most devastating crop diseases of winter wheat, i.e., Septoria Tritici (Zymoseptoria tritici). In advanced decision support systems for crop protection choices, disease models form a major component. They reproduce the biophysical processes of disease development and temporal spread as a set of rules or processes to predict disease risk value. However, the adaptation of these rules or processes to incorporate the effects of climate change is complex and requires extensive rework. To remedy this issue, statistical machine learning techniques have been introduced to model disease severity percentage for some diseases. However, the use of artificial neural networks has been limited (mainly to image data) and is unexplored for Septoria Tritici. This paper explores the use of Feed Forward neural networks on fused tabular data for the task of disease severity modelling. First, ten years of trial data ranging from 2008 to 2018 across Europe is used for the creation of the new tabular dataset with a fusion of all important data sources baring impact on disease development: Field-specific data, weather data, crop growth stages, and disease severity observation made by human trial operators (response variable). Next, two implementation architectures of Feed Forward neural networks on tabular data are employed: a) standard architecture with backpropagation, drop out regularization, and batch normalization and b) advanced architecture with improvements such as cyclic learning rate and cosine annealing. The advanced architecture is able to better model the data and make estimations of disease severity with a difference of +-10% giving a better quantifiable estimate of disease stress. For better outreach to farmers, a technique to incorporate such modelling techniques into the well established Decision Support Systems is also presented. Priyamvada Shankar, Andreas Johnen, Marcus Liwicki |
FUSION | 3 |
| 2020 | Trainable Spectrally Initializable Matrix Transformations in Convolutional Neural NetworksabstractIn this work, we introduce a new architectural component to Neural Network (NN), i.e., trainable and spectrally initializable matrix transformations on feature maps. While previous literature has already demonstrated the possibility of adding static spectral transformations as feature processors, our focus is on more general trainable transforms. We study the transforms in various architectural configurations on four datasets of different nature: from medical (ColorectalHist, HAM10000) and natural (Flowers) images to historical documents (CB55). With rigorous experiments that control for the number of parameters and randomness, we show that networks utilizing the introduced matrix transformations outperform vanilla neural networks. The observed accuracy increases appreciably across all datasets. In addition, we show that the benefit of spectral initialization leads to significantly faster convergence, as opposed to randomly initialized matrix transformations. The transformations are implemented as auto-differentiable PyTorch modules that can be incorporated into any neural network architecture. The entire code base is open-source. Michele Alberti, Angela Botros, Narayan Schütz, Rolf Ingold, Marcus Liwicki, Mathias Seuret |
ICPR | 5 |
| 2020 | Pretraining Image Encoders without Reconstruction via Feature Prediction LossabstractThis work investigates three methods for calculating loss for autoencoder-based pretraining of image encoders: The commonly used reconstruction loss, the more recently introduced deep perceptual similarity loss, and a feature prediction loss proposed here; the latter turning out to be the most efficient choice. Standard auto-encoder pretraining for deep learning tasks is done by comparing the input image and the reconstructed image. Recent work shows that predictions based on embeddings generated by image autoencoders can be improved by training with perceptual loss, i.e., by adding a loss network after the decoding step. So far the autoencoders trained with loss networks implemented an explicit comparison of the original and reconstructed images using the loss network. However, given such a loss network we show that there is no need for the time-consuming task of decoding the entire image. Instead, we propose to decode the features of the loss network, hence the name “feature prediction loss”. To evaluate this method we perform experiments on three standard publicly available datasets (LunarLander-v2, STL-10, and SVHN) and compare six different procedures for training image encoders (pixel-wise, perceptual similarity, and feature prediction losses; combined with two variations of image and feature encoding/decoding). The embedding-based prediction results show that encoders trained with feature prediction loss is as good or better than those trained with the other two losses. Additionally, the encoder is significantly faster to train using feature prediction loss in comparison to the other losses. The method implementation used in this work is available online. Gustav Pihlgren, Fredrik Sandin, Marcus Liwicki |
ICPR | 3 |
| 2020 | Improving Image Autoencoder Embeddings with Perceptual LossabstractAutoencoders are commonly trained using element-wise loss. However, element-wise loss disregards high-level structures in the image which can lead to embeddings that disregard them as well. A recent improvement to autoencoders that helps alleviate this problem is the use of perceptual loss. This work investigates perceptual loss from the perspective of encoder embeddings themselves. Autoencoders are trained to embed images from three different computer vision datasets using perceptual loss based on a pretrained model as well as pixel-wise loss. A host of different predictors are trained to perform object positioning and classification on the datasets given the embedded images as input. The two kinds of losses are evaluated by comparing how the predictors performed with embeddings from the differently trained autoencoders. The results show that, in the image domain, the embeddings generated by autoencoders trained with perceptual loss enable more accurate predictions than those trained with element-wise loss. Furthermore, the results show that, on the task of object positioning of a small-scale feature, perceptual loss can improve the results by a factor 10. The experimental setup is available online1. Gustav Pihlgren, Fredrik Sandin, Marcus Liwicki |
IJCNN | 3 |
| 2019 | Labeling, Cutting, Grouping: An Efficient Text Line Segmentation Method for Medieval ManuscriptsabstractThis paper introduces a new way for text-line extraction by integrating deep-learning based pre-classification and state-of-the-art segmentation methods. Text-line extraction in complex handwritten documents poses a significant challenge, even to the most modern computer vision algorithms. Historical manuscripts are a particularly hard class of documents as they present several forms of noise, such as degradation, bleed-through, interlinear glosses, and elaborated scripts. In this work, we propose a novel method which uses semantic segmentation at pixel level as intermediate task, followed by a text-line extraction step. We measured the performance of our method on a recent dataset of challenging medieval manuscripts and surpassed state-of-the-art results by reducing the error by 80.7%. Furthermore, we demonstrate the effectiveness of our approach on various other datasets written in different scripts. Hence, our contribution is two-fold. First, we demonstrate that semantic pixel segmentation can be used as strong denoising pre-processing step before performing text line extraction. Second, we introduce a novel, simple and robust algorithm that leverages the high-quality semantic segmentation to achieve a text-line extraction performance of 99.42% line IU on a challenging dataset. Michele Alberti, Lars Vögtlin, Vinaychandran Pondenkandath, Mathias Seuret, Rolf Ingold, Marcus Liwicki |
ICDAR | 6 |
| 2019 | Amharic Text Image Recognition: Database, Algorithm, and AnalysisabstractThis paper introduces a dataset for an exotic, but very interesting script, Amharic. Amharic follows a unique syllabic writing system which uses 33 consonant characters with their 7 vowels variants of each. Some labialized characters derived by adding diacritical marks on consonants and or removing part of it. These associated diacritics on consonant characters are relatively smaller in size and challenging to distinguish the derived (vowel and labialized) characters. In this paper we tackle the problem of Amharic text-line image recognition. In this work, we propose a recurrent neural network based method to recognize Amharic text-line images. The proposed method uses Long Short Term Memory (LSTM) networks together with CTC (Connectionist Temporal Classification). Furthermore, in order to overcome the lack of annotated data, we introduce a new dataset that contains 337,332 Amharic text-line images which is made freely available at http://www.dfki.uni-kl.de/~belay/. The performance of the proposed Amharic OCR model is tested by both printed and synthetically generated datasets, and promising results are obtained. Birhanu Belay, Tewodros Habtegebrial, Marcus Liwicki, Gebeyehu Belay, Didier Stricker |
ICDAR | 3 |
| 2019 | ICDAR 2019 Historical Document Reading Challenge on Large Structured Chinese Family RecordsabstractIn this paper, we present a large historical database of Chinese family records with the aim to develop robust systems for historical document analysis. In this direction, we propose a Historical Document Reading Challenge on Large Chinese Structured Family Records (ICDAR 2019 HDRC-CHINESE). The objective of the competition is to recognize and analyze the layout, and finally detect and recognize the textlines and characters of the large historical document image dataset containing more than 100000 pages. Cascade R-CNN, CRNN, and U-Net based architectures were trained to evaluate the performances in these tasks. Error rate of 0.01 has been recorded for textline recognition (Task1) whereas a Jaccard Index of 99:54% has been recorded for layout analysis (Task2). The graph edit distance based total error ratio of 1:5% has been recorded for complete integrated textline detection and recognition (Task3). Rajkumar Saini, Derek Dobson, Jon Morrey, Marcus Liwicki, Foteini Liwicki |
ICDAR | 4 |
| 2019 | A Comprehensive Study of ImageNet Pre-Training for Historical Document Image AnalysisabstractAutomatic analysis of scanned historical documents comprises a wide range of image analysis tasks, which are often challenging for machine learning due to a lack of human-annotated learning samples. With the advent of deep neural networks, a promising way to cope with the lack of training data is to pre-train models on images from a different domain and then fine-tune them on historical documents. In the current research, a typical example of such cross-domain transfer learning is the use of neural networks that have been pre-trained on the ImageNet database for object recognition. It remains a mostly open question whether or not this pre-training helps to analyse historical documents, which have fundamentally different image properties when compared with ImageNet. In this paper, we present a comprehensive empirical survey on the effect of ImageNet pre-training for diverse historical document analysis tasks, including character recognition, style classification, manuscript dating, semantic segmentation, and content-based retrieval. While we obtain mixed results for semantic segmentation at pixel-level, we observe a clear trend across different network architectures that ImageNet pre-training has a positive effect on classification as well as content-based retrieval. Linda Studer, Michele Alberti, Vinaychandran Pondenkandath, Pinar Goktepe, Thomas Kolonko, Andreas Fischer 0002, Marcus Liwicki, Rolf Ingold |
ICDAR | 7 |
| 2019 | Factored Convolutional Neural Network for Amharic Character Image RecognitionabstractIn this paper we propose a novel CNN based approach for Amharic character image recognition. The proposed method is designed by leveraging the structure of Amharic graphemes. Amharic characters could be decomposed in to a consonant and a vowel. As a result of this consonant-vowel combination structure, Amharic characters lie within a matrix structure called 'Fidel Gebeta'. The rows and columns of 'Fidel Gebeta' correspond to a character's consonant and the vowel components, respectively. The proposed method has a CNN architecture with two classifiers that detect the row/consonant and column/vowel components of a character. The two classifiers share a common feature space before they fork-out at their last layers. The method achieves state-of-the-art result on a synthetically generated dataset. The proposed method achieves 94.97% overall character recognition accuracy. Birhanu Belay, Tewodros Habtegebrial, Marcus Liwicki, Gebeyehu Belay, Didier Stricker |
ICIP | 3 |
| 2019 | Bidirectional Learning for Robust Neural NetworksabstractA multilayer perceptron can behave as a generative classifier by applying bidirectional learning (BL). It consists of training an undirected neural network to map input to output and vice-versa; therefore it can produce a classifier in one direction, and a generator in the opposite direction for the same data. The learning process of BL tries to reproduce the neuroplasticity stated in Hebbian theory using only backward propagation of errors. In this paper, two learning techniques are independently introduced which use BL for improving robustness to white noise static and adversarial examples. The first method is bidirectional propagation of errors, which the error propagation occurs in backward and forward directions. Motivated by the fact that its generative model receives as input a constant vector per class, we introduce as a second method the novel hybrid adversarial networks (HAN). Its generative model receives a random vector as input and its training is based on generative adversarial networks (GAN). To assess the performance of BL, we perform experiments using several architectures with fully and convolutional layers, with and without bias. Experimental results show that both methods improve robustness to white noise static and adversarial examples, and even increase accuracy, but have different behavior depending on the architecture and task, being more beneficial to use the one or the other. Nevertheless, HAN using a convolutional architecture with batch normalization presents outstanding robustness, reaching state-of-the-art accuracy on adversarial examples of hand-written digits. Sidney Pontes-Filho, Marcus Liwicki |
IJCNN | 2 |
| 2019 | Subword Semantic Hashing for Intent Classification on Small DatasetsabstractIn this paper, we introduce the use of Semantic Hashing as embedding for the task of Intent Classification and achieve state-of-the-art performance on three frequently used benchmarks. Intent Classification on a small dataset is a challenging task for data-hungry state-of-the-art Deep Learning based systems. Semantic Hashing is an attempt to overcome such a challenge and learn robust text classification. Current word embedding based methods [11], [13], [14] are dependent on vocabularies. One of the major drawbacks of such methods is out-of-vocabulary terms, especially when having small training datasets and using a wider vocabulary. This is the case in Intent Classification for chatbots, where typically small datasets are extracted from internet communication. Two problems arise with the use of internet communication. First, such datasets miss a lot of terms in the vocabulary to use word embeddings efficiently. Second, users frequently make spelling errors. Typically, the models for intent classification are not trained with spelling errors and it is difficult to think about ways in which users will make mistakes. Models depending on a word vocabulary will always face such issues. An ideal classifier should handle spelling errors inherently. With Semantic Hashing, we overcome these challenges and achieve state-of-the-art results on three datasets: Chatbot, Ask Ubuntu, and Web Applications [3]. Our benchmarks are available online. Kumar Shridhar, Ayushman Dash, Amit Sahu, Gustav Pihlgren, Vinaychandran Pondenkandath, György Kovács 0001, Foteini Liwicki, Marcus Liwicki |
IJCNN | 9 |
| 2019 | Examining the Combination of Multi-Band Processing and Channel Dropout for Robust Speech Recognitionabstractsponsorship: Laszlo Toth was supported by the J ' anos Bolyai Research Scholarship of the Hungarian Academy of Sciences and the UNKP19-4 New Excellence Program of the Hungarian Ministry of Innovation and Technology. (Hungarian Academy of Sciences, New Excellence Program of the Hungarian Ministry of Innovation and Technology|UNKP19-4) György Kovács 0001, László Tóth 0001, Dirk Van Compernolle, Marcus Liwicki |
INTERSPEECH | 4 |
| 2019 | Grayification: A meaningful grayscale conversion to improve handwritten historical documents analysis
Manuel Bouillon, Rolf Ingold, Marcus Liwicki |
Pattern Recognit. Lett. | 3 |
| 2019 | Combining graph edit distance and triplet networks for offline signature verification
Paul Maergner, Vinaychandran Pondenkandath, Michele Alberti, Marcus Liwicki, Kaspar Riesen, Rolf Ingold, Andreas Fischer 0002 |
Pattern Recognit. Lett. | 4 |
| 2018 | A Semi-automatized Modular Annotation Tool for Ancient Manuscript AnnotationabstractIn this paper, we present DIVAnnotation, an ancient document annotation tool which is freely available as open source. This software is easily modular thanks to the splitting of the different annotation steps through the use of a tabbed graphical user interface. State-of-the-art document image analysis methods are included through web services, thus allowing users to generate automatically annotations and correct them manually when needed. The annotations are stored into a highly structured TEI file which makes data access and manipulation simple. A Java library for managing TEI files generated by DIVAnnotation is also provided as open source. Mathias Seuret, Manuel Bouillon, Foteini Liwicki, Marcel Gygli, Marcus Liwicki, Rolf Ingold |
DAS | 5 |
| 2018 | Web Services in Document Image Analysis - Recent Developments on DIVAServices and the Importance of Building an EcosystemabstractWeb Services are being adapted into the workflows of many Document Image Analysis researchers. However, so far, there is no common platform for providing access to algorithms in the community. DIVAServices aims to become this by providing a platform that is open to the whole community to provide their own methods as Web Services. In this paper we present updates and enhancements made to the existing DIVAServices platform. This includes a new computational backend, a revamped execution workflow based on asynchronous communication, and the possibility for methods to specify their outputs. Furthermore, we discuss the importance of an ecosystem for such platforms. We argue that only providing a RESTful API is not enough. Users need tools and services around the framework that support them in adapting the Web Services and we introduce some of the tools that we built around DIVAServices. Marcel Gygli, Marcus Liwicki, Rolf Ingold |
DAS | 2 |
| 2018 | DeepDIVA: A Highly-Functional Python Framework for Reproducible ExperimentsabstractWe introduce DeepDIVA: an infrastructure designed to enable quick and intuitive setup of reproducible experiments with a large range of useful analysis functionality. Reproducing scientific results can be a frustrating experience, not only in document image analysis but in machine learning in general. Using DeepDIVA a researcher can either reproduce a given experiment or share their own experiments with others. Moreover, the framework offers a large range of functions, such as boilerplate code, keeping track of experiments, hyper-parameter optimization, and visualization of data and results. To demonstrate the effectiveness of this framework, this paper presents case studies in the area of handwritten document analysis where researchers benefit from the integrated functionality. DeepDIVA is implemented in Python and uses the deep learning framework PyTorch. It is completely open source, and accessible as Web Service through DIVAServices. Michele Alberti, Vinaychandran Pondenkandath, Marcel Gygli, Rolf Ingold, Marcus Liwicki |
ICFHR | 5 |
| 2018 | Recognizing Challenging Handwritten Annotations with Fully Convolutional NetworksabstractThis paper introduces a very challenging dataset of historic German documents and evaluates Fully Convolutional Neural Network (FCNN) based methods to locate handwritten annotations of any kind in these documents. The handwritten annotations can appear in form of underlines and text by using various writing instruments, e.g., the use of pencils makes the data more challenging. We train and evaluate various end-to-end semantic segmentation approaches and report the results. The task is to classify the pixels of documents into two classes: background and handwritten annotation. The best model achieves a mean Intersection over Union (IOU) score of 95.6% on the test documents of the presented dataset. We also present a comparison of different strategies used for data augmentation and training on our presented dataset. For evaluation, we use the Layout Analysis Evaluator for the ICDAR 2017 Competition on Layout Analysis for Challenging Medieval Manuscripts. Andreas Kölsch, Saurabh Varshneya, Muhammad Zeshan Afzal, Marcus Liwicki |
ICFHR | 5 |
| 2018 | Identifying Cross-Depicted Historical MotifsabstractCross-depiction is the problem of identifying the same object even when it is depicted in a variety of manners.This is a common problem in handwritten historical document image analysis, for instance when the same letter or motif is depicted in several different ways. It is a simple task for humans yet conventional computer vision methods struggle to cope with it. In this paper we address this problem using state-of-the-art deep learning techniques on a dataset of historical watermarks containing images created with different methods of reproduction, such as hand tracing, rubbing, and radiography.To study the robustness of deep learning based approaches to the cross-depiction problem, we measure their performance on two different tasks: classification and similarity rankings. For the former we achieve a classification accuracy of 96 % using deep convolutional neural networks. For the latter we have a false positive rate at 95% recall of 0.11. These results outperform state-of-the-art methods by a significant margin. Vinaychandran Pondenkandath, Michele Alberti, Nicole Eichenberger, Rolf Ingold, Marcus Liwicki |
ICFHR | 5 |
| 2018 | Symbol Grounding Association in Multimodal Sequences with Missing ElementsabstractIn this paper, we extend a symbolic association framework for being able to handle missing elements in multimodal sequences. The general scope of the work is the symbolic associations of object-word mappings as it happens in language development in infants. In other words, two different representations of the same abstract concepts can associate in both directions. This scenario has been long interested in Artificial Intelligence, Psychology, and Neuroscience. In this work, we extend a recent approach for multimodal sequences (visual and audio) to also cope with missing elements in one or both modalities. Our method uses two parallel Long Short-Term Memories (LSTMs) with a learning rule based on EM-algorithm. It aligns both LSTM outputs via Dynamic Time Warping (DTW). We propose to include an extra step for the combination with the max operation for exploiting the common elements between both sequences. The motivation behind is that the combination acts as a condition selector for choosing the best representation from both LSTMs. We evaluated the proposed extension in the following scenarios: missing elements in one modality (visual or audio) and missing elements in both modalities (visual and sound). The performance of our extension reaches better results than the original model and similar results to individual LSTM trained in each modality. Federico Raue, Andreas Dengel 0001, Thomas M. Breuel, Marcus Liwicki |
J. Artif. Intell. Res. | 4 |
| 2017 | Character-Level Dialect Identification in Arabic Using Long Short-Term Memory
Karim Sayadi, Mansour Hamidi, Marc Bui, Marcus Liwicki, Andreas Fischer 0002 |
CICLing (2) | 4 |
| 2017 | Classless Association Using Neural Networks
Federico Raue, Sebastian Palacio, Andreas Dengel 0001, Marcus Liwicki |
ICANN (2) | 4 |
| 2017 | Cutting the Error by Half: Investigation of Very Deep CNN and Advanced Training Strategies for Document Image ClassificationabstractWe present an exhaustive investigation of recent Deep Learning architectures, algorithms, and strategies for the task of document image classification to finally reduce the error by more than half. Existing approaches, such as the DeepDoc-Classifier, apply standard Convolutional Network architectures with transfer learning from the object recognition domain. The contribution of the paper is threefold: First, it investigates recently introduced very deep neural network architectures (GoogLeNet, VGG, ResNet) using transfer learning (from real images). Second, it proposes transfer learning from a huge set of document images, i.e. 400; 000 documents. Third, it analyzes the impact of the amount of training data (document images) and other parameters to the classification abilities. We use two datasets, the Tobacco-3482 and the large-scale RVL-CDIP dataset. We achieve an accuracy of 91:13% for the Tobacco-3482 dataset while earlier approaches reach only 77:6%. Thus, a relative error reduction of more than 60% is achieved. For the large dataset RVL-CDIP, an accuracy of 90:97% is achieved, corresponding to a relative error reduction of 11:5%. Muhammad Zeshan Afzal, Andreas Kölsch, Sheraz Ahmed, Marcus Liwicki |
ICDAR | 4 |
| 2017 | Real-Time Document Image Classification Using Deep CNN and Extreme Learning MachinesabstractThis paper presents an approach for real-time training and testing for document image classification. In production environments, it is crucial to perform accurate and (time-)efficient training. Existing deep learning approaches for classifying documents do not meet these requirements, as they require much time for training and fine-tuning the deep architectures. Motivated from Computer Vision, we propose a two-stage approach. The first stage trains a deep network that works as feature extractor and in the second stage, Extreme Learning Machines (ELMs) are used for classification. The proposed approach outperforms all previously reported structural and deep learning based methods with a final accuracy of 83.24% on Tobacco-3482 dataset, leading to a relative error reduction of 25% when compared to a previous Convolutional Neural Network (CNN) based approach (DeepDocClassifier). More importantly, the training time of the ELM is only 1.176 seconds and the overall prediction time for 2,482 images is 3.066 seconds. As such, this novel approach makes deep learning-based document classification suitable for large-scale real-time applications. Andreas Kölsch, Muhammad Zeshan Afzal, Markus Ebbecke, Marcus Liwicki |
ICDAR | 4 |
| 2017 | PCA-Initialized Deep Neural Networks Applied to Document Image AnalysisabstractIn this paper, we present a novel approach for initializing deep neural networks, i.e., by using Principal Component Analysis (PCA) to initialize neural layers. Usually, the initialization of the weights of a deep neural network is done in one of the three following ways: 1) with random values, 2) layer-wise, usually as Deep Belief Network or as auto-encoder, and 3) re-use of layers from another network (transfer learning). Therefore, typically, many training epochs are needed before meaningful weights are learned, or a rather similar dataset is required for seeding a fine-tuning of transfer learning. In this paper, we describe how to turn a PCA into an auto-encoder, by generating an encoder layer of the PCA parameters and furthermore adding a decoding layer. We analyze the initialization technique on real documents. First, we show that a PCA-based initialization is quick and leads to a very stable initialization. Furthermore, for the task of layout analysis we investigate the effectiveness of PCA-based initialization and show that it outperforms state-of-the-art random weight initialization methods. Mathias Seuret, Michele Alberti, Marcus Liwicki, Rolf Ingold |
ICDAR | 3 |
| 2017 | ICDAR2017 Competition on Layout Analysis for Challenging Medieval ManuscriptsabstractThis paper reports on the ICDAR2017 Competition on Layout Analysis for Challenging Medieval Manuscripts (HisDoc-Layout-Comp) and provides further details and discussions. In this competition we introduce a new challenging dataset and state-of-the-art benchmark results for pixel-labelling and text line segmentation. The DIVA-HisDB comprises medieval manuscripts with complex layout in contrast to previous datasets, where rectangular text blocks and only a few decorative elements exist. In particular, the images of this competition contain many interlinear and marginal glosses as well as texts in various sizes and decorated letters. This makes the distinction of the four target labels (text, comment, decoration, and background) more difficult. In addition, to reflect the needs of scholars in the humanities, we request multi-labeling of certain regions (decorated text as text and decoration). Furthermore, we measure not just the accuracy, but the Intersection over Union (IU) of pixel sets, which better reflects the real performance. Indeed, in our results we observe that the accuracy appears to be rather high, but the IU reveals, that there is still room for improvement. For the task of line segmentation, the recognition results are rather low (overall error higher than 5%). Noteworthy, a combination of the best layout analysis method with an adapted seam-carving based method achieves better results than the best contestant. Foteini Liwicki, Manuel Bouillon, Mathias Seuret, Marcel Gygli, Michele Alberti, Rolf Ingold, Marcus Liwicki |
ICDAR | 7 |
| 2017 | Selecting Fine-Tuned Features for Layout Analysis of Historical DocumentsabstractIn this paper, we investigate fine-tuned features learned by deep neural networks in the context of layout analysis. Pre-training and fine-tuning are techniques used in deep neural networks to learn representations (features) of input. However, it is not clear if the fine-tuned features are all useful for a following classification task. We investigate this problem using feature selection. Firstly, features are learned by a deep neural network, where stacked autoencoders are used for pre-training and then the whole network is fine-tuned. Then, a feature selection method is used to select relevant features for classification. We observe that despite fine-tuning, a significant number of the features are still redundant or irrelevant for layout classification. Furthermore, features from the top layer of the stacked autoencoders are generally more relevant for classification than those from lower layers. Hao Wei 0001, Mathias Seuret, Marcus Liwicki, Rolf Ingold, Pei Fu |
ICDAR | 3 |
| 2017 | Transforming sensor data to the image domain for deep learning - An application to footstep detectionabstractConvolutional Neural Networks (CNNs) have become the state-of-the-art in various computer vision tasks, but they are still premature for most sensor data, especially in pervasive and wearable computing. A major reason for this is the limited amount of annotated training data. In this paper, we propose the idea of leveraging the discriminative power of pre-trained deep CNNs on 2-dimensional sensor data by transforming the sensor modality to the visual domain. By three proposed strategies, 2D sensor output is converted into pressure distribution imageries. Then we utilize a pre-trained CNN for transfer learning on the converted imagery data. We evaluate our method on a gait dataset of floor surface pressure mapping. We obtain a classification accuracy of 87.66%, which outperforms the conventional machine learning methods by over 10%. Monit Shah Singh, Vinaychandran Pondenkandath, Bo Zhou 0005, Paul Lukowicz, Marcus Liwicki |
IJCNN | 5 |
| 2016 | Page Segmentation for Historical Document Images Based on Superpixel Classification with Unsupervised Feature LearningabstractIn this paper, we present an efficient page segmentation method for historical document images. Many existing methods either rely on hand-crafted features or perform rather slow as they treat the problem as a pixel-level assignment problem. In order to create a feasible method for real applications, we propose to use superpixels as basic units of segmentation, and features are learned directly from pixels. An image is first oversegmented into superpixels with the simple linear iterative clustering (SLIC) algorithm. Then, each superpixel is represented by the features of its central pixel. The features are learned from pixel intensity values with stacked convolutional autoencoders in an unsupervised manner. A support vector machine (SVM) classifier is used to classify superpixels into four classes: periphery, background, text block, and decoration. Finally, the segmentation results are refined by a connected component based smoothing procedure. Experiments on three public datasets demonstrate that compared to our previous method, the proposed method is much faster and achieves comparable segmentation results. Additionally, much fewer pixels are used for classifier training. Kai Chen 0011, Mathias Seuret, Marcus Liwicki, Jean Hennebert, Rolf Ingold |
DAS | 4 |
| 2016 | Complete System for Text Line Extraction Using Convolutional Neural Networks and Watershed TransformabstractWe present a novel Convolutional Neural Network based method for the extraction of text lines, which consists of an initial Layout Analysis followed by the estimation of the Main Body Area (i.e., the text area between the baseline and the corpus line) for each text line. Finally, a region-based method using watershed transform is performed on the map of the Main Body Area for extracting the resulting lines. We have evaluated the new system on the IAM-HisDB, a publicly available dataset containing historical documents, outperforming existing learning-based text line extraction methods, which consider the problem as pixel labelling problem into text and non-text regions. Joan Pastor-Pellicer, Muhammad Zeshan Afzal, Marcus Liwicki, María José Castro Bleda |
DAS | 3 |
| 2016 | SDK Reinvented: Document Image Analysis Methods as RESTful Web ServicesabstractDocument Image Analysis (DIA) systems become ever more advanced, but also more complex -- computationally, and logically. This increases the difficulty of integrating existing state-of-the-art approaches into new research or into practical workflows. The current approach to sharing software is publishing source code -- leaving the burden to the integrator -- or creating a Software Development Kit (SDK) which is often restricted to one programming language. We present DIVAServices a framework for sharing and accessing DIA methods within the research community and beyond. Using a RESTful web service architecture we provide access to the methods, leading to only one system on which the binaries of methods need to be maintained. All it takes for a developer to use an algorithm is a simple HTTP request with the image data and parameters for the method and they will receive the computed results in a format that allows for seamless integration into any kind of workflow or for further processing. Furthermore, DIVAServices is open-source, enabling other research groups or libraries to host their own instance in their environment. Using this framework, future DIA systems can be built on the shoulders of well tested algorithms, accessible to everyone. Marcel Gygli, Rolf Ingold, Marcus Liwicki |
DAS | 3 |
| 2016 | Symbolic Association Using Parallel Multilayer Perceptron
Federico Raue, Sebastian Palacio, Thomas M. Breuel, Wonmin Byeon, Andreas Dengel 0001, Marcus Liwicki |
ICANN (2) | 6 |
| 2016 | Page Segmentation for Historical Handwritten Document Images Using Conditional Random FieldsabstractIn this paper, we present a Conditional Random Field (CRF) model to deal with the problem of segmenting handwritten historical document images into different regions. We consider page segmentation as a pixel-labeling problem, i.e., each pixel is assigned to one of a set of labels. Features are learned from pixel intensity values with stacked convolutional autoencoders in an unsupervised manner. The features are used for the purpose of initial classification with a multilayer perceptron. Then a CRF model is introduced for modeling the local and contextual information jointly in order to improve the segmentation. For the purpose of decreasing the time complexity, we perform labeling at superpixel level. In the CRF model, graph nodes are represented by superpixels. The label of each pixel is determined by the label of the superpixel to which it belongs. Experiments on three public datasets demonstrate that, compared to previous methods, the proposed method achieves more accurate segmentation results and is much faster. Kai Chen 0011, Mathias Seuret, Marcus Liwicki, Jean Hennebert, Rolf Ingold |
ICFHR | 3 |
| 2016 | KPTI: Katib's Pashto Text Imagebase and Deep Learning BenchmarkabstractThis paper presents the first Pashto text image database for scientific research and thereby the first dataset with complete handwritten and printed text line images which ultimately covers all alphabets of Arabic and Persian languages. Language like Pashto, written in a complex way by calligraphers, still requires a mature Optical Character Recognition (OCR), system. Although 50 million people use this language both for oral and written communication, there is no significant effort which is devoted to the recognition of Pashto Script. A real dataset of 17,015 images having Pashto text lines is introduced. The images are acquired via scanning from hand scribed Pashto books. Further, in this work, we evaluated the performance of deep learning based models like Bidirectional and Multi-Dimensional Long Short Term Memory (BLSTM and MDLSTM) networks for Pashto texts and provide a baseline character error rate of 9.22%. Riaz Ahmad 0001, Muhammad Zeshan Afzal, Sheikh Faisal Rashid, Marcus Liwicki, Thomas M. Breuel, Andreas Dengel 0001 |
ICFHR | 4 |
| 2016 | N-Light-N: A Highly-Adaptable Java Library for Document Analysis with Convolutional Auto-Encoders and Related ArchitecturesabstractThis paper presents a novel, highly-adaptable Java framework N-light-N, for the work with deep neural networks, especially with CAEs. While the most popular deep learning libraries focus on fast processing and high performance, they only implement the main-stream network architectures and network units. In recent research in the document domain, however, we have shown that modified networks, units, and training processes significantly improve the performance in various tasks. To enable the document research community with such capabilities, in this paper we introduce a novel, publicly available Deep Learning framework which is easy to use, adapt, and extend. Furthermore, we present successful applications for three tasks, including two in the domain of handwritten historical documents, and show how the framework can be used for adaptation, optimization, and deeper analysis. Mathias Seuret, Rolf Ingold, Marcus Liwicki |
ICFHR | 3 |
| 2016 | DIVA-HisDB: A Precisely Annotated Large Dataset of Challenging Medieval ManuscriptsabstractThis paper introduces a publicly available historical manuscript database DIVA-HisDB for the evaluation of several Document Image Analysis (DIA) tasks. The database consists of 150 annotated pages of three different medieval manuscripts with challenging layouts. Furthermore, we provide a layout analysis ground-truth which has been iterated on, reviewed, and refined by an expert in medieval studies. DIVA-HisDB and the ground truth can be used for training and evaluating DIA tasks, such as layout analysis, text line segmentation, binarization and writer identification. Layout analysis results of several representative baseline technologies are also presented in order to help researchers evaluate their methods and advance the frontiers of complex historical manuscripts analysis. An optimized state-of-the-art Convolutional Auto-Encoder (CAE) performs with around 95% accuracy, demonstrating that for this challenging layout there is much room for improvement. Finally, we show that existing text line segmentation methods fail due to interlinear and marginal text elements. Foteini Liwicki, Mathias Seuret, Nicole Eichenberger, Angelika Garz, Marcus Liwicki, Rolf Ingold |
ICFHR | 5 |
| 2016 | Optimizing recurrent reservoirs with neuro-evolution
Sebastian Otte, Martin V. Butz, Danil Koryakin, Fabian Becker, Marcus Liwicki, Andreas Zell |
Neurocomputing | 5 |
| 2015 | Scene labeling with LSTM recurrent neural networksabstractThis paper addresses the problem of pixel-level segmentation and classification of scene images with an entirely learning-based approach using Long Short Term Memory (LSTM) recurrent neural networks, which are commonly used for sequence classification. We investigate two-dimensional (2D) LSTM networks for natural scene images taking into account the complex spatial dependencies of labels. Prior methods generally have required separate classification and image segmentation stages and/or pre- and post-processing. In our approach, classification, segmentation, and context integration are all carried out by 2D LSTM networks, allowing texture and spatial model parameters to be learned within a single model. The networks efficiently capture local and global contextual information over raw RGB values and adapt well for complex scene images. Our approach, which has a much lower computational complexity than prior methods, achieved state-of-the-art performance over the Stanford Background and the SIFT Flow datasets. In fact, if no pre- or post-processing is applied, LSTM networks outperform other state-of-the-art approaches. Hence, only with a single-core Central Processing Unit (CPU), the running time of our approach is equivalent or better than the compared state-of-the-art approaches which use a Graphics Processing Unit (GPU). Finally, our networks' ability to visualize feature maps from each layer supports the hypothesis that LSTM networks are overall suited for image processing tasks. Wonmin Byeon, Thomas M. Breuel, Federico Raue, Marcus Liwicki |
CVPR | 4 |
| 2015 | Learning Recurrent Dynamics using Differential Evolution
Sebastian Otte, Fabian Becker, Martin V. Butz, Marcus Liwicki, Andreas Zell |
ESANN | 4 |
| 2015 | Robust Visual Terrain Classification with Recurrent Neural Networks
Sebastian Otte, Stefan Laible, Richard Hanten, Marcus Liwicki, Andreas Zell |
ESANN | 4 |
| 2015 | Quantifying reading habits: counting how many words you readabstractReading is a very common learning activity, a lot of people perform it everyday even while standing in the subway or waiting in the doctors office. However, we know little about our everyday reading habits, quantifying them enables us to get more insights about better language skills, more effective learning and ultimately critical thinking. This paper presents a first contribution towards establishing a reading log, tracking how much reading you are doing at what time. We present an approach capable of estimating the words read by a user, evaluate it in an user independent approach over 3 experiments with 24 users over 5 different devices (e-ink reader, smartphone, tablet, paper, computer screen). We achieve an error rate as low as 5% (using a medical electrooculography system) or 15% (based on eye movements captured by optical eye tracking) over a total of 30 hours of recording. Our method works for both an optical eye tracking and an Electrooculography system. We provide first indications that the method works also on soon commercially available smart glasses. Kai Kunze, Katsutoshi Masai, Masahiko Inami, Ömer Sacakli, Marcus Liwicki, Andreas Dengel 0001, Shoya Ishimaru, Koichi Kise |
UbiComp | 5 |
| 2015 | Deepdocclassifier: Document classification with deep Convolutional Neural NetworkabstractThis paper presents a deep Convolutional Neural Network (CNN) based approach for document image classification. One of the main requirement of deep CNN architecture is that they need huge number of samples for training. To overcome this problem we adopt a deep CNN which is trained using big image dataset containing millions of samples i.e., ImageNet. The proposed work outperforms both the traditional structure similarity methods and the CNN based approaches proposed earlier. The accuracy of the proposed approach with merely 20 images per class outperforms the state-of-the-art by achieving classification accuracy of 68.25%. The best results on Tobbacoo-3428 dataset show that our proposed method outperforms the state-of-the-art method by a significant margin and achieved a median accuracy of 77.6% with 100 samples per class used for training and validation. Muhammad Zeshan Afzal, Samuele Capobianco, Muhammad Imran Malik, Simone Marinai, Thomas M. Breuel, Andreas Dengel 0001, Marcus Liwicki |
ICDAR | 7 |
| 2015 | Scale and rotation invariant OCR for Pashto cursive script using MDLSTM networkabstractOptical Character Recognition (OCR) of cursive scripts like Pashto and Urdu is difficult due the presence of complex ligatures and connected writing styles. In this paper, we evaluate and compare different approaches for the recognition of such complex ligatures. The approaches include Hidden Markov Model (HMM), Long Short Term Memory (LSTM) network and Scale Invariant Feature Transform (SIFT). Current state of the art in cursive script assumes constant scale without any rotation, while real world data contain rotation and scale variations. This research aims to evaluate the performance of sequence classifiers like HMM and LSTM and compare their performance with descriptor based classifier like SIFT. In addition, we also assess the performance of these methods against the scale and rotation variations in cursive script ligatures. Moreover, we introduce a database of 480,000 images containing 1000 unique ligatures or sub-words of Pashto. In this database, each ligature has 40 scale and 12 rotation variations. The evaluation results show a significantly improved performance of LSTM over HMM and traditional feature extraction technique such as SIFT. Riaz Ahmad 0001, Muhammad Zeshan Afzal, Sheikh Faisal Rashid, Marcus Liwicki, Thomas M. Breuel |
ICDAR | 4 |
| 2015 | Recognizable units in Pashto language for OCRabstractAtomic segmentation of cursive scripts into constituent characters is one of the most challenging problems in pattern recognition. To avoid segmentation in cursive script, concrete shapes are considered as recognizable units. Therefore, the objective of this work is to find out the alternate recognizable units in Pashto cursive script. These alternatives are ligatures and primary ligatures. However, we need sound statistical analysis to find the appropriate numbers of ligatures and primary ligatures in Pashto script. In this work, a corpus of 2, 313, 736 Pashto words are extracted from a large scale diversified web sources, and total of 19, 268 unique ligatures have been identified in Pashto cursive script. Analysis shows that only 7000 ligatures represent 91% portion of overall corpus of the Pashto unique words. Similarly, about 7, 681 primary ligatures are also identified which represent the basic shapes of all the ligatures. Riaz Ahmad 0001, Muhammad Zeshan Afzal, Sheikh Faisal Rashid, Marcus Liwicki, Andreas Dengel 0001, Thomas M. Breuel |
ICDAR | 4 |
| 2015 | Combination of multiple aligned recognition outputs using WFST and LSTMabstractThe contribution of this paper is a new strategy of integrating multiple recognition outputs of diverse recognizers. Such an integration can give higher performance and more accurate outputs than a single recognition system. The problem of aligning various Optical Character Recognition (OCR) results lies in the difficulties to find the correspondence on character, word, line, and page level. These difficulties arise from segmentation and recognition errors which are produced by the OCRs. Therefore, alignment techniques are required for synchronizing the outputs in order to compare them. Most existing approaches fail when the same error occurs in the multiple OCRs. If the corrections do not appear in one of the OCR approaches are unable to improve the results. We design a Line-to-Page alignment with edit rules using Weighted Finite-State Transducers (WFST). These edit rules are based on edit operations: insertion, deletion, and substitution. Therefore, an approach is designed using Recurrent Neural Networks with Long Short-Term Memory (LSTM) to predict these types of errors. A Character-Epsilon alignment is designed to normalize the size of the strings for the LSTM alignment. The LSTM returns best voting, especially when the heuristic approaches are unable to vote among various OCR engines. LSTM predicts the correct characters, even if the OCR could not produce the characters in the outputs. The approaches are evaluated on OCR's output from the UWIII and historical German Fraktur dataset which are obtained from state-of-the-art OCR systems. The experiments shows that the error rate of the LSTM approach has the best performance with around 0.40%, while other approaches are between 1.26% and 2.31%. Mayce Ibrahim Ali Al Azawi, Marcus Liwicki, Thomas M. Breuel |
ICDAR | 2 |
| 2015 | Page segmentation of historical document images with convolutional autoencodersabstractIn this paper, we present an unsupervised feature learning method for page segmentation of historical handwritten documents available as color images. We consider page segmentation as a pixel labeling problem, i.e., each pixel is classified as either periphery, background, text block, or decoration. Traditional methods in this area rely on carefully hand-crafted features or large amounts of prior knowledge. In contrast, we apply convolutional autoencoders to learn features directly from pixel intensity values. Then, using these features to train an SVM, we achieve high quality segmentation without any assumption of specific topologies and shapes. Experiments on three public datasets demonstrate the effectiveness and superiority of the proposed approach. Kai Chen 0011, Mathias Seuret, Marcus Liwicki, Jean Hennebert, Rolf Ingold |
ICDAR | 3 |
| 2015 | ICDAR2015 competition on signature verification and writer identification for on- and off-line skilled forgeries (SigWIcomp2015)abstractThis paper presents the results of the ICDAR 2015 competition on signature verification and writer identification for on- and off-line skilled forgeries jointly organized by PR-researchers and Forensic Handwriting Examiners (FHEs). The aim is to bridge the gap between recent technological developments and forensic casework. Two modalities (signatures and handwritten text) are considered and training and evaluation data are collected and provided by FHEs and PR-researchers. Four tasks are defined for four different languages; Bengali off-line signature verification, Italian off-line signature verification, German on-line signature verification, and English handwritten text based writer identification. In total, 40 systems have participated in this competition. The participants of the signatures modality were motivated to report their results in Likelihood Ratios (LRs). This has made the systems even more interesting for application in forensic casework. For evaluating the performance of the systems, we have used the forensically substantial Cost of Log Likelihood Ratios (Ĉllr) in the case of signatures, and the F-measure in the case of handwritten text. Muhammad Imran Malik, Sheraz Ahmed, Angelo Marcelli, Umapada Pal 0001, Michael Blumenstein, Linda Alewijnse, Marcus Liwicki |
ICDAR | 7 |
| 2015 | Sparse radial sampling LBP for writer identificationabstractSampling Local Binary Patterns, a variant of Local Binary Patterns (LBP) for text-as-texture classification. By adapting and extending the standard LBP operator to the particularities of text we get a generic text-as-texture classification scheme and apply it to writer identification. In experiments on CVL and ICDAR 2013 datasets, the proposed feature-set and a simple end-to-end pipeline demonstrate State-Of-the-Art (SOA) performance. Among the SOA, the proposed method is the only one that is based on dense extraction of a single local feature descriptor. This makes it fast and applicable at the earliest stages in a DIA pipeline without the need for segmentation, binarization, or extraction of multiple features. Anguelos Nicolaou, Andrew D. Bagdanov, Marcus Liwicki, Dimosthenis Karatzas |
ICDAR | 3 |
| 2015 | Parallel sequence classification using recurrent neural networks and alignmentabstractThe aim of this work is to investigate Long Short-Term Memory (LSTM) for finding the semantic associations between two parallel text lines of different instances of the same class sequence. In this work, we propose a new model called class-less classifier, which is cognitive motivated by a simplified version of the infants learning. The presented model not only learns the semantic association but also learns the relation between the labels and the classes. In addition, our model uses two parallel class-less LSTM networks and the learning rule is based on the alignment of both networks. For testing purposes, a parallel sequence dataset is generated based on MNIST dataset, which is a standard dataset for handwritten digit recognition. The results of our model were similar to the standard LSTM. Federico Raue, Wonmin Byeon, Thomas M. Breuel, Marcus Liwicki |
ICDAR | 4 |
| 2015 | Gradient-domain degradations for improving historical documents images layout analysisabstractWe present a novel method for adding realistic degradations to historical document images in order to generate more training data. Degradation patches are extracted from other documents and applied to the target document in the gradient domain. Working in the gradient domain has not been done for this purpose in document images analysis so far. It has the advantage to prevent color inconsistencies and allows to efficiently avoid border effects. This paper contains the detailed description of our novel method, with a focus on the mathematical aspect of the transition to and from the gradient domain. Furthermore, we perform quantitative experiments where we investigate the effects of using synthetically generated training data on historical documents with different kind of degradations. Mathias Seuret, Kai Chen 0011, Nicole Eichenberger, Marcus Liwicki, Rolf Ingold |
ICDAR | 4 |
| 2015 | Recognition of historical Greek polytonic scripts using LSTM networksabstractThis paper reports on high-performance Optical Character Recognition (OCR) experiments using Long Short-Term Memory (LSTM) Networks for Greek polytonic script. Even though there are many Greek polytonic manuscripts, the digitization of such documents has not been widely applied, and very limited work has been done on the recognition of such scripts. We have collected a large number of diverse document pages of Greek polytonic scripts in a novel database, called Polyton-DB, containing 15; 689 textlines of synthetic and authentic printed scripts and performed baseline experiments using LSTM Networks. Evaluation results show that the character error rate obtained with LSTM varies from 5.51% to 14.68% (depending on the document) and is better than two well-known OCR engines, namely, Tesseract and ABBYY FineReader. Foteini Liwicki, Adnan Ul-Hasan, Vassilis Papavassiliou, Basilios Gatos, Vassilis Katsouros, Marcus Liwicki |
ICDAR | 6 |
| 2015 | A sequence learning approach for multiple script identificationabstractIn this paper, we present a novel methodology for multiple script identification using Long Short-Term Memory (LSTM) networks' sequence-learning capabilities. Our method is able to identify multiple scripts at text-line level, where two or more scripts are present in the same text-line. Unlike traditional techniques, where either shape features or bounding boxes of individual characters are extracted, the LSTM-based system learns a particular script in a supervised learning framework. Moreover, this system neither needs specific features nor other preprocessing steps other than text-line extraction and text-line normalization. The proposed method works on text-line level, where it identifies each character as belonging to a particular script. We have developed a database consisting of English and Greek script, and our system achieved a script recognition accuracy of 98.186% on this dataset. Adnan Ul-Hasan, Muhammad Zeshan Afzal, Faisal Shafait, Marcus Liwicki, Thomas M. Breuel |
ICDAR | 4 |
| 2015 | Curriculum learning for printed text line recognition of ligature-based scriptsabstractThis paper introduces a novel curriculum learning strategy for ligature-based scripts. Long Short-Term Memory Networks require thousands or even millions of iterations on target symbols, depending upon the complexity of the target data, to converge when trained for sequence transcription because they have to localize the individual symbols along with the recognition. Curriculum learning reduces the number of target symbols to be visited before the network converges. In this paper, we propose a ligature-based complexity measure to define the sampling order of the training data. Experiments performed on UPTI database show that the curriculum learning using our strategy can reduce the total number of target symbols before convergence for printed Urdu Nastaleeq OCR task. Adnan Ul-Hasan, Faisal Shafait, Marcus Liwicki |
ICDAR | 3 |
| 2015 | An analysis of Dynamic Cortex Memory networksabstractThe recently introduced Dynamic Cortex Memory (DCM) is an extension of the Long Short Term Memory (LSTM) providing a systematic inter-gate connection infrastructure. In this paper the behavior of DCM networks is studied in more detail and their potential in the field of gradient-based sequence learning is investigated. Hereby, DCM networks are analyzed regarding particular key features of neural signal processing systems, namely, their robustness to noise and their ability of time warping. Throughout all experiments we show that DCMs converge faster and yield better results than LSTMs. Hereby, DCM networks require overall less weights than pure LSTM networks to achieve the same or even better results. Besides, a promising neurally implemented just-in-time online signal filter approach is presented, which is latency-free and still provides an accurate filtering performance much better than conventional low-pass filters. We also show that the neural networks can do explicit time warping even better than the Dynamic Time Warping (DTW) algorithm, which is a specialized method developed for this task. Sebastian Otte, Andreas Zell, Marcus Liwicki |
IJCNN | 3 |
| 2015 | Parallel Multi-Dimensional LSTM, With Application to Fast Biomedical Volumetric Image SegmentationabstractConvolutional Neural Networks (CNNs) can be shifted across 2D images or 3D videos to segment them. They have a fixed input size and typically perceive only small local contexts of the pixels to be classified as foreground or background. In contrast, Multi-Dimensional Recurrent NNs (MD-RNNs) can perceive the entire spatio-temporal context of each pixel in a few sweeps through all pixels, especially when the RNN is a Long Short-Term Memory (LSTM). Despite these theoretical advantages, however, unlike CNNs, previous MD-LSTM variants were hard to parallelise on GPUs. Here we re-arrange the traditional cuboid order of computations in MD-LSTM in pyramidal fashion. The resulting PyraMiD-LSTM is easy to parallelise, especially for 3D data such as stacks of brain slice images. PyraMiD-LSTM achieved best known pixel-wise brain image segmentation results on MRBrainS13 (and competitive results on EM-ISBI12). Marijn F. Stollenga, Wonmin Byeon, Marcus Liwicki, Jürgen Schmidhuber |
NIPS | 3 |
| 2015 | Scene analysis by mid-level attribute learning using 2D LSTM networks and an application to web-image tagging
Wonmin Byeon, Marcus Liwicki, Thomas M. Breuel |
Pattern Recognit. Lett. | 2 |
| 2014 | A Combined System for Text Line Extraction and Handwriting Recognition in Historical DocumentsabstractAutomated reading of historical handwriting is needed to search and browse ancient manuscripts in digital libraries based on their textual content. In this paper, we present a combined system for text localization and transcription in page images. It includes flexible learning-based methods for layout analysis and handwriting recognition, which were developed in the context of the Swiss research project HisDoc. A comprehensive experimental evaluation is provided for the medieval Parzival database, demonstrating a promising word recognition accuracy of 93.0% with closed vocabulary. In order to harmonize the evaluation of the two document analysis tasks, we introduce a novel evaluation measure for text line extraction that takes substitution, deletion, as well as insertion errors into account. Andreas Fischer 0002, Micheal Baechler, Angelika Garz, Marcus Liwicki, Rolf Ingold |
Document Analysis Systems | 4 |
| 2014 | Local Binary Patterns for Arabic Optical Font RecognitionabstractOptical Font Recognition (OFR) has been proven to increase Optical Character Recognition (OCR) accuracy, but it can also help in harvesting semantic information from documents. It therefore becomes a part of many Document Image Analysis (DIA) pipelines. Our work is based on the hypothesis that Local Binary Patterns (LBP), as a generic texture classification method, can address several distinct DIA problems at the same time such as OFR, script detection, writer identification, etc. In this paper we strip down the Redundant Oriented LBP (RO-LBP) method, previously used in writer identification, and apply it for OFR with the goal of introducing a generic method that classifies text as oriented texture. We focus on Arabic OFR and try to perform a thorough comparison of our method and the leading Gaussian Mixture Model method that is developed specifically for the task. Depending on the nature of proposed OFR method, each method's performance is usually evaluated on different data and with different evaluation protocols. The proposed experimental procedure addresses this problem and allows us to compare OFR methods that are fundamentally different by adapting them to a common measurement protocol. In performed experiments LBP method achieves perfect results on large text blocks generated from the APTI database, while preserving its very broad generic attributes as proven by secondary experiments. Anguelos Nicolaou, Fouad Slimane, Volker Märgner, Marcus Liwicki |
Document Analysis Systems | 4 |
| 2014 | Dynamic Cortex Memory: Enhancing Recurrent Neural Networks for Gradient-Based Sequence Learning
Sebastian Otte, Marcus Liwicki, Andreas Zell |
ICANN | 2 |
| 2014 | Page Segmentation for Historical Handwritten Document Images Using Color and Texture FeaturesabstractIn this paper we present a physical structure detection method for historical handwritten document images. We considered layout analysis as a pixel labeling problem. By classifying each pixel as either periphery, background, text block, or decoration, we achieve high quality segmentation without any assumption of specific topologies and shapes. Various color and texture features such as color variance, smoothness, Laplacian, Local Binary Patterns, and Gabor Dominant Orientation Histogram are used for classification. Some of these features have so far not got many attentions for document image layout analysis. By applying an Improved Fast Correlation-Based Filter feature selection algorithm, the redundant and irrelevant features are removed. Finally, the segmentation results are refined by a smoothing post-processing procedure. The proposed method is demonstrated by experiments conducted on three different historical handwritten document image datasets. Experiments show the benefit of combining various color and texture features for classification. The results also show the advantage of using a feature selection method to choose optimal feature subset. By applying the proposed method we achieve superior accuracy compared with earlier work on several datasets, e.g., We achieved 93% accuracy compared with 91% of the previous method on the Parzival dataset which contains about 100 million pixels. Kai Chen 0011, Hao Wei 0001, Jean Hennebert, Rolf Ingold, Marcus Liwicki |
ICFHR | 5 |
| 2014 | Online Signature Verification Based on Kolmogorov-Smirnov Distribution DistanceabstractOnline signature verification methods examine the dynamics of the handwriting process to decide whether a signature is probably genuine or forged. Most of the previously proposed methods for online signature verification apply Neural Networks, Dynamic Time Warping, or Hidden Markov Model for classification and they consider several aspects, like planar coordinates, pressure, velocity, and acceleration with respect to time. Here we apply a non-parametric statistical test for a comparison of features and the verification of signatures. Erika Griechisch, Muhammad Imran Malik, Marcus Liwicki |
ICFHR | 3 |
| 2014 | Automatic Signature Stability Analysis and Verification Using Local FeaturesabstractThe purpose of writing this paper is two-fold. First, it presents a novel signature stability analysis based on signature's local / part-based features. The Speeded Up Local features (SURF) are used for local analysis which give various clues about the potential areas from whom the features should be exclusively considered while performing signature verification. Second, based on the results of the local stability analysis we present a novel signature verification system and evaluate this system on the publicly available dataset of forensic signature verification competition, 4NSigComp2010, which contains genuine, forged, and disguised signatures. The proposed system achieved an EER of 15%, which is considerably very low when compared against all the participants of the said competition. Furthermore, we also compare the proposed system with some of the earlier reported systems on the said data. The proposed system also outperforms these systems. Muhammad Imran Malik, Marcus Liwicki, Andreas Dengel 0001, Seiichi Uchida, Volkmar Frinken |
ICFHR | 2 |
| 2014 | Pixel Level Handwritten and Printed Content Discrimination in Scanned DocumentsabstractClassification of the content of a scanned document as either printed or handwritten is typically tackled as a segmentation problem of pages into text lines or words. However these methods are not applicable on documents where handwritten annotations overlay printed text. In this paper we propose to treat the task as a pixel classification task, i.e., To classify individual foreground pixels into either printed or handwritten pixels. Our method uses various features of diverse nature taking the surrounding window into account. The influence of the features and their parameters are investigated and optimized on a validation set. Each foreground pixel is then classified by a multilayer perceptron using feature vectors based on a pixel neighborhood. Finally, a post-processing step corrects typical misclassifications, i.e., It removes outliers based on several heuristics. We evaluated our method on printed documents with real handwritten annotations and reached an accuracy of 96.10% on the test set. This is significantly higher than a previously published methods based on local features. Mathias Seuret, Marcus Liwicki, Rolf Ingold |
ICFHR | 2 |
| 2014 | Hybrid Feature Selection for Historical Document Layout AnalysisabstractIn this paper we propose a novel hybrid feature selection method for historical Document Image Analysis (DIA). Adapted greedy forward selection and genetic selection are used in a cascading way. We apply the proposed method to the task of historical document layout analysis on three handwritten datasets of diverse nature. The documents contain complex layouts, different handwriting styles, and several results of decay. The task is to segment each page into four areas: periphery, background, text block, and decoration. The proposed method selected significantly less features and resulted in significantly lower error rates than using all features. Compared to several conventional feature selection methods, the proposed method is competitive with respect to the number of selected features and the resultant error rates. In addition, we found that some features, e.g., Gradient, Laplacian, and local binary patterns (LBP), are selected by most of the feature selection methods and we give some explanations. This finding suggests a clue for the layout analysis on handwritten documents in general. Hao Wei 0001, Kai Chen 0011, Rolf Ingold, Marcus Liwicki |
ICFHR | 4 |
| 2014 | Texture Classification Using 2D LSTM NetworksabstractIn this paper, we investigate the ability of the Long short term memory (LSTM) recurrent neural network architecture to perform texture classification on images. Existing approaches to texture classification rely on manually designed preprocessing steps or selected feature extractors. Since LSTM networks are able to bridge over long time lags, we propose applying them directly on the image, circumventing any handcrafted pre-processing. We investigate different approaches with several input and output representations. In our experiments on a number of widely used texture benchmarking tasks (KTH-TIPS, OuTex, VisTexL, VisTexP, and Newmarket), we show that the performance is comparable to, or better than, existing state-of-the-art methods for texture classification. Wonmin Byeon, Marcus Liwicki, Thomas M. Breuel |
ICPR | 2 |
| 2014 | Robust Text Line Segmentation for Historical Manuscript Images Using Color and TextureabstractIn this paper we present a novel text line segmentation method for historical manuscript images. We use a pyramidal approach where at the first level, pixels are classified into: text, background, decoration, and out of page, at the second level, text regions are split into text line and non text line. Color and texture features based on Local Binary Patterns and Gabor Dominant Orientation are used for classification. By applying a modified Fast Correlation-Based Filter feature selection algorithm, redundant and irrelevant features are removed. Finally, the text line segmentation results are refined by a smoothing post-processing procedure. Unlike other projection profile or connected components methods, the proposed algorithm does not use any script-specific knowledge and is applicable to color images. The proposed algorithm is evaluated on three historical manuscript image datasets of diverse nature and achieved an average precision of 91% and recall of 84%. Experiments also show that the proposed algorithm is robust with respect to changes of the writing style, page layout, and noise on the image. Kai Chen 0011, Hao Wei 0001, Marcus Liwicki, Jean Hennebert, Rolf Ingold |
ICPR | 3 |
| 2014 | LSTM-Based Early Recognition of Motion PatternsabstractIn this paper a method for Early Recognition (ER) of Motion Templates (MTs) is presented. We define ER as an algorithm to provide recognition results before a motion sequence is completed. In our experiments we apply Long Short-Term Memory (LSTM) and optimize the training for the task of recognizing the motion template as early as possible. The evaluation has shown that the recognition accuracy for a frame-by-frame classification the LSTM achieves a recognition accuracy of 88% if no training data of the person him/herself is included, and 92% if the training data also contains motion sequences of the person. Furthermore, the average earliness - the number of time frames it takes before the LSTM correctly classifies a motion pattern - is around 24.77 frames, which is less than a second with the used tracking technology, i.e., the Microsoft Kinect. Marcus Liwicki, Didier Stricker, Christopher Schölzel, Seiichi Uchida |
ICPR | 2 |
| 2014 | Statistical segmentation and structural recognition for floor plan interpretation - Notation invariant structural element recognition
Lluís-Pere de las Heras, Sheraz Ahmed, Marcus Liwicki, Ernest Valveny, Gemma Sánchez |
Int. J. Document Anal. Recognit. | 3 |
| 2014 | Automatic analysis and sketch-based retrieval of architectural floor plans
Sheraz Ahmed, Marcus Liwicki, Christoph Langenhan, Andreas Dengel 0001, Frank Petzold |
Pattern Recognit. Lett. | 3 |
| 2014 | Bridging the gap between handwriting recognition and knowledge management
Marcus Liwicki, Sebastian Ebert, Andreas Dengel 0001 |
Pattern Recognit. Lett. | 1 |
| 2014 | More than ink - Realization of a data-embedding pen
Marcus Liwicki, Seiichi Uchida, Akira Yoshida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Pattern Recognit. Lett. | 1 |
| 2013 | Automatic Ground Truth Generation of Camera Captured Documents Using Document Image RetrievalabstractIn this paper a novel method for automatic ground truth generation of camera captured document images is proposed. Currently, no dataset is available for camera captured documents. It is very difficult to build these datasets manually, as it is very laborious and costly. The proposed method is fully automatic, allowing building the very large scale (i.e., millions of images) labeled camera captured documents dataset, without any human intervention. Evaluation of samples generated by the proposed approach shows that 99.98% of the images are correctly labeled. Novelty of the proposed approach lies in the use of document image retrieval for automatic labeling, especially for camera captured documents, which contain different distortions specific to camera, e.g., blur, occlusion, perspective distortion, etc. Sheraz Ahmed, Koichi Kise, Masakazu Iwamura, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 4 |
| 2013 | A Generic Method for Stamp Segmentation Using Part-Based FeaturesabstractTraditionally, stamps are considered as a seal of authenticity for documents. For automatic processing and verification, segmentation of stamps from documents is pivotal. Existing methods for stamp extraction mostly employ color and/or shape based techniques, thereby limiting their applicability to only colored and specific shape stamps. In this paper, a novel, generic method based on part-based features is presented for segmentation of stamps from document images. The proposed method can segment black, colored, unseen, arbitrary shaped, textual, as well as graphical stamps. The proposed method is evaluated on a publicly available dataset for stamp detection and verification and achieved recall and precision of 73% and 83% respectively, for black stamps which were not addressed in the past. Sheraz Ahmed, Faisal Shafait, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 3 |
| 2013 | Text Line Extraction Using DMLP Classifiers for Historical ManuscriptsabstractThis paper proposes a novel text line extraction method for historical documents. The method works in two steps. In the first step, layout analysis is performed to recognize the physical structure of a given document using a classification technique, more precisely the pixels of a coloured document image are classified into five classes: text-block, core-text-line, decoration, background, and periphery. This layout recognition is achieved by a cascade of two Dynamic Multilayer Perceptron (DMLP) classifiers and works without binarisation. In the second step, an algorithm takes the layout recognition results as an input, extracts the text lines, and groups them into blocks using the connected components approach. Finally, the algorithm refines the boundaries of the text lines using the binary image and the layout recognition results. Our system is evaluated on three historical manuscripts with a test set of 49 pages. The best obtained hit rate for text lines is 96.3%. Micheal Baechler, Marcus Liwicki, Rolf Ingold |
ICDAR | 2 |
| 2013 | Online Signature Analysis Based on Accelerometric and Gyroscopic Pens and Legendre SeriesabstractIn this paper we compare two captured databases which contain local acceleration and angle information recorded during the signing process. Approximately a year passed between the capturing of the two databases and they contain several signatures from the same writers. We analyze the expedience of the proposed devices and examine the overlap of the databases using Legendre approximation for feature computation and Support Vector Machine for classification. In addition we plan to make the concerned databases publicly available for research purposes. Erika Griechisch, Muhammad Imran Malik, Marcus Liwicki |
ICDAR | 3 |
| 2013 | FREAK for Real Time Forensic Signature VerificationabstractThis paper presents a novel signature verification system based on local features of signatures. The proposed system uses Fast Retina Key points (FREAK) which represent local features and are inspired by the human visual system, particularly the retina. To locate local points of interest in signatures, two local key point detectors, i.e., Features from Accelerated Segment Test (FAST) and Speeded-up Robust Features (SURF), have been used and their performance comparison in terms of Equal Error Rate (EER) and time is presented. The proposed system has been evaluated on publicly available dataset of forensic signature verification competition, 4NSigComp2010, which contains genuine, forged, and disguised signatures. The proposed system achieved an EER of 30%, which is considerably very low when compared against all the participants of the said competition. In addition to EER, the proposed system requires only 0.6 seconds on average to verify a 3000*1500 scanned signature. This shows that the proposed system has a potential and suitability for forensic signature verification as well as real time applications. Muhammad Imran Malik, Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 3 |
| 2013 | ICDAR 2013 Competitions on Signature Verification and Writer Identification for On- and Offline Skilled Forgeries (SigWiComp 2013)abstractThis paper presents the results of the ICDAR2013 competitions on signature verification and writer identification for on- and offline skilled forgeries jointly organized by PR researchers and Forensic Handwriting Examiners (FHEs). The aim is to bridge the gap between recent technological developments and forensic casework. Two modalities (signatures, and handwritten text) are considered where training and evaluation data (in Dutch and Japanese) were collected and provided by FHEs and PR-researchers. Four tasks were defined where the systems had to perform Dutch offline signature verification, Japanese offline signature verification, Japanese online signature verification, and Dutch writer identification. The participants of the signatures modality were motivated to report their results in Likelihood Ratios (LR). This has made the systems even more interesting for application in forensic casework. For evaluation of signatures modality, we used both the traditional Equal Error Rate (EER) and forensically substantial Cost of Log Likelihood Ratios (Ĉllr). The system having the smallest value of the Minimum Cost of Log Likelihood Ratio (Ĉllrmin) is declared winner. For evaluation of the handwritten text modality, we used the precision and accuracy measures and winners are announced on the basis of best F-measure value. Muhammad Imran Malik, Marcus Liwicki, Linda Alewijnse, Wataru Ohyama, Michael Blumenstein, Bryan Found |
ICDAR | 2 |
| 2013 | Part-Based Automatic System in Comparison to Human Experts for Forensic Signature VerificationabstractThe purpose of writing this paper is three-fold. First, it presents a novel local / part-based automatic system for forensic signature verification involving disguised signatures. Disguised signatures are written by authentic authors but with the intention of later denial. The proposed system reaches an equal error rate of 3.36% in classifying disguised and genuine signatures. Second, it compares the performance of the proposed system with various state-of-the-art signature verification systems on the same data, i.e., the publicly available dataset of 4NSigComp2010 signature verification competition. Third, it presents a performance comparison of the proposed system with human forensic handwriting examiners. It is important as it highlights the potential of the proposed system to assist humans in solving real world forensic signature verification cases. Muhammad Imran Malik, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 2 |
| 2013 | Continuous Partial-Order Planning for Multichannel Document Analysis: A Process-Driven ApproachabstractWith the rise of email communication, enterprises strive to manage incoming documents from all input channels for achieving customer satisfaction. Their overall goal is to reduce request processing time and to increase processing quality. Previously, we proposed the approach of process-driven document analysis (DA) using the concepts of Attentive Tasks (ATs) and the Specialist Board (SB). The ATs formalize information expectations of the processes toward an incoming document, whereas the SB describes all available DA methods. In this paper, we propose to apply continuous partial order planning (CPOP) for guiding DA with the goal of optimal extraction accuracy and runtime. To our knowledge, this approach provides a novel method for integrating knowledge management with DA, in particular for processes. Since planning has not been applied to this field yet, we explore learning the suitability function (SF) and the adaptation of the DA plan. First evaluations indicate the applicability of the approach and preferences for calibration. Kristin Stamm, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 2 |
| 2013 | Part-Based Recognition of Arbitrary FontsabstractIn this paper, the part-based recognition method is introduced and applied to the arbitrary font recognition. The principle of the part-based method is to represent the character image as a set of parts and then recognize the image by finding the most possible parts set from the reference database. Since the part-based method does not rely on the global structure of a character, it is supposed to be robust against the variant appearances of the character. The experiment results indicate that it is possible to apply the part-based method to the font recognition, which is always considered as a difficult task by most of the researchers. Seiichi Uchida, Marcus Liwicki |
ICDAR | 3 |
| 2013 | Graph-based retrieval of building information models for supporting the early design stages
Christoph Langenhan, Marcus Liwicki, Frank Petzold, Andreas Dengel 0001 |
Adv. Eng. Informatics | 3 |
| 2013 | Part-based methods for handwritten digit recognition
Seiichi Uchida, Marcus Liwicki, Yaokai Feng |
Frontiers Comput. Sci. | 3 |
| 2012 | Extraction of Text Touching Graphics Using SURFabstractIn this paper we propose a novel part-based method for the extraction of text touching graphic components. The Speeded Up Robust Features (SURF) are used to localize the text components and distinguish them from graphics. We introduce several post-processing steps to finally detect the text. We have tested our method on a publicly available data set of architectural floor plans and on real geographical maps. On floor plans we have located more than 95% of the text components which were not identified as text beforehand because they were touching graphic components. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
Document Analysis Systems | 2 |
| 2012 | Automatic Room Detection and Room Labeling from Architectural Floor PlansabstractThis paper presents an automatic system for analyzing and labeling architectural floor plans. In order to detect the locations of the rooms, the proposed systems extracts both, structural and semantic information from given floor plans. Furthermore, OCR is applied on the text layer to retrieve the meaningful room labeling. Finally, a novel post-processing is proposed to split rooms into several sub-regions if several semantic rooms share the same physical room. Our fully automatic system is evaluated on a publicly available dataset of architectural floor plans. In our experiments, we could clearly outperform other state-of-the-art approaches for room detection. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
Document Analysis Systems | 2 |
| 2012 | Koios++: A Query-Answering System for Handwritten InputabstractIn this paper we propose KOIOS++, which automatically processes natural language queries provided by handwritten input. The system integrates several recent achievements in the area of handwriting recognition, natural language processing, information retrieval, and human computer interaction. It uses a knowledge base described by the resource description framework (RDF). Our generic approach first generates a lexicon as background information for the handwritten text recognition. After recognizing a handwritten query, several output hypotheses are sent to a natural language processing system in order to generate a structured query (SPARQL query). Subsequently, the query is applied to the given knowledge base and a result graph visualizes the retrieved information. At all stages, the user can easily adjust the intermediate results if there is any undesired outcome. The system is implemented as a web-service and therefore works for handwritten input on digital paper as well as on input on Pen-enabled interactive surfaces. Furthermore, we build on the generic RDF-representation of semantic knowledge which is also used by the linked open data (LOD) initiative. As such, our system works well in various scenarios. We have implemented prototypes for querying company knowledge bases, the DBPedia1, the DBLP computer science bibliography2, and a knowledge base of the DAS 2012. Marcus Liwicki, Björn Forcher, Philipp Jaeger, Andreas Dengel 0001 |
Document Analysis Systems | 1 |
| 2012 | Seamless Integration of Handwriting Recognition into Pen-Enabled Displays for Fast User InteractionabstractThis paper proposes a framework for the integration of handwriting recognition into natural user interfaces. As more and more pen-enabled touch displays are available, we make use of the distinction between touch actions and pen actions. Furthermore, we apply a recently introduced mode detection approach to distinguish between handwritten strokes and graphics drawn with the pen. These ideas are implemented in the Touch & Write SDK which can be used for various applications. In order to evaluate the effectiveness of our approach, we have conducted experiments for an annotation scenario. We asked several users to mark and label several objects in videos. We have measured the labeling time when using our novel user interaction system and compared it to the time needed when using common labeling tools. Furthermore, we compare our handwritten input paradigm to other existing systems. It turns out that the annotation is performed much faster when using our method and the user experience is also much better. Marcus Liwicki, Tobias Zimmermann, Andreas Dengel 0001 |
Document Analysis Systems | 1 |
| 2012 | A Signature Verification Framework for Digital Pen ApplicationsabstractIn this paper we present a framework for real-time online signature verification scenarios. The proposed framework is based on state-of-the-art feature extraction and Gaussian Mixture Model (GMM) classification. While our signature verification library is generally applicable to any input device using digital pens, we have implemented verification scenarios using the Anoto digital pen. As such our automated signature verification framework becomes an interesting commodity for industry, because the Anoto SDK is easy to apply and the GMM-based classification can be seamlessly integrated. The novelty of this work is the application of our framework that takes real-time online signature verification to every scenario where digital pens may potentially be used. In this paper we describe several scenarios where our framework has been applied, including signatures in financial contracts or ordering processes. We also propose a general approach to integrate the GMM-descriptions into electronic ID-cards in order to also store behavioral biometrics on these cards. In experiments we have measured the performance of the signature verification system when skilled forgeries were present. The interest shown by our partner financial institutions and the results of our initial evaluations indicate that our signature verification framework suits exactly the demands of our clients. Muhammad Imran Malik, Sheraz Ahmed, Andreas Dengel 0001, Marcus Liwicki |
Document Analysis Systems | 4 |
| 2012 | Toward Part-Based Document Image DecodingabstractDocument image decoding (DID) is a trial to understand the contents of a whole document without any reference information about font, language, etc. Typically, DID approaches assume the correct segmentation of the document and some a priori knowledge about the language or the script. Unfortunately, this assumption will not hold if we deal with various documents, such as documents with various sized fonts, camera-captured documents, free-layout documents, or historical documents. In this paper, we propose a part-based character identification method where no segmentation into characters is necessary and no a priori information about the document is needed. The approach clusters similar key points and groups frequent neighboring key point clusters. Then a second iteration is performed, i.e., the groups are again clustered and optionally pairs frequent group clusters are detected. Our first experimental results on multi font-size documents look already very promising. We could find nearly perfect correspondences between characters and detected group clusters. Wang Song, Seiichi Uchida, Marcus Liwicki |
Document Analysis Systems | 3 |
| 2012 | Signature Segmentation from Document ImagesabstractIn this paper we propose a novel method for the extraction of signatures from document images. Instead of using a human defined set of features a part-based feature extraction method is used. In particular, we use the Speeded Up Robust Features (SURF) to distinguish the machine printed text from signatures. Using SURF features makes the approach generally more useful and reliable for different resolution documents. We have evaluated our system on the publicly available Tobacco-800 dataset in order to compare it to previous work. Finally, all signatures were found in the images and less than half of the found signatures are false positives. Therefore, our system can be applied for practical use. Sheraz Ahmed, Muhammad Imran Malik, Marcus Liwicki, Andreas Dengel 0001 |
ICFHR | 3 |
| 2012 | ICFHR 2012 Competition on Automatic Forensic Signature Verification (4NsigComp 2012)abstractThis paper presents the results of the ICFHR2012 Competition on Automatic Forensic Signature Verification jointly organized by PR-researchers and Forensic Handwriting Examiners (FHEs). The aim is to bridge the gap between recent technological developments and forensic casework. A forensic like training set containing disguised signatures along with skilled forgeries and genuine signatures was provided to the participants. They were motivated to report the results in Likelihood Ratios (LR). This has made the systems even more interesting for application in forensic casework. For evaluation we used both the traditional Equal Error Rate (EER) and forensically substantial Cost of Log Likelihood Ratios (Ĉllr). The system having the best Minimum Cost of Log Likelihood Ratio ( Ĉllrmin) is declared winner. Various experiments both including and excluding disguised signatures from the test set are reported. Marcus Liwicki, Muhammad Imran Malik, Linda Alewijnse, C. Elisa van den Heuvel, Bryan Found |
ICFHR | 1 |
| 2012 | From Terminology to Evaluation: Performance Assessment of Automatic Signature Verification SystemsabstractThis paper is an effort towards the development of a shared conceptualization regarding automatic signature verification systems. The requirements of both communities, Pattern Recognition and Forensic Handwriting Examiners, are explicitly focused. This is required because an increasing gap regarding evaluation of automatic verification systems is observed in the recent past. The paper addresses three major areas. First, it highlights how signature verification is taken differently in the above mentioned communities and why this gap is increasing. Various factors that widen this gap are discussed with reference to some of the recent signature verification studies and probable solutions are suggested. Second, it discusses the state-of-the-art evaluation and its problems as seen by FHEs. The real evaluation issues faced by FHEs, when trying to incorporate automatic signature verification systems in their routine casework, are presented. Third, it reports a standardized evaluation scheme capable of fulfilling the requirements of both PR researchers and FHEs. Muhammad Imran Malik, Marcus Liwicki |
ICFHR | 2 |
| 2012 | Local Feature Based Online Mode Detection with Recurrent Neural NetworksabstractIn this paper we propose a novel approach for online mode detection, where the task is to classify ink traces into several categories. In contrast to previous approaches working on global features, we introduce a system completely relying on local features. For classification, standard recurrent neural networks (RNNs) and the recently introduced long short-term memory (LSTM) networks are used. Experiments are performed on the publicly available IAMonDo-database which serves as a benchmark data set for several researches. In the experiments we investigate several RNN structures and classification sub-tasks of different complexities. The final recognition rate on the complete test set is 98.47% in average, which is significantly higher than the 97% achieved with an MCS in previous work. Further interesting results on different subsets are also reported in this paper. Sebastian Otte, Dirk Krechel, Marcus Liwicki, Andreas Dengel 0001 |
ICFHR | 3 |
| 2012 | Online Signature Verification Based on Legendre Series Representation: Robustness Assessment of Different Feature CombinationsabstractIn this paper, orthogonal polynomials series are used to approximate the time functions associated to the signatures. The coefficients in these series expansions, computed resorting to least squares estimation techniques, are then used as features to model the signatures. Different combinations of several time functions (pen coordinates, incremental variation of pen coordinates and pen pressure), related to the signing process, are analyzed in this paper for two different signature styles, namely, Western signatures and Chinese signatures of a publicly available Signature Database. Two state-of-the-art classification methods, namely, Support Vector Machines and Random Forests are used in the verification experiments. The proposed online signature verification system delivers error rates comparable to results reported over the same signature datasets in a previous signature verification competition. Marianela Parodi, Juan Carlos Gómez, Marcus Liwicki |
ICFHR | 3 |
| 2012 | Part-based method on handwritten texts
Seiichi Uchida, Marcus Liwicki |
ICPR | 3 |
| 2012 | Unsupervised motion pattern learning for motion segmentation
Gabriele Bleser-Taetz, Marcus Liwicki, Didier Stricker |
ICPR | 3 |
| 2012 | Semantic E-Ink: Knowledge-Based Assistance for Making Mental Models ExplicitabstractIn this paper we describe a system which assists knowledge workers in making notes of their thoughts and transferring them to the computer. Our system processes handwritten notes written down with a digital pen. These notes are processed in order to recognize and understand their meaning. To realize such systems, two novel processing stages are proposed for the first time in literature. The first stage is the inclusion of knowledge bases into the Handwriting Recognition (HWR) process, where we make use of a person's mental model. The second stage is the transition from pure HWR to understanding of the handwritten notes, i.e. the system extracts knowledge in form of ontologies. For both novel approaches we performed a set of experiments on various data. With the proposed techniques, the recognition rate of the HWR system as well as the performance of the information extraction system are significantly increased. Andreas Dengel 0001, Marcus Liwicki |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | Faster subgraph isomorphism detection by well-founded total order indexing
Marcus Liwicki, Andreas Dengel 0001 |
Pattern Recognit. Lett. | 2 |
| 2011 | Fast Subgraph Isomorphism Detection for Graph-Based Retrieval
Christoph Langenhan, Thomas Roth-Berghofer, Marcus Liwicki, Andreas Dengel 0001, Frank Petzold |
ICCBR | 4 |
| 2011 | Improved Automatic Analysis of Architectural Floor PlansabstractThis paper proposes a novel complete system for automated floor plan analysis. Besides applying and improving state-of-the-art processing methods, we introduce novel preprocessing methods, e.g., the differentiation between thick, medium, and thin lines and the removal of components outside the convex hull of the outer walls. Especially the latter method increases the performance of the final system. In our experiments on a reference data set we compare our approach to other approaches available in the literature. We show that our system outperforms previous systems. The final room recognition accuracy is 79% that is 10% higher than the 69% achieved by a state-of-the-art approach from the literature. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 2 |
| 2011 | Text/Graphics Segmentation in Architectural Floor PlansabstractIn this paper, we propose an improved method for text/graphics segmentation. Text/graphics separation is a crucial preprocessing step in document analysis before further analysis and recognition can be applied. Our proposed system extends the method of Tombre et al. with a number of improvements to make it more suitable for architectural floor plans. A crucial novel preprocessing step is the detection and removal of walls before the actual segmentation. Furthermore, text components are then extracted by analyzing connected components and even considering text overlapping with graphics. Finally, a smearing approach is used to remove noise and extract the final text components. Evaluation results over the series of 90 floor plans which has also been used in reference work shows that our method has a recall of almost 99% and a precision greater then 97%. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 3 |
| 2011 | Reliable Online Stroke Recovery from Offline Data with the Data-Embedding PenabstractIn this paper we propose a complete system for online stroke recovery from offline data. The key idea of our approach is to use a novel pen device which is able to embed meta information into the ink during writing the strokes. This pen-device overcomes the need to get access to any memory on the pen when trying to recover the information, which is especially useful in multi-writer or multi-pen scenarios. The actual data-embedding is achieved by an additional ink dot sequence along a handwritten pattern during writing. We design the ink-dot sequence in such a way that it is possible to retrieve the writing direction from a scanned image. Furthermore, we propose novel processing steps in order to retrieve the original writing direction and finally the embedded data. In our experiments we show that we can reliably recover the writing direction of various patterns. Our system is able to determine the writing direction of straight lines, simple patterns with crossings (e.g., "x" and "II"), and even more complex patterns like handwritten words and symbols. Marcus Liwicki, Akira Yoshida, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICDAR | 1 |
| 2011 | Signature Verification Competition for Online and Offline Skilled Forgeries (SigComp2011)abstractThe Netherlands Forensic Institute and the Institute for Forensic Science in Shanghai are in search of a signature verification system that can be implemented in forensic casework and research to objectify results. We want to bridge the gap between recent technological developments and forensic casework. In collaboration with the German Research Center for Artificial Intelligence we have organized a signature verification competition on datasets with two scripts (Dutch and Chinese) in which we asked to compare questioned signatures against a set of reference signatures. We have received 12 systems from 5 institutes and performed experiments on online and offline Dutch and Chinese signatures. For evaluation, we applied methods used by Forensic Handwriting Examiners (FHEs) to assess the value of the evidence, i.e., we took the likelihood ratios more into account than in previous competitions. The data set was quite challenging and the results are very interesting. Marcus Liwicki, Muhammad Imran Malik, C. Elisa van den Heuvel, Xiaohong Chen 0001, Charles Berger 0002, Reinoud Stoel, Michael Blumenstein, Bryan Found |
ICDAR | 1 |
| 2011 | Look Inside the World of Parts of Handwritten CharactersabstractPart-based recognition is expected to be robust in difficult handwritten character recognition tasks. This is because part-based recognition is based on aggregation of independent recognition results at individual local parts without considering their global relations and thus is robust against various deformations, such as partial occlusion, overlap, broken stroke, etc. Since part-based recognition is a new approach, there are still several open problems toward its practical use. For example, compared with entire images, local parts are more ambiguous, i.e., less discriminative. For better recognition accuracy and less computations, we need to know the characteristics of local parts and then, for example, discard less discriminative parts. The purpose of this paper is to conduct some experiments in order to observe and analyze how the local parts of multiple classes are distributed in feature spaces. By handling parts appropriately based on the analysis, we will be able to enhance the usefulness of the part-based method. Wang Song, Seiichi Uchida, Marcus Liwicki |
ICDAR | 3 |
| 2011 | Comparative Study of Part-Based Handwritten Character Recognition MethodsabstractThe purpose of this paper is to introduce three part-based methods for handwritten character recognition and then compare their performances experimentally. All of those methods decompose handwritten characters into "parts". Then some recognition processes are done in a part-wise manner and, finally, the recognition results at all the parts are combined via voting to have the recognition result of the entire character. Since part-based methods do not rely on the global structure of the character, we can expect their robustness against various deformations. Three voting methods have been investigated for the combination: single voting, multiple voting, and class distance. All of them use different strategies for voting. Experimental results on the MNIST database showed the relative superiority of the class distance method and the robustness of the multiple voting method against the reduction of training set. Wang Song, Seiichi Uchida, Marcus Liwicki |
ICDAR | 3 |
| 2011 | MCS for Online Mode Detection: Evaluation on Pen-Enabled Multi-touch InterfacesabstractThis paper proposes a new approach for drawing mode detection in online handwriting. The system classifies groups of ink traces into several categories. The main contributions of this work are as follows. First, we improve and optimize several state-of-the-art recognizers by adding new features and applying feature selections. Second, we use several classifiers for the recognition. Third, we perform multiple classifier combination strategies for combining the outputs. Finally, a large experimental evaluation on two data sets is performed: the publicly available Touch&Write database which has been acquired on a pen-enabled multi-touch surface, and the publicly available IAMonDo-database which serves as a benchmark. In our experiments on the IAM-OnDo-database we achieved a recognition rate of 97%, which is much higher than other results reported in the literature. On the more balanced multi-touch surface data set we achieved a recognition rate of close to 98%. Marcus Liwicki, Yannik T. H. Schelske, Christopher Schölzel, Florian Strauß, Andreas Dengel 0001 |
ICDAR | 2 |
| 2011 | Digital pen in mammography patient formsabstractWe present a digital pen based interface for clinical radiology reports in the field of mammography. It is of utmost importance in future radiology practices that the radiology reports be uniform, comprehensive, and easily managed. This means that reports must be "readable" to humans and machines alike. In order to improve reporting practices in mammography, we allow the radiologist to write structured reports with a special pen on paper with an invisible dot pattern. A handwriting software takes care of the interpretation of the written report which is transferred into an ontological representation. In addition, a gesture recogniser allows radiologists to encircle predefined annotation suggestions which turns out to be the most beneficial feature. The radiologist can (1) provide the image and image region annotations mapped to a FMA, RadLex, or ICD10 code, (2) provide free text entries, and (3) correct/select annotations while using multiple gestures on the forms and sketch regions. The resulting, automatically generated PDF report is then stored in a semantic backend system for further use and contains all transcribed annotations as well as all free form sketches. Daniel Sonntag, Marcus Liwicki |
ICMI | 2 |
| 2011 | Interactive paper for radiology findingsabstractThis paper presents a pen-based interface for clinical radiologists. It is of utmost importance in future radiology practices that the radiology reports be uniform, comprehensive, and easily managed. This means that reports must be "readable" to humans and machines alike. In order to improve reporting practices, we allow the radiologist to write structured reports with a special pen on normal paper. A handwriting recognition and interpretation software takes care of the interpretation of the written report which is transferred into an ontological representation. The resulting report is then stored in a semantic backend system for further use. We will focus on the pen-based interface and new interaction possibilities with gestures in this scenario. Daniel Sonntag, Marcus Liwicki |
IUI | 2 |
| 2011 | From Handwriting Recognition to Ontologie-Based Information Extraction of Handwritten Notes
Marcus Liwicki, Sebastian Ebert, Andreas Dengel 0001 |
KES (4) | 1 |
| 2011 | An Intelligent Shopping List - Combining Digital Paper with Product Ontologies
Marcus Liwicki, Sandra Thieme, Gerrit Kahl, Andreas Dengel 0001 |
KES (4) | 1 |
| 2011 | Handwriting on Paper as a Cybermedium
Akira Yoshida, Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
KES (4) | 2 |
| 2011 | Combining diverse systems for handwritten text line recognition
Marcus Liwicki, Horst Bunke, James A. Pittman, Stefan Knerr |
Mach. Vis. Appl. | 1 |
| 2011 | Automatic gender detection using on-line and off-line information
Marcus Liwicki, Andreas Schlapbach, Horst Bunke |
Pattern Anal. Appl. | 1 |
| 2010 | IAMonDo-database: an online handwritten document database with non-uniform contentsabstractIn this paper we present a new database of online handwritten documents with different contents such as text, drawings, diagrams, formulas, tables, lists, and markings. It was designed to serve as a standard dataset for the development, training, testing and comparison of methods in the field of handwritten document analysis. The database can serve as a basis for layout analysis, and different segmentation and recognition tasks considering online or just offline information. Its size is 1,000 documents produced by approximately 200 writers including a total of 329,849 online strokes. Few constraints were imposed on the writers when creating the documents. Nonetheless, the database has a stable distribution of the different content types. A software tool was developed to allow easy access to the documents which are stored in InkML. In this paper we also present two experiments which show the challenge this database poses. They may figure as references for further research in this area. Emanuel Indermühle, Marcus Liwicki, Horst Bunke |
Document Analysis Systems | 2 |
| 2010 | Improving handwriting recognition by the use of semantic informationabstractThis paper proposes a first attempt to include real semantic information into the process of handwriting recognition. We take advantage of the fact that the main topic of handwritten notes is often known beforehand like in annotation or reviewing tasks. Using state-of-the-art technologies from the knowledge management research area it is possible to store a semantic representation of the user's knowledge in a Personal Information Model (PIMO). This PIMO stores the relations between semantic concepts and documents on the computer. In this paper we extract texts from related documents and concepts of the PIMO. The vocabulary of these texts is then used to aid the recognizer. In our multi-writer experiments, a significant improvement of the recognition accuracy by 8% on the text line level has been achieved. Marcus Liwicki, Hassan Mohamed Abou Eisha, Andreas Dengel 0001 |
Document Analysis Systems | 1 |
| 2010 | Touch & Write: a multi-touch table with pen-inputabstractIn this paper we present a novel rear-projection tabletop called Touch & Write. It combines the FTIR technology for touching with the Anoto-technology for handwriting. This allows an implicit switch between the modes object manipulation, and content editing. Our system incorporates real-time gesture and handwriting recognition. Drawn objects and written concepts can be converted to digital information immediately. We introduce a functional application, the LeCoOnt concept mapping software makes use of the full capability of the Touch & Write table. Touching actions are used for arranging the concepts like sheets on a normal table, and to recognizes guestures like zooming. Pen-actions are used for drawing, connecting concepts, and handwriting. The handwritten strokes are automatically recognized and converted into a machine-readable string. This system provides a reliable alternative to common approaches which try to reconstruct the information from photographs. Marcus Liwicki, Oleg Rostanin, Saher Mohamed El-Neklawy, Andreas Dengel 0001 |
Document Analysis Systems | 1 |
| 2010 | Data-embedding pen: augmenting ink strokes with meta-informationabstractIn this paper we present the first operational version of the data-embedding pen. During writing a pattern, this pen produces an additional ink-dot sequence along the ink stroke of the pattern. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. Since the information is placed on the paper, it can be extracted just by scanning or photographing the paper. There is no need to get access to any memory on the pen to recover the information. This is useful especially in multi-writer or multi-pen scenarios. The experiments using an encoding scheme and a decoding algorithm showed very promising results. For example, it was proved that we can embed 28 or more bits of information on simple handwritten patterns and decode them with a high reliability. Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
Document Analysis Systems | 1 |
| 2010 | a.SCatch: Semantic Structure for Architectural Floor Plan Retrieval
Christoph Langenhan, Thomas Roth-Berghofer, Marcus Liwicki, Andreas Dengel 0001, Frank Petzold |
ICCBR | 4 |
| 2010 | Ontology-Based Information Extraction from Handwritten DocumentsabstractIn this paper we introduce a new layer for the task of handwriting recognition. We add semantic information by means of ontologies. The task of our recognizer therefore is not only to recognize the ASCII transcription of the handwritten document, but also to identify the semantic concepts which appear in the text. This task is called ontology-based information extraction (OBIE), which has been applied to electronic documents recently. OBIE methods first segment the text into tokens, then identify their values and their corresponding instances of the ontology, and finally try to generate new facts based on the text. To the authors' knowledge, in this paper OBIE is proposed for the first time in handwriting literature. In our experiments we have evaluated the process up to the instantiation. We have found that using not only the top alternative, but also the k-best alternatives increases the performance of information extraction. Furthermore, the use of an ontology-based lexicon results in another performance increase. Sebastian Ebert, Marcus Liwicki, Andreas Dengel 0001 |
ICFHR | 2 |
| 2010 | Forensic Signature Verification Competition 4NSigComp2010 - Detection of Simulated and Disguised SignaturesabstractThis competition scenario aims at a performance comparison of several automated systems for the task of signature verification. The systems have to rate the probability of authorship and non-authorship of signatures. In particular they have to determine whether questioned signatures are simulated disguised or the normal signature of the reference writer. Furthermore, the results will be compared to forensic handwriting examiners (FHEs) opinions on the same tasks. As such, to the best of the authors’ knowledge, this scenario will be the first attempt in literature to relate system performances to the performance of FHEs who gave their opinion on exactly the the same signatures. Marcus Liwicki, C. Elisa van den Heuvel, Bryan Found, Muhammad Imran Malik |
ICFHR | 1 |
| 2010 | Embedding Meta-Information in Handwriting -- Reed-Solomon for Reliable Error CorrectionabstractIn this paper a more compact and more reliable coding scheme for the data-embedding pen is proposed. The data-embedding pen produces an additional ink-dot sequence along a handwritten pattern during writing. The ink-dot sequence represents, for example, meta-information (such as the writer's name and the date of writing) and thus drastically increases the value of the handwriting on a physical paper. There is no need to get access to any memory on the pen to recover the information, which is especially useful in multi-writer or multi-pen scenarios. In this paper we focus on the compactness of the encoded information. The aim of this paper is to encode as much information as possible in short stroke sequences. In our experiments we show that we can embed more information in shorter strokes than in previous work. In straight lines as short as 5 cm, 32 bits can successfully be embedded. Furthermore, the new encoding scheme also works reliably on more complex patterns. Marcus Liwicki, Seiichi Uchida, Masakazu Iwamura, Shinichiro Omachi, Koichi Kise |
ICFHR | 1 |
| 2010 | Part-Based Recognition of Handwritten CharactersabstractIn the part-based recognition method proposed in this paper, a handwritten character image is represented by just a set of local parts. Then, each local part of the input pattern is recognized by a nearest-neighbor classifier. Finally, the category of the input pattern is determined by aggregating the local recognition results. This approach is opposed to conventional character recognition approaches which try to benefit from the global structure information as much as possible. Despite a pessimistic expectation, we have reached recognition rates much higher than 90% for a digit recognition task. In this paper we provide a detailed analysis in order to understand the results and find the merits of the local approach. Seiichi Uchida, Marcus Liwicki |
ICFHR | 2 |
| 2010 | a.SCAtch - A Sketch-Based Retrieval for Architectural Floor PlansabstractArchitects' daily routine means working with drawings. They use either a pen or a computer sketching their ideas or drawing to scale. When beginning a new project they often have to search for similar projects in the past. In this paper a sketch-based approach is proposed to query the floor plan repository. The user searches for semantically similar floor plans just by drawing the new plan. An algorithm extracts the semantic structure sketched by the architect on DFKI's Touch & Write table and compares the structure of the sketch with the ones from the floor plan repository. The a SCatch system enables the user to easily access knowledge from past projects. While in the current prototype only sketches with a predefined structure are recognized, we will extend the system to work with normal floor plans. Marcus Liwicki, Andreas Dengel 0001 |
ICFHR | 2 |
| 2010 | Analysis of Local Features for Handwritten Character RecognitionabstractThis paper investigates a part-based recognition method of handwritten digits. In the proposed method, the global structure of digit patterns is discarded by representing each pattern by just a set of local feature vectors. The method is then comprised of two steps. First, each of J local feature vectors of a target pattern is recognized into one of ten categories ("0''-"9'') by the nearest neighbor discrimination with a large database of reference vectors. Second, the category of the target pattern is determined by the majority voting on the J local recognition results. Despite a pessimistic expectation, we have reached recognition rates much higher than 90% for the task of digit recognition. Seiichi Uchida, Marcus Liwicki |
ICPR | 2 |
| 2009 | Combining Alignment Results for Historical Handwritten Document AnalysisabstractIn this paper we propose a new strategy for combining the outputs of several alignment systems. Based on the word boundaries retrieved from a number of individual alignment systems, the new boundaries are estimated. We investigate three strategies for this estimation. First, the mean value of the individual boundaries is taken, second the median is selected, and third, confidence values of the alignment systems are considered. We apply the combination strategies on a word mapping system for historical handwritten manuscripts. After some preprocessing and normalizing steps, three differently trained hidden Markov model based handwriting recognizers are applied to the text lines in forced alignment mode. As a result, the positions of the word boundaries are obtained. In in a number of experiments it is shown that a combination strategy based on the median outperforms the others and all individual alignment systems with a word mapping rate of about 95%. Emanuel Indermühle, Marcus Liwicki, Horst Bunke |
ICDAR | 2 |
| 2009 | Language Model Integration for the Recognition of Handwritten Medieval DocumentsabstractBuilding recognition systems for historical documents is a difficult task. Especially, when it comes to medieval scripts. The complexity is mainly affected by the poor quality and the small quantity of the data available. In this paper we apply an HMM based recognition system to medieval manuscripts from the 13th century written in Middle High German. The recognition system, which was originally developed for modern scripts, has been adapted to medieval scripts. Beside the data processing, one of the major challenges is to create a suitable language model. Because of the lack of appropriate independent text corpora for medieval languages, the language model has to be created on the base of a rather small number of manuscripts only. Due to the small size of the corpus, optimizing the language model parameters can quickly lead to the problem of overfitting. In this paper we describe a strategy to integrate all available information into the language model and to optimize the language model parameters without suffering from this problem. Markus Wüthrich, Marcus Liwicki, Andreas Fischer 0002, Emanuel Indermühle, Horst Bunke, Gabriel Viehhauser, Michael Stolz |
ICDAR | 2 |
| 2009 | Feature Selection for HMM and BLSTM Based Handwriting Recognition of Whiteboard NotesabstractIn this paper, we describe feature selection experiments for online handwriting recognition. We investigated a set of 25 online and pseudo-offline features to find out which features are important and which features may be redundant. To analyze the saliency of the features, we applied a sequential forward and a sequential backward search on the feature set. A hidden Markov model and a neural network based recognizer have been used as recognition engines. In our experiments, we obtained interesting results. Using a set of only five features, we achieved a performance similar to that of the reference system that uses all 25 features. The five selected features have a low correlation and have been the top choices during the first iterations of the forward search with both recognizers. Furthermore, for both recognizers, subsets have been identified that outperform the reference system with statistical significance. In order to assess the results more rigorously, we have compared our recognizer with the widely used commercial recognizer from Microsoft. Marcus Liwicki, Horst Bunke |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2009 | A Novel Connectionist System for Unconstrained Handwriting RecognitionabstractRecognizing lines of unconstrained handwritten text is a challenging task. The difficulty of segmenting cursive or overlapping characters, combined with the need to exploit surrounding context, has led to low recognition rates for even the best current recognizers. Most recent progress in the field has been made either through improved preprocessing or through advances in language modeling. Relatively little work has been done on the basic recognition algorithms. Indeed, most systems rely on the same hidden Markov models that have been used for decades in speech and handwriting recognition, despite their well-known shortcomings. This paper proposes an alternative approach based on a novel type of recurrent neural network, specifically designed for sequence labeling tasks where the data is hard to segment and contains long-range bidirectional interdependencies. In experiments on two large unconstrained handwriting databases, our approach achieves word recognition accuracies of 79.7 percent on online data and 74.1 percent on offline data, significantly outperforming a state-of-the-art HMM-based system. In addition, we demonstrate the network's robustness to lexicon size, measure the individual influence of its hidden layers, and analyze its use of context. Last, we provide an in-depth discussion of the differences between the network and HMMs, suggesting reasons for the network's superior performance. Alex Graves, Marcus Liwicki, Santiago Fernández, Roman Bertolami, Horst Bunke, Jürgen Schmidhuber |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Combining diverse on-line and off-line systems for handwritten text line recognition
Marcus Liwicki, Horst Bunke |
Pattern Recognit. | 1 |
| 2008 | Writer-Dependent Recognition of Handwritten Whiteboard Notes in Smart Meeting Room EnvironmentsabstractIn this paper we present a writer-dependent handwriting recognition system based on hidden Markov models (HMMs). This system, which has been developed in the context of research on smart meeting rooms, operates in two stages. First, a Gaussian mixture model (GMM)-based writer identification system developed for smart meeting rooms identifies the person writing on the whiteboard. Then a recognition system adapted to the individual writer is applied. Two different methods for obtaining writer-dependent recognizers are proposed. The first method uses the available writer-specific data to train an individual recognition system for each writer from scratch, while the second method takes a writer-independent recognizer and adapts it with the data from the considered writer. The experiments have been performed on the IAM-OnDB. In the first stage,the writer identification system produces a perfect identification rate. In the second stage, the writer-specific recognition system gets significantly better recognition results, compared to the writer-independent recognizer. The final word recognition rate on the IAM-OnDB-t1 benchmark task is close to 80 %. Marcus Liwicki, Andreas Schlapbach, Horst Bunke |
Document Analysis Systems | 1 |
| 2008 | A writer identification system for on-line whiteboard data
Andreas Schlapbach, Marcus Liwicki, Horst Bunke |
Pattern Recognit. | 2 |
| 2007 | Combining On-Line and Off-Line Systems for Handwriting RecognitionabstractIn this paper we present a new multiple classifier system (MCS)for recognizing notes written on a whiteboard. This MCS combines one off-line and two on-line handwriting recognition systems derived from previous work. The recognizers are all based on Hidden Markov Models but vary in the way of preprocessing and normalization. To combine the output sequences of the recognizers, we incrementally align the word sequences using a standard string matching algorithm. For deriving the final decision a voting strategy is applied. With the combination we could increase the system performance over the best individual recognizer by about 2%. Marcus Liwicki, Horst Bunke |
ICDAR | 1 |
| 2007 | On-Line Handwritten Text Line Detection Using Dynamic ProgrammingabstractIn this paper we propose a novel approach to th tion of on-line handwritten text lines based on dynamic programming. We try to find the paths with the minimum cost between two consecutive text lines. Most steps of the proposed algorithm are based on off-line information. Hence the method can also be applied to off-line documents after a few minor changes. In our experiments we show that this dynamic programming based approach is better than a common on-line segmentation procedure. Marcus Liwicki, Emanuel Indermühle, Horst Bunke |
ICDAR | 1 |
| 2007 | Unconstrained On-line Handwriting Recognition with Recurrent Neural NetworksabstractOn-line handwriting recognition is unusual among sequence labelling tasks in that the underlying generator of the observed data, i.e. the movement of the pen, is recorded directly. However, the raw data can be difficult to interpret because each letter is spread over many pen locations. As a consequence, sophisticated pre-processing is required to obtain inputs suitable for conventional sequence labelling algorithms, such as HMMs. In this paper we describe a system capable of directly transcribing raw on-line handwriting data. The system consists of a recurrent neural network trained for sequence labelling, combined with a probabilistic language model. In experiments on an unconstrained on-line database, we record excellent results using either raw or pre-processed data, well outperforming a benchmark HMM in both cases. Alex Graves, Santiago Fernández, Marcus Liwicki, Horst Bunke, Jürgen Schmidhuber |
NIPS | 3 |
| 2007 | Handwriting Recognition of Whiteboard Notes - Studying the Influence of Training Set Size and TypeabstractThis paper presents a system for the recognition of online whiteboard notes. Notes written on a whiteboard is a new modality in handwriting recognition research that has received relatively little attention in the past. For the recognition we use an offline HMM-recognizer, which is supplemented with methods for processing the online data and generating offline images. The system consists of six main modules: online preprocessing, transformation of online to offline data, offline preprocessing, feature extraction, classification and post-processing. The recognition rate of our basic recognizer in a writer independent experiment is 59.5%. By applying state-of-the-art methods, such as optimizing the number of states and Gaussian components, and by including a language model we could achieve a statistically significant increase of the recognition rate to 64.3%. To further improve the system performance we increased the size of the training set. For that we investigated two different strategies. First, we used another existing database of offline handwritten text. Second, we used a recently collected whiteboard database, called the IAM-OnDB. By means of these strategies the recognition rate could be further increased up to 68.5%. Marcus Liwicki, Horst Bunke |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Writer Identification for Smart Meeting Room Systems
Marcus Liwicki, Andreas Schlapbach, Horst Bunke, Samy Bengio, Johnny Mariéthoz, Jonas Richiardi |
Document Analysis Systems | 1 |
| 2006 | Chalklets: Developing Applications for a Board EnvironmentabstractThis paper presents a novel software framework and methodology to run applications on an interactive whiteboard. A new type of application called "Chalklets" has been designed. Chalklets use the chalkboard as an interface metaphor replacing the desktop. This allows the seamless integration of educational mini applications into electronic-whiteboard-supported lectures. This article describes the framework, presents a number of Chalklets, and discusses the concept, using the example of a Chalklet for simulating logic circuits sketched on the whiteboard Lars Knipping, Marcus Liwicki |
ISM | 2 |
| 2005 | Enhancing Training Data for Handwriting Recognition of Whiteboard Notes with Samples from a Different DatabaseabstractRecognition of unconstrained handwritten text is still a challenge. In this paper we consider a new problem, which is the recognition of notes written on a whiteboard. Our recognizer is based on hidden Markov models (HMMs). As it is difficult to acquire sufficient amounts of training data for the HMMs we propose two strategies for enlarging the training set. Both strategies are based on an existing database of offline handwritten text, which includes handwriting samples different from whiteboard data. The two proposed strategies are MAP adaptation and merging of training sets. With these methods we can achieve improvements of the word recognition rate of up to 5.7%. Marcus Liwicki, Horst Bunke |
ICDAR | 1 |
| 2005 | IAM-OnDB - an On-Line English Sentence Database Acquired from Handwritten Text on a WhiteboardabstractIn this paper we present IAM-OnDB - a new large online handwritten sentences database. It is publicly available and consists of text acquired via an electronic interface from a whiteboard. The database contains about 86 K word instances from an 11 K dictionary written by more than 200 writers. We also describe a recognizer for unconstrained English text that was trained and tested using this database. This recognizer is based on hidden Markov models (HMMs). In our experiments we show that by using larger training sets we can significantly increase the word recognition rate. This recognizer may serve as a benchmark reference for future research. Marcus Liwicki, Horst Bunke |
ICDAR | 1 |
| 2005 | Recognizing and Simulating Sketched Logic Circuits
Marcus Liwicki, Lars Knipping |
KES (3) | 1 |