EDBT 2026 Demo / reviewers in the wild / expert
Lambert Schomaker
dblp:27/2292 · also Lambert R. B. Schomaker
· DBLP profile ↗
100ranked-venue papers
9as first author
16since 2021 · last 2025
0000-0003-2351-930XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 20 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-HTR: A Novel Self-supervised Handwritten Text Recognition Framework Using Generative Adversarial Networks
Lisa Koopmans, Maruf A. Dhali, Lambert Schomaker |
ICDAR (3) | 3 |
| 2024 | Optimizing and interpreting the latent space of the conditional text-to-image GANsabstractAbstract Text-to-image generation intends to automatically produce a photo-realistic image, conditioned on a textual description. To facilitate the real-world applications of text-to-image synthesis, we focus on studying the following three issues: (1) How to ensure that generated samples are believable, realistic or natural? (2) How to exploit the latent space of the generator to edit a synthesized image? (3) How to improve the explainability of a text-to-image generation framework? We introduce two new data sets for benchmarking, i.e., the Good & Bad, bird and face, data sets consisting of successful as well as unsuccessful generated samples. This data set can be used to effectively and efficiently acquire high-quality images by increasing the probability of generating Good latent codes with a separate, new classifier. Additionally, we present a novel algorithm which identifies semantically understandable directions in the latent space of a conditional text-to-image GAN architecture by performing independent component analysis on the pre-trained weight values of the generator. Furthermore, we develop a background-flattening loss (BFL), to improve the background appearance in the generated images. Subsequently, we introduce linear-interpolation analysis between pairs of text keywords. This is extended into a similar triangular ‘linguistic’ interpolation. The visual array of interpolation results gives users a deep look into what the text-to-image synthesis model has learned within the linguistic embeddings. Experimental results on the recent DiverGAN generator, pre-trained on three common benchmark data sets demonstrate that our classifier achieves a better than 98% accuracy in predicting Good/Bad classes for synthetic samples and our proposed approach is able to derive various interpretable semantic properties for the text-to-image GAN model. Lambert Schomaker |
Neural Comput. Appl. | 2 |
| 2024 | Fusion-s2igan: an efficient and effective single-stage framework for speech-to-image generationabstractAbstract The goal of a speech-to-image transform is to produce a photo-realistic picture directly from a speech signal. Current approaches are based on a stacked modular framework that suffers from three vital issues: (1) Training separate networks is time-consuming, inefficient and the convergence of the final generative model depends on the previous generators; (2) The quality of precursor images is ignored; (3) Multiple discriminator networks need to be trained. We propose an efficient and effective single-stage framework called Fusion-S2iGan to yield perceptually plausible and semantically consistent image samples on the basis of spoken descriptions. Fusion-S2iGan introduces a visual+speech fusion module (VSFM), with a pixel-attention module (PAM), a speech-modulation module (SMM) and a weighted-fusion module (WFM), to inject the speech embedding from a speech encoder into the generator while improving the quality of synthesized pictures. The PAM module models the semantic affinities between pixel regions and by assigning larger weights to significant locations. The VSFM module adopts SMM to modulate visual feature maps using fine-grained linguistic cues present in the speech vector. Subsequently, the weighted-fusion model (WFM) captures the semantic importance of the image-attention mask and the speech-modulation module at the level of the channels, in an adaptive manner. Fusion-S2iGan spreads the bimodal information over all layers of the generator network to reinforce the visual feature maps at various hierarchical levels in the architecture. A series of experiments is conducted on four benchmark data sets: CUB birds, Oxford-102, Flickr8k and Places-subset. Results demonstrate the superiority of Fusion-S2iGan compared to the state-of-the-art models with a multi-stage architecture and a performance level that is close to traditional text-to-image approaches. Lambert Schomaker |
Neural Comput. Appl. | 2 |
| 2023 | The Effects of Character-Level Data Augmentation on Style-Based Dating of Historical ManuscriptsabstractIdentifying the production dates of historical manuscripts is one of the main goals for paleographers when studying ancient documents. Automatized methods can provide paleographers with objective tools to estimate dates more accurately. Previously, statistical features have been used to date digitized historical manuscripts based on the hypothesis that handwriting styles change over periods. However, the sparse availability of such documents poses a challenge in obtaining robust systems. Hence, the research of this article explores the influence of data augmentation on the dating of historical manuscripts. Linear Support Vector Machines were trained with k-fold cross-validation on textural and grapheme-based features extracted from historical manuscripts of different collections, including the Medieval Paleographical Scale, early Aramaic manuscripts, and the Dead Sea Scrolls. Results show that training models with augmented data improve the performance of historical manuscripts datin g by 1% - 3% in cumulative scores. Additionally, this indicates further enhancement possibilities by considering models specific to the features and the documents’ scripts Lisa Koopmans, Maruf A. Dhali, Lambert Schomaker |
ICPRAM | 3 |
| 2023 | Image-Based Material Analysis of Ancient Historical DocumentsabstractResearchers continually perform corroborative tests to classify ancient historical documents based on the physical materials of their writing surfaces. However, these tests, often performed on-site, requires actual access to the manuscript objects. The procedures involve a considerable amount of time and cost, and can damage the manuscripts. Developing a technique to classify such documents using only digital images can be very useful and efficient. In order to tackle this problem, this study uses images from a famous historical collection, the Dead Sea Scrolls, to propose a novel method to classify the materials of the manuscripts. The proposed classifier uses the two-dimensional Fourier Transform to identify patterns within the manuscript surfaces. Combining a binary classification system employing the transform with a majority voting process is shown to be effective for this classification task. This pilot study shows a successful classification percentage of up to 97% for a confi ned amount of manuscripts produced from either parchment or papyrus material. Feature vectors based on Fourier-space grid representation outperformed a concentric Fourier-space format. Thomas Reynolds, Maruf A. Dhali, Lambert Schomaker |
ICPRAM | 3 |
| 2023 | Correction to: A Limited-size ensemble of homogeneous CNN/LSTMS for high-performance word classificationabstractIn the original publication of this article, in Table 11 and Fig. 9, there was an error in the calculation of the weighted average of the word-accuracy values. The correct figure and table results are provided in this erratum report and turned out to be slightly higher. These weighted-average rates are dominated by the large KdK data set and are not the focus of the interpretation of the results: The differences within the individual data sets are more important to understand the effects of the conditions, i.e., dictionary size and ensemble application. Therefore, the miscalculation has no effect on the Discussion section. (Figure presented.) Comparison of the effect of the two label-coding schemes (Plain vs Extra-separator) and dictionary application on the single architecture and ensemble voting on the RIMES, the KdK, and GW data sets showing the weighted average word accuracy taking test-set sizes into account. Averaging was done over sets Weighted average of word accuracy (%) on the RIMES, KdK and GW data sets, using the dual-state word-beam search applying the Concise dictionary and the Extra-separator label-coding scheme, for the two CTC methods and single vs ensemble voting. Averaging was carried out over sets CTC decoder Framework Single Ensemble Best path 89.1 92.2 Dual-state word-beam search 96.2 97.0 The raw counts can be found in the Zenodo repository, "Erratum to: A limited-size ensemble of homogeneous CNN/LSTMs for high-performance word classification, doi: 10.1007/s00521-020-05612-0", (). Mahya Ameryan, Lambert Schomaker |
Neural Comput. Appl. | 2 |
| 2022 | Optimized latent-code selection for explainable conditional text-to-image GANsabstractThe task of text-to-image generation has achieved remarkable progress due to the advances in conditional generative adversarial networks (GANs). However, existing conditional text-to-image GANs approaches mostly concentrate on improving both image quality and semantic relevance but ignore the explainability of the model which plays a vital role in real-world applications. In this paper, we present a variety of techniques to take a deep look into the latent space and semantic space of a conditional text-to-image GANs model. We introduce pairwise linear interpolation of latent codes and ‘linguistic’ linear interpolation to study what the model has learned within the latent space and ‘linguistic’ embeddings. Subsequently, we extend linear interpolation to triangular interpolation conditioned on three corners to further analyze the model. After that, we build a Good/Bad data set containing unsuccessfully and successfully synthesized samples and corresponding latent codes for the image-quality research. Based on this data set, we propose a framework for finding good latent codes by utilizing a linear SVM. Experimental results on the recent DiverGAN generator trained on two benchmark data sets qualitatively prove the effectiveness of our presented techniques, with a better than 94% accuracy in predicting Good/Bad classes for latent vectors. The Good/Bad data set is publicly available at https://zenodo.org/record/5850224#.YeGMwP7MKUk. Lambert Schomaker |
IJCNN | 2 |
| 2022 | DiverGAN: An Efficient and Effective Single-Stage Framework for Diverse Text-to-Image GenerationabstractIn this paper, we concentrate on the text-to-image synthesis task that aims at automatically producing perceptually realistic pictures from text descriptions. Recently, several single-stage methods have been proposed to deal with the problems of a more complicated multi-stage modular architecture. However, they often suffer from the lack-of-diversity issue, yielding similar outputs given a single textual sequence. To this end, we present an efficient and effective single-stage framework (DiverGAN) to generate diverse, plausible and semantically consistent images according to a natural-language description. DiverGAN adopts two novel word-level attention modules, i.e., a channel-attention module (CAM) and a pixel-attention module (PAM), which model the importance of each word in the given sentence while allowing the network to assign larger weights to the significant channels and pixels semantically aligning with the salient words. After that, Conditional Adaptive Instance-Layer Normalization (CAdaILN) is introduced to enable the linguistic cues from the sentence embedding to flexibly manipulate the amount of change in shape and texture, further improving visual-semantic representation and helping stabilize the training. Also, a dual-residual structure is developed to preserve more original visual features while allowing for deeper networks, resulting in faster convergence speed and more vivid details. Furthermore, we propose to plug a fully-connected layer into the pipeline to address the lack-of-diversity problem, since we observe that a dense layer will remarkably enhance the generative capability of the network, balancing the trade-off between a low-dimensional random latent code contributing to variants and modulation modules that use high-dimensional and textual contexts to strength feature maps. Inserting a linear layer after the second residual block achieves the best variety and quality. Both qualitative and quantitative results on benchmark data sets demonstrate the superiority of our DiverGAN for realizing diversity, without harming quality and semantic consistency. Lambert Schomaker |
Neurocomputing | 2 |
| 2021 | An Investigation Into the Effect of the Learning Rate on Overestimation Bias of Connectionist Q-learningabstractIn Reinforcement learning, Q-learning is the best-known algorithm but it suffers from overestimation bias, which may lead to poor performance or unstable learning. In this paper, we present a novel analysis of this problem using various control tasks. For solving these tasks, Q-learning is combined with a multilayer perceptron (MLP), experience replay, and a target network. We focus our analysis on the effect of the learning rate when training the MLP. Furthermore, we examine if decaying the learning rate over time has advantages over static ones. Experiments have been performed using various maze-solving problems involving deterministic or stochastic transition functions and 2D or 3D grids and two Open-AI gym control problems. We conducted the same experiments with Double Q-learning using two MLPs with the same parameter settings, but without target networks. The results on the maze problems show that for Q-learning combined with the MLP, the overestimation occurs when higher learning rates are used and not when lower learning rates are used. The Double Q-learning variant becomes much less stable with higher learning rates and with low learning rates the overestimation bias may still occur. Overall, decaying learning rates clearly improves the performances of both Q-learning and Double Q-learning. Yifei Chen 0007, Lambert Schomaker, Marco A. Wiering |
ICAART (2) | 2 |
| 2021 | Self-Imitation Learning by PlanningabstractImitation learning (IL) enables robots to acquire skills quickly by transferring expert knowledge, which is widely adopted in reinforcement learning (RL) to initialize exploration. However, in long-horizon motion planning tasks, a challenging problem in deploying IL and RL methods is how to generate and collect massive, broadly distributed data such that these methods can generalize effectively. In this work, we solve this problem using our proposed approach called self-imitation learning by planning (SILP), where demonstration data are collected automatically by planning on the visited states from the current policy. SILP is inspired by the observation that successfully visited states in the early reinforcement learning stage are collision-free nodes in the graph-search based motion planner, so we can plan and relabel robot's own trials as demonstrations for policy learning. Due to these self-generated demonstrations, we relieve the human operator from the laborious data preparation process required by IL and RL methods in solving complex motion planning tasks. The evaluation results show that our SILP method achieves higher success rates and enhances sample efficiency compared to selected baselines, and the policy learned in simulation performs well in a real-world placement task with changing goals and obstacles. Sha Luo, Hamidreza Kasaei 0001, Lambert Schomaker |
ICRA | 3 |
| 2021 | Reinforcement Learning with Potential Functions Trained to Discriminate Good and Bad StatesabstractReward shaping is an efficient way to incorporate domain knowledge into a reinforcement learning agent. Nev-ertheless, it is unpractical and inconvenient to require prior knowledge for designing shaping rewards. Therefore, learning the shaping reward function by the agent during training could be more effective. In this paper, based on the potential-based reward shaping framework, which guarantees policy invariance, we propose to learn a potential function concurrently with training an agent using a reinforcement learning algorithm. In the proposed method, the potential function is trained by examining states that occur in good and in bad episodes. We apply the proposed adaptive potential function while training an agent with Q-learning and develop two novel algorithms. One is APF-QMLP, which applies the good/bad state potential function combined with Q-learning and multi-layer perceptrons (MLPs) to estimate the Q-function. The other is APF-Dueling-DQN, which combines the novel potential function with Dueling DQN. In particular, an autoencoder is adopted in APF-Dueling-DQN to map image states from Atari games to hash codes. We evaluated the created algorithms empirically in four environments: a six-room maze, CartPole, Acrobot, and Ms-Pacman, involving low-dimensional or high-dimensional state spaces. The experimental results showed that the proposed adaptive potential function improved the performances of the selected reinforcement learning algorithms. Yifei Chen 0007, Hamidreza Kasaei 0001, Lambert Schomaker, Marco A. Wiering |
IJCNN | 3 |
| 2021 | DTGAN: Dual Attention Generative Adversarial Networks for Text-to-Image GenerationabstractMost existing text-to-image generation methods adopt a multi-stage modular architecture which has three significant problems: 1) Training multiple networks increases the run time and affects the convergence and stability of the generative model; 2) These approaches ignore the quality of early-stage generator images; 3) Many discriminators need to be trained. To this end, we propose the Dual Attention Generative Adversarial Network (DTGAN) which can synthesize high-quality and semantically consistent images only employing a single generator/discriminator pair. The proposed model introduces channel-aware and pixel-aware attention modules that can guide the generator to focus on text-relevant channels and pixels based on the global sentence vector and to fine-tune original feature maps using attention weights. Also, Conditional Adaptive Instance-Layer Normalization (CAdaILN) is presented to help our attention modules flexibly control the amount of change in shape and texture by the input natural-language description. Furthermore, a new type of visual loss is utilized to enhance the image resolution by ensuring vivid shape and perceptually uniform color distributions of generated images. Experimental results on benchmark datasets demonstrate the superiority of our proposed method compared to the state-of-the-art models with a multi-stage framework. Visualization of the attention maps shows that the channel-aware attention module is able to localize the discriminative regions, while the pixel-aware attention module has the ability to capture the globally visual contents for the generation of an image. Lambert Schomaker |
IJCNN | 2 |
| 2021 | CentroidNetV2: A hybrid deep neural network for small-object segmentation and counting
Klaas Dijkstra, Jaap van de Loosdrecht, Waatze A. Atsma, Lambert Schomaker, Marco A. Wiering |
Neurocomputing | 4 |
| 2021 | A limited-size ensemble of homogeneous CNN/LSTMs for high-performance word classificationabstractAbstract The strength of long short-term memory neural networks (LSTMs) that have been applied is more located in handling sequences of variable length than in handling geometric variability of the image patterns. In this paper, an end-to-end convolutional LSTM neural network is used to handle both geometric variation and sequence variability. The best results for LSTMs are often based on large-scale training of an ensemble of network instances. We show that high performances can be reached on a common benchmark set by using proper data augmentation for just five such networks using a proper coding scheme and a proper voting scheme. The networks have similar architectures (convolutional neural network (CNN): five layers, bidirectional LSTM (BiLSTM): three layers followed by a connectionist temporal classification (CTC) processing step). The approach assumes differently scaled input images and different feature map sizes. Three datasets are used: the standard benchmark RIMES dataset (French); a historical handwritten dataset KdK (Dutch); the standard benchmark George Washington (GW) dataset (English). Final performance obtained for the word-recognition test of RIMES was 96.6%, a clear improvement over other state-of-the-art approaches which did not use a pre-trained network. On the KdK and GW datasets, our approach also shows good results. The proposed approach is deployed in the Monk search engine for historical-handwriting collections. Mahya Ameryan, Lambert Schomaker |
Neural Comput. Appl. | 2 |
| 2021 | GR-RNN: Global-context residual recurrent neural networks for writer identification
Lambert Schomaker |
Pattern Recognit. | 2 |
| 2021 | CT-Net: Cascade T-shape deep fusion networks for document binarizationabstractDocument binarization is a key step in most document analysis tasks. However, historical-document images usually suffer from various degradations, making this a very challenging processing stage. The performance of document image binarization has improved dramatically in recent years by the use of Convolutional Neural Networks (CNNs). In this paper, a dual-task, T-shaped neural network is proposed that has the main task of binarization and an auxiliary task of image enhancement. The neural network for enhancement learns the degradations in document images and the specific CNN-kernel features can be adapted towards the binarization task in the training process. In addition, the enhancement image can be considered as an improved version of the input image, which can be fed into the network for fine-tuning, making it possible to design a chained-cascade network (CT-Net). Experimental results on document binarization competition datasets (DIBCO datasets) and MCS dataset show that our proposed method outperforms competing state-of-the-art methods in most cases. Lambert Schomaker |
Pattern Recognit. | 2 |
| 2020 | Improving the robustness of LSTMs for word classification using stressed word endings in dual-state word-beam searchabstractIn recent years, long short-term memory neural networks (LSTMs) followed by a connectionist temporal classification (CTC) have shown strength in solving handwritten text recognition problems. Such networks can handle not only sequence variability but also geometric variation by using a convolutional front end, at the input side. Although different approaches have been introduced for decoding activations in the CTC output layer, only limited consideration is given to the use of proper label-coding schemes. In this paper, we use a limited-size ensemble of end-to-end convolutional LSTM Neural Networks to evaluate four label-coding schemes. Additionally, we evaluate two CTC search techniques: Best-path search vs dual-state word-beam search (DSWBS). The classifiers in the ensemble have comparable architectures but variable numbers of hidden units. We tested the coding and search approaches on three datasets: A standard benchmark IAM dataset (English) and two more difficult historical handwritten datasets (diaries and field notes, highly multilingual). Results show that stressing the word endings in the label-coding scheme yields a higher performance, especially for DSWBS. However, stressing the start-of-word shapes with a token appears to be disadvantageous. Mahya Ameryan, Lambert Schomaker |
ICFHR | 2 |
| 2020 | Recognizing Bengali Word Images - A Zero-Shot Learning PerspectiveabstractZero-Shot Learning(ZSL) techniques could classify a completely unseen class, which it has never seen before during training. Thus, making it more apt for any real-life classification problem, where it is not possible to train a system with annotated data for all possible class types. This work investigates recognition of word images written in Bengali Script in a ZSL framework. The proposed approach performs Zero-Shot word recognition by coupling deep learned features procured from various CNN architectures along with 13 basic shapes/stroke primitives commonly observed in Bengali script characters. As per the notion of ZSL framework those 13 basic shapes are termed as “Signature/Semantic Attributes”. The obtained results are promising while evaluation was carried out in a Five-Fold cross-validation setup dealing with samples from 250 word classes. Sukalpa Chanda, Daniel Haitink, Prashant Kumar Prasad, Jochem Baas, Umapada Pal 0001, Lambert Schomaker |
ICPR | 6 |
| 2020 | Accelerating Reinforcement Learning for Reaching Using Continuous Curriculum LearningabstractReinforcement learning has shown great promise in the training of robot behavior due to the sequential decision making characteristics. However, the required enormous amount of interactive and informative training data provides the major stumbling block for progress. In this study, we focus on accelerating reinforcement learning (RL) training and improving the performance of multi-goal reaching tasks. Specifically, we propose a precision-based continuous curriculum learning (PCCL) method in which the requirements are gradually adjusted during the training process, instead of fixing the parameter in a static schedule. To this end, we explore various continuous curriculum strategies for controlling a training process. This approach is tested using a Universal Robot 5e in both simulation and real-world multi-goal reach experiments. Experimental results support the hypothesis that a static training schedule is suboptimal, and using an appropriate decay function for curriculum learning provides superior results in a faster way. Sha Luo, Hamidreza Kasaei 0001, Lambert Schomaker |
IJCNN | 3 |
| 2020 | IMU-based Deep Neural Networks for Locomotor Intention PredictionabstractThis paper focuses on the design and comparison of different deep neural networks for the real-time prediction of locomotor intentions by using data from inertial measurement units. The deep neural network architectures are convolutional neural networks, recurrent neural networks, and convolutional recurrent neural networks. The input to the architectures are features in the time domain, which have been derived either from one inertial measurement unit placed on the upper right leg of ten healthy subjects, or two inertial measurement units placed on both the upper and lower right leg of ten healthy subjects. The study shows that a WaveNet, i.e., a full convolutional neural network, achieves a peak F1-score of 87.17% in the case of one IMU, and a peak of 97.88% in the case of two IMUs, with a 5-fold cross-validation. Huaitian Lu, Lambert Schomaker, Raffaella Carloni |
IROS | 2 |
| 2020 | Learning to Grasp 3D Objects using Deep Residual U-NetsabstractGrasp synthesis is one of the challenging tasks for any robot object manipulation task. In this paper, we present a new deep learning-based grasp synthesis approach for 3D objects. In particular, we propose an end-to-end 3D Convolutional Neural Network to predict the objects' graspable areas. We named our approach Res-U-Net since the architecture of the network is designed based on U-Net structure and residual network-styled blocks. It devised to plan 6-DOF grasps for any desired object, be efficient to compute and use, and be robust against varying point cloud density and Gaussian noise. We have performed extensive experiments to assess the performance of the proposed approach concerning graspable part detection, grasp success rate, and robustness to varying point cloud density and Gaussian noise. Experiments validate the promising performance of the proposed architecture in all aspects. A video showing the performance of our approach in the simulation environment can be found at http://youtu.be/5_yAJCc8owo. Lambert Schomaker, Hamidreza Kasaei 0001 |
RO-MAN | 2 |
| 2020 | One-vs-One classification for deep neural networksabstractFor performing multi-class classification, deep neural networks almost always employ a One-vs-All (OvA) classification scheme with as many output units as there are classes in a dataset. The problem of this approach is that each output unit requires a complex decision boundary to separate examples from one class from all other examples. In this paper, we propose a novel One-vs-One (OvO) classification scheme for deep neural networks that trains each output unit to distinguish between a specific pair of classes. This method increases the number of output units compared to the One-vs-All classification scheme but makes learning correct decision boundaries much easier. In addition to changing the neural network architecture, we changed the loss function, created a code matrix to transform the one-hot encoding to a new label encoding, and changed the method for classifying examples. To analyze the advantages of the proposed method, we compared the One-vs-One and One-vs-All classification methods on three plant recognition datasets (including a novel dataset that we created) and a dataset with images of different monkey species using two deep architectures. The two deep convolutional neural network (CNN) architectures, Inception-V3 and ResNet-50, are trained from scratch or pre-trained weights. The results show that the One-vs-One classification method outperforms the One-vs-All method on all four datasets when training the CNNs from scratch. However, when using the two classification schemes for fine-tuning pre-trained CNNs, the One-vs-All method leads to the best performances, which is presumably because the CNNs had been pre-trained using the One-vs-All scheme. Pornntiwa Pawara, Emmanuel Okafor, Marc Groefsema, Lambert Schomaker, Marco A. Wiering |
Pattern Recognit. | 5 |
| 2020 | Feature-extraction methods for historical manuscript dating based on writing style developmentabstractPaleographers and philologists perform significant research in finding the dates of ancient manuscripts to understand the historical contexts. To estimate these dates, the traditional process of using classical paleography is subjective, tedious, and often time-consuming. An automatic system based on pattern recognition techniques that infers these dates would be a valuable tool for scholars. In this study, the development of handwriting styles over time in the Dead Sea Scrolls, a collection of ancient manuscripts, is used to create a model that predicts the date of a query manuscript. In order to extract the handwriting styles, several dedicated feature-extraction techniques have been explored. Additionally, a self-organizing time map is used as a codebook. Support vector regression is used to estimate a date based on the feature vector of a manuscript. The date estimation from grapheme-based technique outperforms other feature-extraction techniques in identifying the chronological style development of handwriting in this study of the Dead Sea Scrolls. Maruf A. Dhali, Camilo Nathan Jansen, Jan Willem de Wit, Lambert Schomaker |
Pattern Recognit. Lett. | 4 |
| 2020 | FragNet: Writer Identification Using Deep Fragment NetworksabstractWriter identification based on a small amount of text is a challenging problem. In this paper, we propose a new benchmark study for writer identification based on word or text block images which approximately contain one word. In order to extract powerful features on these word images, a deep neural network, named FragNet, is proposed. The FragNet has two pathways: feature pyramid which is used to extract feature maps and fragment pathway which is trained to predict the writer identity based on fragments extracted from the input image and the feature maps on the feature pyramid. We conduct experiments on four benchmark datasets, which show that our proposed method can generate efficient and robust deep representations for writer identification based on both word and page images. Lambert Schomaker |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Continuous Learning in Large-scale Problems: The Case of Multi-script Historical Handwritten Document Collections
Lambert Schomaker |
ICAART (1) | 1 |
| 2019 | No Padding Please: Efficient Neural Handwriting RecognitionabstractNeural handwriting recognition (NHR) is the recognition of handwritten text with deep learning models, such as multi-dimensional long short-term memory (MDLSTM) recurrent neural networks. Models with MDLSTM layers have achieved state-of-the art results on handwritten text recognition tasks. While multi-directional MDLSTM-layers have an unbeaten ability to capture the complete context in all directions, this strength limits the possibilities for parallelization, and therefore comes at a high computational cost. In this work we develop methods to create efficient MDLSTM-based models for NHR, particularly a method aimed at eliminating computation waste that results from padding. This proposed method, called example packing, replaces wasteful stacking of padded examples with efficient tiling in a 2-dimensional grid. For word-based NHR this yields a speed improvement of factor 6.6 over an already efficient baseline of minimal padding for each batch separately. For line-based NHR the savings are more modest, but still significant. In addition to example packing, we propose: 1) a technique to optimize parallelization for dynamic graph definition frameworks including PyTorch, using convolutions with grouping, 2) a method for parallelization across GPUs for variable-length example batches. All our techniques are thoroughly tested on our own PyTorch re-implementation of MDLSTM-based NHR models. A thorough evaluation on the IAM dataset shows that our models are performing similar to earlier implementations of state-of-the art models. Our efficient NHR model and some of the reusable techniques discussed with it offer ways to realize relatively efficient models for the omnipresent scenario of variable-length inputs in deep learning. Gideon Maillette de Buy Wenniger, Lambert Schomaker, Andy Way |
ICDAR | 2 |
| 2019 | Multi-script text versus non-text classification of regions in scene imagesabstractText versus non-text region classification is an essential but difficult step in scene-image analysis due to the considerable shape complexity of text and background patterns. There exists a high probability of confusion between background elements and letter parts. This paper proposes a feature-based classification of image blocks using the color autocorrelation histogram (CAH) and the scale-invariant feature transform (SIFT) algorithm, yielding a combined scale and color-invariant feature suitable for scene-text classification. For the evaluation, features were extracted from different color spaces, applying color-histogram autocorrelation. The color features are adjoined with a SIFT descriptor. Parameter tuning is performed and evaluated. For the classification, a standard nearest-neighbor (1NN) and a support-vector machine (SVM) were compared. The proposed method appears to perform robustly and is especially suitable for Asian scripts such as Kannada and Thai, where urban scene-text fonts are characterized by a high curvature and salient color variations. Bowornrat Sriman, Lambert Schomaker |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Hyperspectral demosaicking and crosstalk correction using deep learning
Klaas Dijkstra, Jaap van de Loosdrecht, Lambert Schomaker, Marco A. Wiering |
Mach. Vis. Appl. | 3 |
| 2019 | Deep adaptive learning for writer identification based on single handwritten word images
Lambert Schomaker |
Pattern Recognit. | 2 |
| 2019 | DeepOtsu: Document enhancement and binarization using iterative deep learningabstractThis paper presents a novel iterative deep learning framework and applies it to document enhancement and binarization. Unlike the traditional methods that predict the binary label of each pixel on the input image, we train the neural network to learn the degradations in document images and produce uniform images of the degraded input images, which in turn allows the network to refine the output iteratively. Two different iterative methods have been studied in this paper: recurrent refinement (RR) that uses the same trained neural network in each iteration for document enhancement and stacked refinement (SR) that uses a stack of different neural networks for iterative output refinement. Given the learned nature of the uniform and enhanced image, the binarization map can be easily obtained through use of a global or local threshold. The experimental results on several public benchmark data sets show that our proposed method provides a new, clean version of the degraded image, one that is suitable for visualization and which shows promising results for binarization using Otsu’s global threshold, based on enhanced images learned iteratively by the neural network. Lambert Schomaker |
Pattern Recognit. | 2 |
| 2018 | Deep Learning for Classification and as Tapped-Feature Generator in Medieval Word-Image RecognitionabstractHistorical manuscripts are the main source of information about past. In recent years, digitization of large quantities of historical handwritten documents is in vogue. This trend gives access to a plethora of information about our medieval past. Such digital archives can be more useful if automatic indexing and retrieval of document images can be provided to the end users of a digital library. An automatic transcription of the full digital archive using traditional Optical Character Recognition (OCR) is still not possible with sufficient accuracy. If full transcription is not available, the end users are interested in indexing and retrieving of particular document pages of their interest. Hence recognition of certain keywords from within the corpus will be sufficient to meet the end users needs. Recently, deep-learning based methods have shown competence in image classification problems. However, one bottleneck with deep-learning based techniques is that it requires a huge amount of training samples per class. Since the number of samples per word class is scarce for collections that are freshly scanned, this is a serious hindrance for direct usage of the deep-learning technique for the purpose of word image recognition in historical document images. This paper aims to investigate the problem of recognizing words from historical document images using a deep-learning based framework for feature extraction and classification while countering the problem of the low amount of image samples using off-line data augmentation techniques. Encouraging results (highest accuracy of 90.03%) were obtained while dealing with 365 different word classes. Sukalpa Chanda, Emmanuel Okafor, Sébastien Hamel, Dominique Stutzmann, Lambert Schomaker |
DAS | 5 |
| 2018 | Deep Learning Policy QuantizationabstractWe introduce a novel type of actor-critic approach for deep reinforcement learning which is based on learning vector quantization. We replace the softmax operator of the policy with a more general and more flexible operator that is similar to the robust soft learning vector quantization algorithm. We compare our approach to the default A3C architecture on three Atari 2600 games and a simplistic game called Catch. We show that the proposed algorithm outperforms the softmax architecture on Catch. On the Atari games, we observe a nonunanimous pattern in terms of the best performing model. Jos van de Wolfshaar, Marco A. Wiering, Lambert Schomaker |
ICAART (2) | 3 |
| 2018 | Detection and Recognition of Badgers Using Deep Learning
Emmanuel Okafor, Gerard Berendsen, Lambert Schomaker, Marco A. Wiering |
ICANN (3) | 3 |
| 2018 | Zero-Shot Learning Based Approach For Medieval Word Recognition using Deep-Learned FeaturesabstractHistorical manuscripts reflect our past. Recently digitization of large quantities of historical handwritten documents is taking place in every corner of the world, and are being archived. From those digital repositories, automatic text indexing and retrieval system fetch only those documents to an end user that they are interested in. A regular OCR technology is not capable of rendering this service to an end user in a reliable manner. Instead, a word recognition/spotting algorithm performs the task. Word recognition based systems require enough labelled data per class to train the system. Moreover, all word classes need to be taught beforehand. Though word spotting could evade this drawback of prior training, these systems often need to have additional overheads like a language model to deal with "out of lexicon" words. Zero-shot learning could be a possible alternative to counter such situation. A Zero-shot learning algorithm is capable of handling unseen classes, provided the algorithm has been fortified with rich discriminating features and reliable "attribute description" per class during training. Since deeply learned features have enough discriminating power, a deep learning framework has been used here for feature extraction purpose. To the best of our knowledge, this is probably the first work on "out of lexicon" medieval word recognition using a Zero-Shot Learning framework. We obtained very encouraging results(accuracy ≈57% for "out of lexicon" classes) while dealing with 166 training classes and 50 unseen test classes. Sukalpa Chanda, Jochem Baas, Daniel Haitink, Sébastien Hamel, Dominique Stutzmann, Lambert Schomaker |
ICFHR | 6 |
| 2018 | A Deep Convolutional Neural Network for Location Recognition and Geometry based InformationabstractIn this paper we propose a new approach to Deep Neural Networks (DNNs) based on the particular needs of navigation tasks. To investigate these needs we created a labeled image dataset of a test environment and we compare classical computer vision approaches with the state of the art in image classification. Based on these results we have developed a new DNN architecture that outperforms previous architectures in recognizing locations, relying on the geometrical features of the images. In particular we show the negative effects of scale, rotation, and position invariance properties of the current state of the art DNNs on the task. We finally show the results of our proposed architecture that preserves the geometrical properties. Our experiments show that our method outperforms the state of the art image classification networks in recognizing locations. Francesco Bidoia, Matthia Sabatelli, Amirhossein Shantia, Marco A. Wiering, Lambert Schomaker |
ICPRAM | 5 |
| 2018 | CentroidNet: A Deep Neural Network for Joint Object Localization and Counting
Klaas Dijkstra, Jaap van de Loosdrecht, Lambert Schomaker, Marco A. Wiering |
ECML/PKDD (3) | 3 |
| 2017 | Data Augmentation for Plant Classification
Pornntiwa Pawara, Emmanuel Okafor, Lambert Schomaker, Marco A. Wiering |
ACIVS | 3 |
| 2017 | Hyper-spectral frequency selection for the classification of vegetation diseases
Klaas Dijkstra, Jaap van de Loosdrecht, Lambert Schomaker, Marco A. Wiering |
ESANN | 3 |
| 2017 | A Digital Palaeographic Approach towards Writer Identification in the Dead Sea ScrollsabstractTo understand the historical context of an ancient manuscript, scholars rely on the prior knowledge of writer and date of that document. In this paper, we study the Dead Sea Scrolls, a collection of ancient manuscripts with immense historical, religious, and linguistic significance, which was discovered in the mid-20th century near the Dead Sea. Most of the manuscripts of this collection have become digitally available only recently and techniques from the pattern recognition field can be applied to revise existing hypotheses on the writers and dates of these scrolls. This paper presents our ongoing work which aims to introduce digital palaeography to the field and generate fresh empirical data by means of pattern recognition and artificial intelligence. Challenges in analyzing the Dead Sea Scrolls are highlighted by a pilot experiment identifying the writers using several dedicated features. Finally, we discuss whether to use specifically-designed shape features for writer identification or to use the Deep Learning methods on a relatively limited ancient manuscript collection which is degraded over the course of time and is not labeled, as in the case of the Dead Sea Scrolls. Maruf A. Dhali, Mladen Popovic, Eibert Tigchelaar, Lambert Schomaker |
ICPRAM | 5 |
| 2017 | Comparing Local Descriptors and Bags of Visual Words to Deep Convolutional Neural Networks for Plant RecognitionabstractThe use of machine learning and computer vision methods for recognizing different plants from images has attracted lots of attention from the community. This paper aims at comparing local feature descriptors and bags of visual words with different classifiers to deep convolutional neural networks (CNNs) on three plant datasets; AgrilPlant, LeafSnap, and Folio. To achieve this, we study the use of both scratch and fine-tuned versions of the GoogleNet and the AlexNet architectures and compare them to a local feature descriptor with k-nearest neighbors and the bag of visual words with the histogram of oriented gradients combined with either support vector machines and multi-layer perceptrons. The results shows that the deep CNN methods outperform the hand-crafted features. The CNN techniques can also learn well on a relatively small dataset, Folio. Pornntiwa Pawara, Emmanuel Okafor, Olarik Surinta, Lambert Schomaker, Marco A. Wiering |
ICPRAM | 4 |
| 2017 | Operational data augmentation in classifying single aerial images of animalsabstractIn deep learning, data augmentation is important to increase the amount of training images to obtain higher classification accuracies. Most data-augmentation methods adopt the use of the following techniques: cropping, mirroring, color casting, scaling and rotation for creating additional training images. In this paper, we propose a novel data-augmentation method that transforms an image into a new image containing multiple rotated copies of the original image in the operational classification stage. The proposed method creates a grid of n×n cells, in which each cell contains a different randomly rotated image and introduces a natural background in the newly created image. This algorithm is used for creating new training and testing images, and enhances the amount of information in an image. For the experiments, we created a novel dataset with aerial images of cows and natural scene backgrounds using an unmanned aerial vehicle, resulting in a binary classification problem. To classify the images, we used a convolutional neural network (CNN) architecture and compared two loss functions (Hinge loss and cross-entropy loss). Additionally, we compare the CNN to classical feature-based techniques combined with a k-nearest neighbor classifier or a support vector machine. The results show that the pre-trained CNN with our proposed data-augmentation technique yields significantly higher accuracies than all other approaches. Emmanuel Okafor, Rik Smit, Lambert Schomaker, Marco A. Wiering |
INISTA | 3 |
| 2017 | Beyond OCR: Multi-faceted understanding of handwritten document characteristicsabstractHandwritten document understanding is a fundamental research problem in pattern recognition and it relies on the effective features. In this paper, we propose a joint feature distribution (JFD) principle to design novel discriminative features which could be the joint distribution of features on adjacent positions or the joint distribution of different features on the same location. Following the proposed JFD principle, we introduce seventeen features, including twelve textural-based and five grapheme-based features. We evaluate these features for different applications from four different perspectives to understand handwritten documents beyond OCR (optical character reognition), by writer identification, script recognition, historical manuscript dating and localization. Extensive experimental results demonstrate that our novel QuadHinge and CoHinge features following the JFD principle provide promising results on these four applications. Lambert Schomaker |
Pattern Recognit. | 2 |
| 2017 | Writer identification using curvature-free features
Lambert Schomaker |
Pattern Recognit. | 2 |
| 2016 | General Pattern Run-Length Transform for Writer IdentificationabstractIn this paper we present a novel textural-based feature for writer identification: the General Pattern Run-Length Transform (GPRLT), which is the histogram of the run-length of any complex patterns. The GPRLT can be computed on the binary images (GPRLT bin) or on the gray scale images (GPRLT gray) without using any binarization or segmentation methods. Experimental results show that the GPRLT gray achieves even higher performance than the GPRLT bin for writer identification. The writer identification performance on the challenging CERUG-EN data set demonstrates that the proposed methods outperform state-of-the-art algorithms. Our source code and data set are available on www.ai.rug.nl/~sheng/dflib. Lambert Schomaker |
DAS | 2 |
| 2016 | Historical Document Dating Using Unsupervised Attribute LearningabstractThe date of historical documents is an important metadata for scholars using them, as they need to know the historical context of the documents. This paper presents a novel attribute representation for medieval documents to automatically estimate the date information, which are the years they had been written. Non-semantic attributes are discovered in the low-level feature space using an unsupervised attribute learning method. A negative data set is involved in the attribute learning to make sure that our system rejects the documents which are not from the Middle Ages nor from the same archives. Experimental results on the basis of the Medieval Paleographic Scale (MPS) data set demonstrate that the proposed method achieves the state-of-the-art result. Petros Samara, Jan Burgers, Lambert Schomaker |
DAS | 4 |
| 2016 | Co-occurrence Features for Writer IdentificationabstractIn this paper, we propose two novel textural-based features for writer identification: CoHinge and QuadHinge which are based on the spatial and attribute co-occurrence of the Hinge kernel. The CoHinge feature is the joint distribution of the Hinge kernel on two different pixels of writing contours and the QuadHinge feature is the joint distribution of angles and curvature information of contour fragments. We evaluate the proposed features on five benchmark data sets and their combined large set and the experimental results demonstrate the discriminative and powerful of the proposed features. Lambert Schomaker |
ICFHR | 2 |
| 2016 | Discovering Visual Element Evolutions for Historical Document DatingabstractDiscovering visual elements correlated with temporal information in images is a challenging problem. In this paper, we study this problem with regard to handwritten historical document dating. We propose a novel stroke descriptor based on a scale-invariant log-polar space using the stroke width as the scale factor. Furthermore, the primary stroke shapes in documents are generated and termed stroke shape elements (also called visual elements in this paper). To discover the changes in visual elements over time, the Evolutionary Self-Organizing Map (ESOM) is proposed with a new time dimension based on the standard Kohonen's map to preserve the time topology. The proposed ESOM is a weakly-supervised learning method, integrating the visual elements mining and the evolution learning into one framework to preserve the date topology and time topology simultaneously. The dating of historical documents is performed by voting stroke shape elements based on their labels estimated from the the trained ESOM codebook to the label space, yielding a probability distribution. Experimental results demonstrate the effectiveness of the proposed approach for historical document dating. Petros Samara, Jan Burgers, Lambert Schomaker |
ICFHR | 4 |
| 2016 | Evaluating automatically parallelized versions of the support vector machineabstractSummary The support vector machine (SVM) is a supervised learning algorithm used for recognizing patterns in data. It is a very popular technique in machine learning and has been successfully used in applications such as image classification, protein classification, and handwriting recognition. However, the computational complexity of the kernelized version of the algorithm grows quadratically with the number of training examples. To tackle this high computational complexity, we have developed a directive‐based approach that converts a gradient‐ascent based training algorithm for the CPU to an efficient graphics processing unit (GPU) implementation. We compare our GPU‐based SVM training algorithm to the standard LibSVM CPU implementation, a highly optimized GPU‐LibSVM implementation, as well as to a directive‐based OpenACC implementation. The results on different handwritten digit classification datasets demonstrate an important speed‐up for the current approach when compared to the CPU and OpenACC versions. Furthermore, our solution is almost as fast and sometimes even faster than the highly optimized CUBLAS‐based GPU‐LibSVM implementation, without sacrificing the algorithm's accuracy. Copyright © 2014 John Wiley & Sons, Ltd. Valeriu Codreanu, Bob Dröge, David Williams 0002, Burhan Yasar, Po Yang 0001, Baoquan Liu, Feng Dong 0005, Olarik Surinta, Lambert Schomaker, Jos B. T. M. Roerdink, Marco A. Wiering |
Concurr. Comput. Pract. Exp. | 9 |
| 2016 | Historical manuscript dating based on temporal pattern codebookabstractManuscript dating is an essential part of historical scholarship. This paper proposes a framework for image-based historical manuscript dating based on handwritten pattern analysis in scanned historical manuscript images. We first use a singular structural feature to extract the mid-level handwritten patterns in historical document images and then encode the discovered handwritten patterns based on a codebook which contains the temporal information. We evaluate our method on the Medieval Paleographic Scale (MPS) data set and experimental results demonstrate that the feature representation based on the codebook which contains temporal information is more discriminative and powerful for dating. In addition, our proposed method can also visualize the evolution of handwritten patterns over time. Petros Samara, Jan Burgers, Lambert Schomaker |
Comput. Vis. Image Underst. | 4 |
| 2016 | Musicologist-driven writer identification in early music manuscripts
Masahiro Niitsuma, Lambert Schomaker, Jean-Paul van Oosten, Yo Tomita, David Bell |
Multim. Tools Appl. | 2 |
| 2016 | Image-based historical manuscript dating using contour and stroke fragments
Petros Samara, Jan Burgers, Lambert Schomaker |
Pattern Recognit. | 4 |
| 2016 | Bangla Handwritten Character Segmentation Using Structural Features: A Supervised and Bootstrapping ApproachabstractIn this article, we propose a new framework for segmentation of Bangla handwritten word images into meaningful individual symbols or pseudo-characters. Existing segmentation algorithms are not usually treated as a classification problem. However, in the present study, the segmentation algorithm is looked upon as a two-class supervised classification problem. The method employs an SVM classifier to select the segmentation points on the word image on the basis of various structural features. For training of the SVM classifier, an unannotated training set is prepared first using candidate segmenting points. The training set is then clustered, and each cluster is labeled manually with minimal manual intervention. A semi-automatic bootstrapping technique is also employed to enlarge the training set from new samples. The overall architecture describes a basic step toward building an annotation system for the segmentation problem, which has not so far been investigated. The experimental results show that our segmentation method is quite efficient in segmenting not only word images but also handwritten texts. As a part of this work, a database of Bangla handwritten word images has also been developed. Considering our data collection method and a statistical analysis of our lexicon set, we claim that the relevant characteristics of an ideal lexicon set are present in our handwritten word image database. Tapan Kumar Bhowmik, Swapan K. Parui, Utpal Roy, Lambert Schomaker |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2016 | A Multiple-Label Guided Clustering Algorithm for Historical Document Dating and LocalizationabstractIt is of essential importance for historians to know the date and place of origin of the documents they study. It would be a huge advancement for historical scholars if it would be possible to automatically estimate the geographical and temporal provenance of a handwritten document by inferring them from the handwriting style of such a document. We propose a multiple-label guided clustering algorithm to discover the correlations between the concrete low-level visual elements in historical documents and abstract labels, such as date and location. First, a novel descriptor, called histogram of orientations of handwritten strokes, is proposed to extract and describe the visual elements, which is built on a scale-invariant polar-feature space. In addition, the multi-label self-organizing map (MLSOM) is proposed to discover the correlations between the low-level visual elements and their labels in a single framework. Our proposed MLSOM can be used to predict the labels directly. Moreover, the MLSOM can also be considered as a pre-structured clustering method to build a codebook, which contains more discriminative information on date and geography. The experimental results on the medieval paleographic scale data set demonstrate that our method achieves state-of-the-art results. Petros Samara, Jan Burgers, Lambert Schomaker |
IEEE Trans. Image Process. | 4 |
| 2015 | Recognizing Handwritten Characters with Local Descriptors and Bags of Visual Words
Olarik Surinta, Mahir Faik Karaaba, Tusar Kanti Mishra, Lambert Schomaker, Marco A. Wiering |
EANN | 4 |
| 2015 | A Polar Stroke Descriptor for classification of historical documentsabstractThis paper presents a method for extracting a local scale- and rotation-invariant descriptor called Polar Stroke Descriptor (PSD) from the strokes in handwritten documents. Our descriptor captures the stroke-length distribution on a reference point and results in a robust feature vector. Furthermore, we develop a compact stroke representation termed as strokelets inspired by the classical Bag of Words model. We apply the proposed representation to historical document dating and the experimental results show that the proposed method achieves state-of-the-art performance on the MPS database. Lambert Schomaker |
ICDAR | 2 |
| 2015 | Object Attention Patches for Text Detection and Recognition in Scene Images using SIFT
Bowornrat Sriman, Lambert Schomaker |
ICPRAM (1) | 2 |
| 2015 | Indoor localization by denoising autoencoders and semi-supervised learning in 3D simulated environmentabstractRobotic mapping and localization methods are mostly dominated by using a combination of spatial alignment of sensory inputs, loop closure detection, and a global fine-tuning step. This requires either expensive depth sensing systems, or fast computational hardware at run-time to produce a 2D or 3D map of the environment. In a similar context, deep neural networks are used extensively in scene recognition applications, but are not yet applied to localization and mapping problems. In this paper, we adopt a novel approach by using denoising autoencoders and image information for tackling robot localization problems. We use semi-supervised learning with location values that are provided by traditional mapping methods. After training, our method requires much less run-time computations, and therefore can perform real-time localization on normal processing units. We compare the effects of different feature vectors such as plain images, the scale invariant feature transform and histograms of oriented gradients on the localization precision. The best system can localize with an average positional error of ten centimeters and an angular error of four degrees in 3D simulation. Amirhosein Shantia, Rik Timmers, Lambert Schomaker, Marco A. Wiering |
IJCNN | 3 |
| 2015 | Recognition of handwritten characters using local gradient feature descriptors
Olarik Surinta, Mahir Faik Karaaba, Lambert Schomaker, Marco A. Wiering |
Eng. Appl. Artif. Intell. | 3 |
| 2015 | Junction detection in handwritten documents and its application to writer identification
Marco A. Wiering, Lambert Schomaker |
Pattern Recognit. | 3 |
| 2014 | Towards Style-Based Dating of Historical DocumentsabstractEstimating the date of undated medieval manuscripts by evaluating the script they contain, using document image analysis, is helpful for scholars of various disciplines studying the Middle Ages. However, there are, as yet, no systems to automatically and effectively infer the age of historical scripts using machine learning methods. To build a system to date medieval documents is a challenging problem in several aspects: 1) As yet, no suitable reference dataset of medieval handwriting exists, 2) relatively little is known about the evolution of writing styles in the Middle Ages, and especially in the later Middle Ages. Our Medieval Paleographic Scale (MPS) project aims at solving these problems. We have collected a corpus of charters from the Medieval Dutch language area, dating from the period 1300 to 1550. A global and local regression method is proposed for learning and estimating the year in which these documents were written, using several features which have been successfully used in writer identification. The proposed system can serve as a basic tool for the medievalist or paleographer. The experimental results of the proposed method demonstrate its effectiveness. Petros Samara, Jan Burgers, Lambert Schomaker |
ICFHR | 4 |
| 2014 | A Reevaluation and Benchmark of Hidden Markov ModelsabstractHidden Markov models are frequently used in handwriting-recognition applications. While a large number of methodological variants have been developed to accommodate different use cases, the core concepts have not been changed much. In this paper, we develop a number of datasets to benchmark our own implementation as well as various other tool kits. We introduce a gradual scale of difficulty that allows comparison of datasets in terms of separability of classes. Two experiments are performed to review the basic HMM functions, especially aimed at evaluating the role of the transition probability matrix. We found that the transition matrix may be far less important than the observation probabilities. Furthermore, the traditional training methods are not always able to find the proper (true) topology of the transition matrix. These findings support the view that the quality of the features may require more attention than the aspect of temporal modelling addressed by HMMs. Jean-Paul van Oosten, Lambert Schomaker |
ICFHR | 2 |
| 2014 | A Path Planning for Line Segmentation of Handwritten DocumentsabstractThis paper describes the use of a novel A path-planning algorithm for performing line segmentation of handwritten documents. The novelty of the proposed approach lies in the use of a smart combination of simple soft cost functions that allows an artificial agent to compute paths separating the upper and lower text fields. The use of soft cost functions enables the agent to compute near-optimal separating paths even if the upper and lower text parts are overlapping in particular places. We have performed experiments on the Saint Gall and Monk line segmentation (MLS) datasets. The experimental results show that our proposed method performs very well on the Saint Gall dataset, and also demonstrate that our algorithm is able to cope well with the much more complicated MLS dataset. Olarik Surinta, Michiel Holtkamp, Mahir Faik Karaaba, Jean-Paul van Oosten, Lambert Schomaker, Marco A. Wiering |
ICFHR | 5 |
| 2014 | Delta-n Hinge: Rotation-Invariant Features for Writer IdentificationabstractThis paper presents a method for extracting rotation-invariant features from images of handwriting samples that can be used to perform writer identification. The proposed features are based on the Hinge feature [1], but incorporating the derivative between several points along the ink contours. Finally, we concatenate the proposed features into one feature vector to characterize the writing styles of the given handwritten text. The proposed method has been evaluated using Fire maker and IAM datasets in writer identification, showing promising performance gains. Lambert Schomaker |
ICPR | 2 |
| 2014 | Machine learning for multi-view eye-pair detection
Mahir Faik Karaaba, Lambert Schomaker, Marco A. Wiering |
Eng. Appl. Artif. Intell. | 2 |
| 2014 | Separability versus prototypicality in handwritten word-image retrieval
Jean-Paul van Oosten, Lambert Schomaker |
Pattern Recognit. | 2 |
| 2013 | Writer Identification in Old Music Manuscripts Using Contour-Hinge Feature and Dimensionality Reduction with an Autoencoder
Masahiro Niitsuma, Lambert Schomaker, Jean-Paul van Oosten, Yo Tomita |
CAIP (2) | 2 |
| 2013 | A Comparison of Feature and Pixel-Based Methods for Recognizing Handwritten Bangla DigitsabstractWe propose a novel handwritten character recognition method for isolated handwritten Bangla digits. A feature is introduced for such patterns, the contour angular technique. It is compared to other methods, such as the hotspot feature, the gray-level normalized character image and a basic low-resolution pixel-based method. One of the goals of this study is to explore performance differences between dedicated feature methods and the pixel-based methods. The four methods are compared with support vector machine (SVM) classifiers on the collection of handwritten Bangla digit images. The results show that the fast contour angular technique outperforms the other techniques when not very many training examples are used. The fast contour angular technique captures aspects of curvature of the handwritten image and results in much faster character classification than the gray pixel-based method. Still, this feature obtains a similar recognition compared to the gray pixel-based method when a large training set is used. In order to investigate further whether the different feature methods represent complementary aspects of shape, the effect of majority voting is explored. The results indicate that the majority voting method achieves the best recognition performance on this dataset. Olarik Surinta, Lambert Schomaker, Marco A. Wiering |
ICDAR | 2 |
| 2012 | Separability versus Prototypicality in Handwritten Word RetrievalabstractUser appreciation of a word-image retrieval system is based on the quality of a hit list for a query. Using support vector machines for ranking in large scale, handwritten document collections, we observed that many hit lists suffered from bad instances in the top ranks. An analysis of this problem revealed that two functions needed to be optimised concerning both separability and prototypicality. By ranking images in two stages, the number of distracting images is reduced, making the method very convenient for massive scale, continuously trainable retrieval engines. Instead of cumbersome SVM training, we present a nearest-centroid method and show that precision improvements of up to 35 percentage points can be achieved, yielding up to 100% precision in data sets with a large amount of instances, while maintaining high recall performances. Jean-Paul van Oosten, Lambert Schomaker |
ICFHR | 2 |
| 2012 | Handwritten Character Classification using the Hotspot Feature Extraction Technique
Olarik Surinta, Lambert Schomaker, Marco A. Wiering |
ICPRAM (1) | 2 |
| 2012 | Writer identification using directional ink-trace width measurements
Axel Brink, J. Smit, Marius Bulacu, Lambert Schomaker |
Pattern Recognit. | 4 |
| 2011 | Reinforcement learning algorithms for solving classification problemsabstractWe describe a new framework for applying reinforcement learning (RL) algorithms to solve classification tasks by letting an agent act on the inputs and learn value functions. This paper describes how classification problems can be modeled using classification Markov decision processes and introduces the Max-Min ACLA algorithm, an extension of the novel RL algorithm called actor-critic learning automaton (ACLA). Experiments are performed using 8 datasets from the UCI repository, where our RL method is combined with multi-layer perceptrons that serve as function approximators. The RL method is compared to conventional multi-layer perceptrons and support vector machines and the results show that our method slightly outperforms the multi-layer perceptron and performs equally well as the support vector machine. Finally, many possible extensions are described to our basic method, so that much future research can be done to make the proposed method even better. Marco A. Wiering, Hado van Hasselt, Auke-Dirk Pietersma, Lambert Schomaker |
ADPRL | 4 |
| 2011 | Towards robust writer verification by correcting unnatural slant
Axel Brink, Ralph Niels, Roeland A. van Batenburg, C. Elisa van den Heuvel, Lambert Schomaker |
Pattern Recognit. Lett. | 5 |
| 2009 | Recognition of Handwritten Numerical Fields in a Large Single-Writer Historical CollectionabstractThis paper presents a segmentation-based handwriting recognizer and the performance that it achieves on the numerical fields extracted from a large single-writer historical collection. Our recognizer has the particularity that it uses morphing during training: random elastic deformations are applied to fabricate synthetic training character patterns yielding an improved final recognition performance. Two different digit recognizers are evaluated, a multilayer perceptron (MLP) and radial basis function network (RBF), by plugging them into the same left-to-right Viterbi search framework with a tree organization of there cognition lexicon. We also compare with the performance obtained when no dictionary is used to constrain the recognition results. Marius Bulacu, Axel Brink, Tijn van der Zant, Lambert Schomaker |
ICDAR | 4 |
| 2009 | Using Local Symmetry for Landmark Selection
Gert Kootstra, Sjoerd de Jong, Lambert Schomaker |
ICVS | 3 |
| 2009 | Using symmetrical regions of interest to improve visual SLAMabstractSimultaneous Localization and Mapping (SLAM) based on visual information is a challenging problem. One of the main problems with visual SLAM is to find good quality landmarks, that can be detected despite noise and small changes in viewpoint. Many approaches use SIFT interest points as visual landmarks. The problem with the SIFT interest points detector, however, is that it results in a large number of points, of which many are not stable across observations. We propose the use of local symmetry to find regions of interest instead. Symmetry is a stimulus that occurs frequently in everyday environments where our robots operate in, making it useful for SLAM. Furthermore, symmetrical forms are inherently redundant, and can therefore be more robustly detected. By using regions instead of points-of-interest, the landmarks are more stable. To test the performance of our model, we recorded a SLAM database with a mobile robot, and annotated the database by manually adding ground-truth positions. The results show that symmetrical regions-of-interest are less susceptible to noise, are more stable, and above all, result in better SLAM performance. Gert Kootstra, Lambert Schomaker |
IROS | 2 |
| 2008 | How much handwritten text is needed for text-independent writer verification and identificationabstractThe performance of off-line text-independent writer verification and identification increases when the documents contain more text. This relation was examined by repeatedly conducting writer verification and identification performance tests while gradually increasing the amount of text on the pages. The experiment was performed on the datasets Firemaker and IAM using four different features. It was also determined what the influence of an unequal amount of text in the documents is. For the best features, it appears that the minimum amount of needed text is about 100 characters. Axel Brink, Marius Bulacu, Lambert Schomaker |
ICPR | 3 |
| 2008 | Handwritten-Word Spotting Using Biologically Inspired FeaturesabstractFor quick access to new handwritten collections, current handwriting recognition methods are too cumbersome. They cannot deal with the lack of labeled data and would require extensive laboratory training for each individual script, style, language and collection. We propose a biologically inspired whole-word recognition method which is used to incrementally elicit word labels in a live, web-based annotation system, named Monk. Since human labor should be minimized given the massive amount of image data, it becomes important to rely on robust perceptual mechanisms in the machine. Recent computational models of the neuro-physiology of vision are applied to isolated word classification. A primate cortex-like mechanism allows to classify text-images that have a low frequency of occurrence. Typically these images are the most difficult to retrieve and often contain named entities and are regarded as the most important to people. Usually standard pattern-recognition technology cannot deal with these text-images if there are not enough labeled instances. The results of this retrieval system are compared to normalized word-image matching and appear to be very promising. Tijn van der Zant, Lambert Schomaker, Koen V. Haak |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Towards Explainable Writer Verification and Identification Using Vantage WritersabstractIn this paper, a new method for off-line writer verification and identification is proposed which encodes writer features as a mix of typical handwriting styles, written by so-called vantage writers. Since their handwriting can be shown to the user, the method provides a degree of transparency that is usually not present in automatic verification and identification systems. It acts as a dimensionality reduction of a precomputed basic feature vector. In this experiment, the hinge feature was used. The method was tested with unconstrained connected cursive text in two datasets: a known dataset of 252 writers and 1074 writers from a new, forensic dataset. Axel Brink, Lambert Schomaker, Marius Bulacu |
ICDAR | 2 |
| 2007 | Layout Analysis of Handwritten Historical Documents for Searching the Archive of the Cabinet of the Dutch QueenabstractIn this paper, we describe the structure and the performance of a layout analysis system developed for processing the handwritten documents contained in a large historical collection of very high importance in the Netherlands. We introduce a method based on contour tracing that generates curvilinear separation paths between text lines in order to preserve the ascenders and descenders. Our methods are relevant to research on digitization and retrieval of handwritten historical documents. Marius Bulacu, Rutger van Koert, Lambert Schomaker, Tijn van der Zant |
ICDAR | 3 |
| 2007 | Text-Independent Writer Identification and Verification on Offline Arabic HandwritingabstractIn this paper, we evaluate the performance on Arabic handwriting of the text-independent writer identification methods that we developed and tested on Western script in recent years. We use the IFN/ENIT data in the experiments reported here and our tests involve 350 writers. The results show that our methods are very effective and the conclusions drawn in previous studies remain valid also on Arabic script. High performance is achieved by combining textural features (joint directional probability distributions) with allographic features (grapheme-emission distributions). Marius Bulacu, Lambert Schomaker, Axel Brink |
ICDAR | 2 |
| 2007 | Retrieval of Handwritten Lines in Historical DocumentsabstractThis study describes methods for the retrieval of handwritten lines of text in a historical administrative collection. The goal is to develop generic methods for bootstrapping the retrieval system from a tabula rasa starting condition, i.e., the virtual absence of labeled samples. By exploiting the currently available computing power and the fact that computation takes place off line, it should be possible to provide a good starting point for statistical learning methods. In this manner, a closed collection can be incrementally indexed. A cross-correlation method on line-strip images is presented and results are compared to feature-based methods. Lambert Schomaker |
ICDAR | 1 |
| 2007 | Advances in Writer Identification and VerificationabstractThe behavioral-biometrics methods of writer identification and verification are currently enjoying renewed interest, with very promising results. This paper presents a general background and basis for handwriting biometrics. A range of current methods and applications is given. Results on a number of methods are summarized and a more in-depth example of two combined approaches is presented. By combining textural, allographic and placement features, modern systems are starting to display useful performance levels. However, user acceptance will be largely determined by explainability of system results and the integration of system decisions within a (Bayesian) framework of reasoning that is currently becoming forensic practice. Lambert Schomaker |
ICDAR | 1 |
| 2007 | Automatic Allograph Matching in Forensic Writer IdentificationabstractA well-established task in forensic writer identification focuses on the comparison of prototypical character shapes (allographs) present in handwriting. In order for a computer to perform this task convincingly, it should yield results that are plausible and understandable to the human expert. Trajectory matching is a well-known method to compare two allographs. This paper assesses a promising technique for so-called human-congruous trajectory matching, called Dynamic Time Warping (DTW). In the first part of the paper, an experiment is described that shows that DTW yields results that correspond to the expectations of human users. Since DTW requires the dynamics of the handwritten trace, the "online" dynamic allograph trajectories need to be extracted from the "offline" scanned documents. In the second part of the paper, an automatic procedure to perform this task is described. Images were generated from a large online dataset that provides the true trajectories. This allows for a quantitative assessment of the trajectory extraction techniques rather than a qualitative discussion of a small number of examples. Our results show that DTW can significantly improve the results from trajectory extraction when compared to traditional techniques. Ralph Niels, Louis Vuurpijl, Lambert Schomaker |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | Text-Independent Writer Identification and Verification Using Textural and Allographic FeaturesabstractThe identification of a person on the basis of scanned images of handwriting is a useful biometric modality with application in forensic and historic document analysis and constitutes an exemplary study area within the research field of behavioral biometrics. We developed new and very effective techniques for automatic writer identification and verification that use probability distribution functions (PDFs) extracted from the handwriting images to characterize writer individuality. A defining property of our methods is that they are designed to be independent of the textual content of the handwritten samples. Our methods operate at two levels of analysis: the texture level and the character-shape (allograph) level. At the texture level, we use contour-based joint directional PDFs that encode orientation and curvature information to give an intimate characterization of individual handwriting style. In our analysis at the allograph level, the writer is considered to be characterized by a stochastic pattern generator of ink-trace fragments, or graphemes. The PDF of these simple shapes in a given handwriting sample is characteristic for the writer and is computed using a common shape codebook obtained by grapheme clustering. Combining multiple features (directional, grapheme, and run-length PDFs) yields increased writer identification and verification performance. The proposed methods are applicable to free-style handwriting (both cursive and isolated) and have practical feasibility, under the assumption that a few text lines of handwritten material are available in order to obtain reliable probability estimates. Marius Bulacu, Lambert Schomaker |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Using codebooks of fragmented connected-component contours in forensic and historic writer identification
Lambert Schomaker, Katrin Franke, Marius Bulacu |
Pattern Recognit. Lett. | 1 |
| 2006 | Cognitive Developmental Pattern Recognition: Learning to learnabstractIt can be very difficult to manually create software systems which capture the knowledge of an expert. It is an expensive and laborious process that often results in a suboptimal solution. In this paper we propose a new approach which does not require complete manual knowledge construction. The described system is relevant and usable for the end user from the beginning of its development. It is continuously being trained by the experts while they are using it. The study pays attention to the development of the hypotheses that the developing system uses as it adapts to feedback it receives from its own actions on itself, on objects, on the users of the system and the reflection on this feedback. This paper explains how the computer learns from the experts using an example system which performs handwriting recognition. In order to provide a general mechanism, the principle of learning to learn is proposed and applied to the problem of handwriting recognition.The goal is the creation of a system that learns to 'google' through handwritten documents, starting from scratch with a pile of raw images. Tijn van der Zant, Lambert Schomaker, Marco A. Wiering, Axel Brink |
SMC | 2 |
| 2005 | A Comparison of Clustering Methods for Writer Identification and VerificationabstractAn effective method for writer identification and verification is based on assuming that each writer acts as a stochastic generator of ink-trace fragments, or graphemes. The probability distribution of these simple shapes in a given handwriting sample is characteristic for the writer and is computed using a common codebook of graphemes obtained by clustering. In previous studies we used contours to encode the graphemes, in the current paper we explore a complementary shape representation using normalized bitmaps. The most important aim of the current work is to compare three different clustering methods for generating the grapheme codebook: k-means Kohonen SOM 1D and 2D. Large scale computational experiments show that the proposed method is robust to the underlying shape representation used (whether contours or normalized bitmaps), to the size of codebook used (stable performance for sizes from 10/sup 2/ to 2.5 /spl times/ 10/sup 3/) and to the clustering method used to generate the codebook (essentially the same performance was obtained for all three clustering methods). Marius Bulacu, Lambert Schomaker |
ICDAR | 2 |
| 2005 | Improved Text-Detection Methods for a Camera-based Text Reading System for Blind PersonsabstractAutomatic text recognition from natural images receives a growing attention because of potential applications in image retrieval, robotics and intelligent transport system. Camera-based document analysis becomes a real possibility with the increasing resolution and availability of digital cameras. Our research objective is a system that reads the text encountered in natural scenes with the aim to provide assistance to visually impaired persons. In the case of a blind person, finding the text region is the first important problem that must be addressed, because it cannot be assumed that the acquired image contains only characters. In a previous paper (N. Ezaki et al., 2004), we propose four text-detection methods based on connected components. Finding small characters needed significant improvement. This paper describes a new text-detection method geared for small text characters. This method uses Fisher's discriminant rate (FDR) to decide whether an image area should be binarized using local or global thresholds. Fusing the new method with a previous morphology-based one yields improved results. Using a controllable Webcam and a laptop PC, we developed a prototype that works in real time. At first, our system tries to find in the image areas with small characters. Then it zooms into the found areas to retake higher resolution images necessary for character recognition. Going from this proof-of-concept to a complete system requires further research effort. Nobuo Ezaki, Kimiyasu Kiyota, Bui Truong Minh, Marius Bulacu, Lambert Schomaker |
ICDAR | 5 |
| 2004 | Automatic Writer Identification Using Connected-Component Contours and Edge-Based Features of Uppercase Western ScriptabstractIn this paper, a new technique for offline writer identification is presented, using connected-component contours (COCOCOs or CO3s) in uppercase handwritten samples. In our model, the writer is considered to be characterized by a stochastic pattern generator, producing a family of connected components for the uppercase character set. Using a codebook of CO3s from an independent training set of 100 writers, the probability-density function (PDF) of CO3s was computed for an independent test set containing 150 unseen writers. Results revealed a high-sensitivity of the CO3 PDF for identifying individual writers on the basis of a single sentence of uppercase characters. The proposed automatic approach bridges the gap between image-statistics approaches on one end and manually measured allograph features of individual characters on the other end. Combining the CO3 PDF with an independent edge-based orientation and curvature PDF yielded very high correct identification rates. Lambert Schomaker, Marius Bulacu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Writer Style from Oriented Edge Fragments
Marius Bulacu, Lambert Schomaker |
CAIP | 2 |
| 2003 | WANDA: A generic Framework applied in Forensic Handwriting Analysis and Writer Identification
Katrin Franke, Lambert Schomaker, Christian Veenhuis, C. Taubenheim, Isabelle Guyon, Louis Vuurpijl, Merijn van Erp, G. Zwarts |
HIS | 2 |
| 2003 | Writer Identification Using Edge-Based Directional FeaturesabstractThis paper evaluates the performance of edge-based directional probability distributions as features in writer identification in comparison to a number of non-angular features. It is noted that the joint probability distribution of the angle combination of two "hinged" edge fragments outperforms all other individual features. Combining features may improve the performance. Limitations of the method pertain to the amount of handwritten material needed in order to obtain reliable distribution estimates. The global features treated in this study are sensitive to major style variation (upper- vs lower case), slant, and forged styles, which necessitates the use of other features in realistic forensic writer identification procedures. Marius Bulacu, Lambert Schomaker, Louis Vuurpijl |
ICDAR | 2 |
| 2003 | Sparse-parametric writer identification using heterogeneous feature groupsabstractThis paper evaluates the performance of edge-based directional probability distributions as features in writer identification in comparison to a number of nonangular features. It is noted that angular features outperform all other features. However, the nonangular features provide additional valuable information. Rank-combination was used to realize a sparse-parametric combination scheme based on nearest-neighbor search. Limitations of the proposed methods pertain to the amount of handwritten material needed in order to obtain reliable distribution estimates. The global features treated in this study are sensitive to major style variation (upper- vs lower case), slant, and forged styles, which necessitates the use of other features in realistic forensic writer identification procedures. Lambert Schomaker, Marius Bulacu, Merijn van Erp |
ICIP (1) | 1 |
| 2003 | Architectures for detecting and solving conflicts: two-stage classification and support vector classifiers
Louis Vuurpijl, Lambert Schomaker, Merijn van Erp |
Int. J. Document Anal. Recognit. | 2 |
| 1999 | New Use for the Pen: Outline-based Image QueriesabstractA method for image-based queries and search is proposed which is based on the generation of object outlines in images by using the pen, e.g., on color pen computers. By exploiting the actual presence of the human users with their perceptual-motor abilities and by storing textually annotated queries, an incrementally learning image retrieval system can be developed. As an initial test domain, sets of photographs of motor bicycles were used. Classification performances are given for outline and bitmap-derived feature sets, based on nearest-neighbour matching, with promising results. The next step is to use outlines for edge matching in raw, non-annotated images, for which preliminary results are presented. The benefit of the approach will be a user-based multimodal annotation of image database, yielding a gradual improvement in performance over time. Lambert Schomaker, Louis Vuurpijl, Edward de Leau |
ICDAR | 1 |
| 1999 | Supporting Content Retrieval from WWW via "Basic Level Categories" (poster abstract)abstractNo abstract available. Eduard Hoenkamp, Onno Stegeman, Lambert Schomaker |
SIGIR | 3 |
| 1999 | Finding features used in the human reading of cursive handwriting
Lambert Schomaker, Eliane Segers |
Int. J. Document Anal. Recognit. | 1 |
| 1997 | Finding structure in diversity: a hierarchical clustering method for the categorization of allographs in handwritingabstractThe paper introduces a variant of agglomerative hierarchical clustering techniques. The new technique is used for categorizing character shapes (allographs) in large data sets of handwriting into a hierarchical structure. Such a technique may be used as the basis for a systematic naming scheme of character shapes. Problems with existing methods are described and the proposed method is explained. After application of the method to a very large set of characters, separately for all the letters of the alphabet, relevant clusters are identified and given a unique name. Each cluster represents an allograph prototype. Louis Vuurpijl, Lambert Schomaker |
ICDAR | 2 |
| 1994 | UNIPEN project of on-line data exchange and recognizer benchmarksabstractWe report the status of the UNIPEN project of data exchange and recognizer benchmarks started two years ago at the initiative of the International Association of Pattern Recognition (Technical Committee 11). The purpose of the project is to propose and implement solutions to the growing need of handwriting samples for online handwriting recognizers used by pen-based computers. Researchers from several companies and universities have agreed on a data format, a platform of data exchange and a protocol for recognizer benchmarks. The online handwriting data of concern may include handprint and cursive from various alphabets (including Latin and Chinese), signatures and pen gestures. These data will be compiled and distributed by the Linguistic Data Consortium. The benchmarks will be arbitrated the US National Institute of Standards and Technologies. We give a brief introduction to the UNIPEN format. We explain the protocol of data exchange and benchmarks. Isabelle Guyon, Lambert Schomaker, Réjean Plamondon, Mark Y. Liberman, Stan Janet |
ICPR (2) | 2 |
| 1993 | Using stroke- or character-based self-organizing maps in the recognition of on-line, connected cursive script
Lambert Schomaker |
Pattern Recognit. | 1 |