Gernot A. Fink

dblp:98/4099 · DBLP profile ↗
← Back
125ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0002-7446-7813ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 87 · 8 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 31 · 3 first-author · 11 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Recent Advances in Information Extraction from Historical Archival Records
Arthur Matei, Tim Hallyburton, Lukas Hennies, Christoph Rass, Gernot A. Fink
ICDAR (3)5
2026 Writer Retrieval at Scale
Tim Raven, Tim Hallyburton, Gernot A. Fink
ICDAR (3)3
2025 Interpretable Writer Recognition via Vectors of Locally Aggregated Characters
Tim Raven, Vincent Christlein, Gernot A. Fink
ICDAR (4)3
2025 CM1 - A Dataset for Evaluating Few-Shot Information Extraction with Large Vision Language Models
Fabian Wolf, Oliver Tüselmann, Arthur Matei, Lukas Hennies, Christoph Rass, Gernot A. Fink
ICDAR (2)6
2024 Self-supervised Vision Transformers for Writer Retrieval
Tim Raven, Arthur Matei, Gernot A. Fink
ICDAR (2)3
2024 Anonymisation for Time-Series Human Activity Data
Tim Hallyburton, Nilah Ravi Nair, Fernando Moya Rueda, René Grzeszick, Gernot A. Fink
ICPR (15)5
2024 Augmentation of Human Activity Data: Convert, Generate, Transform
Nilah Ravi Nair, Arthur Matei, Dennis Krön, Fernando Moya Rueda, Christopher Reining, Gernot A. Fink
ICPR (10)6
2024 Representation Biases in Time-Series Human Activity Recognition with Small Sample Sizes
Nilah Ravi Nair, Lena Schmid, Christopher Reining, Fernando Moya Rueda, Markus Pauly, Gernot A. Fink
ICPR (15)6
2024 Neural models for semantic analysis of handwritten document images
abstract
Abstract Semantic analysis of handwritten document images offers a wide range of practical application scenarios. A sequential combination of handwritten text recognition (HTR) and a task-specific natural language processing system offers an intuitive solution in this domain. However, this HTR-based approach suffers from the problem of error propagation. An HTR-free model, which avoids explicit text recognition and solves the task end-to-end, tackles this problem, but often produces poor results. A possible reason for this is that it does not incorporate largely pre-trained semantic word embeddings, which turn out to be one of the most powerful advantages in the textual domain. In this work, we propose an HTR-based and an HTR-free model and compare them on a variety of segmentation-based handwritten document image benchmarks including semantic word spotting, named entity recognition, and question answering. Furthermore, we propose a cross-modal knowledge distillation approach to integrate semantic knowledge from textually pre-trained word embeddings into HTR-free models. In a series of experiments, we investigate optimization strategies for robust semantic word image representation. We show that the incorporation of semantic knowledge is beneficial for HTR-free approaches in achieving state-of-the-art results on a variety of benchmarks.
Oliver Tüselmann, Gernot A. Fink
Int. J. Document Anal. Recognit.2
2024 Self-training for handwritten word recognition and retrieval
abstract
Abstract Handwritten text recognition and Word Retrieval, also known as Word Spotting, are traditional problems in the document analysis community. While the use of increasingly large neural network architectures has led to a steady improvement of performances it comes with the drawback of requiring manually annotated training data. This poses a tremendous problem considering their application to new document collections. To overcome this drawback, we propose a self-training approach that allows to train state-of-the-art models for HTR and word spotting. Self-training is a common technique in semi-supervised learning and usually relies on a small labeled dataset and training on pseudo-labels generated by an initial model. In this work, we show that it is feasible to train models on synthetic data that are sufficiently performant to serve as initial models for self-training. Therefore, the proposed training method does not rely on any manually annotated samples. We further investigate visual and language properties of the synthetic datasets. In order to improve performance and robustness of the self-training approach, we propose different confidence measures for both models that allow to identify and remove erroneous pseudo-labels. The presented training approach clearly outperforms other learning-free methods or adaptation strategies under the absence of manually annotated data.
Fabian Wolf, Gernot A. Fink
Int. J. Document Anal. Recognit.2
2023 Exploring Semantic Word Representations for Recognition-Free NLP on Handwritten Document Images
Oliver Tüselmann, Gernot A. Fink
ICDAR (4)2
2023 Editorial for special issue on "advanced topics in document analysis and recognition"
Koichi Kise, Richard Zanibbi, Rajiv Jain, Gernot A. Fink
Int. J. Document Anal. Recognit.4
2022 Named Entity Linking on Handwritten Document Images
Oliver Tüselmann, Gernot A. Fink
DAS2
2022 A Weighted Combination of Semantic and Syntactic Word Image Representations
Oliver Tüselmann, Kai Brandenbusch, Gernot A. Fink
ICFHR4
2022 Recognition-Free Question Answering on Handwritten Document Collections
Oliver Tüselmann, Friedrich Müller, Fabian Wolf, Gernot A. Fink
ICFHR4
2022 Combining Self-training and Minimal Annotations for Handwritten Word Recognition
Fabian Wolf, Gernot A. Fink
ICFHR2
2022 Video-based Pose-Estimation Data as Source for Transfer Learning in Human Activity Recognition
abstract
Human Activity Recognition (HAR) using on-body devices identifies specific human actions in unconstrained environments. HAR is challenging due to the inter and intra-variance of human movements; moreover, annotated datasets from on-body devices are scarce. This problem is mainly due to the difficulty of data creation, i.e., recording, expensive annotation, and lack of standard definitions of human activities. Previous works demonstrated that transfer learning is a good strategy for addressing scenarios with scarce data. However, the scarcity of annotated on-body device datasets remains. This paper proposes using datasets intended for human-pose estimation as a source for transfer learning; specifically, it deploys sequences of annotated pixel coordinates of human joints from video datasets for HAR and human pose estimation. We pre-train a deep architecture on four benchmark video-based source datasets. Finally, an evaluation is carried out on three on-body device datasets improving HAR performance.
Shrutarv Awasthi, Fernando Moya Rueda, Gernot A. Fink
ICPR3
2022 Image Augmentations in Planetary Science: Implications in Self-Supervised Learning and Weakly-Supervised Segmentation on Mars
abstract
Research on the use of augmentations in physically constrained remote sensing scenarios, like the analysis of Martian surface data, is largely unexplored. In this work we present an analysis on how reasonable augmentation strategies can be selected which are class agnostic and respect physical plausibility in supervised and weakly-supervised tasks. Additionally, we present the first results of self-supervised learning on Martian surface data, discuss the importance of physically plausible augmentations in the context of self-supervised learning, specifically contrastive learning, and provide a comprehensive overview of the generalization properties induced by different augmentation strategies with the help of geomorphic maps.
Dominik Koßmann, Arthur Matei, Thorsten Wilhelm, Gernot A. Fink
ICPR4
2022 Self-Training of Handwritten Word Recognition for Synthetic-to-Real Adaptation
abstract
Performances of Handwritten Text Recognition (HTR) models are largely determined by the availability of labeled and representative training samples. However, in many application scenarios, labeled samples are scarce or costly to obtain. In this work, we propose a self-training approach to train a HTR model solely on synthetic samples and unlabeled data. The proposed training scheme uses an initial model trained on synthetic data to make predictions for the unlabeled target dataset. Starting from this initial model with rather poor performance, we show that a considerable adaptation is possible by training against the iteratively predicted pseudo-labels. Therefore, the investigated self-training method does not require any manually annotated training samples. We evaluate the proposed method on four benchmark datasets and show its effectiveness on reducing the gap to a model trained in a fully-supervised manner.
Fabian Wolf, Gernot A. Fink
ICPR2
2022 UQGAN: A Unified Model for Uncertainty Quantification of Deep Classifiers trained via Conditional GANs
abstract
We present an approach to quantifying both aleatoric and epistemic uncertainty for deep neural networks in image classification, based on generative adversarial networks (GANs). While most works in the literature that use GANs to generate out-of-distribution (OoD) examples only focus on the evaluation of OoD detection, we present a GAN based approach to learn a classifier that produces proper uncertainties for OoD examples as well as for false positives (FPs). Instead of shielding the entire in-distribution data with GAN generated OoD examples which is state-of-the-art, we shield each class separately with out-of-class examples generated by a conditional GAN and complement this with a one-vs-all image classifier. In our experiments, in particular on CIFAR10, CIFAR100 and Tiny ImageNet, we improve over the OoD detection and FP detection performance of state-of-the-art GAN-training based classifiers. Furthermore, we also find that the generated GAN examples do not significantly affect the calibration error of our classifier and result in a significant gain in model accuracy.
Philipp Oberdiek, Gernot A. Fink, Matthias Rottmann
NeurIPS2
2021 Context Aware Generation of Cuneiform Signs
Kai Brandenbusch, Eugen Rusakov, Gernot A. Fink
ICDAR (1)3
2021 Embedded Attributes for Cuneiform Sign Spotting
Eugen Rusakov, Turna Somel, Gerfrid G. W. Müller, Gernot A. Fink
ICDAR (2)4
2021 Are End-to-End Systems Really Necessary for NER on Handwritten Document Images?
Oliver Tüselmann, Fabian Wolf, Gernot A. Fink
ICDAR (2)3
2021 Graph Convolutional Neural Networks for Learning Attribute Representations for Word Spotting
Fabian Wolf, Andreas Fischer 0002, Gernot A. Fink
ICDAR (1)3
2021 Generation of Attributes for Highly Imbalanced Land Cover Data
abstract
Through the rise of new remote sensing datasets with sufficient size for current deep learning models, land cover classification results have improved in recent years. Unfortunately, earth exhibits a natural imbalance in different land cover classes, also visible in these datasets. In the domain of zero-shot learning, image attributes have enabled better results in transfer learning by creation of a mid level representation between different classes. This representation can also be constructed for land cover data and used to detect minority classes through shared features. We propose a way for generation of attributes by text mining for one of the biggest land cover datasets. With these attributes we achieve state-of-the-art performance in land cover classification and improve results especially for minority classes. Further, we show that these attributes have great potential in weakly supervised land cover segmentation.
Dominik Koßmann, Thorsten Wilhelm, Gernot A. Fink
IGARSS3
2021 Guest Editorial: Special Issue: Computer Vision and Pattern Recognition (DAGM GCPR 2019)
Simone Frintrop, Gernot A. Fink, Xiaoyi Jiang 0001
Int. J. Comput. Vis.2
2021 Annotation-Free Word Spotting with Bag-of-Features HMMs
abstract
The annotation-free word spotting method that is proposed in this paper makes document images searchable without requiring any labeled training data. Thus, our method supports the exploration of a document collection directly without demanding any manual efforts from the users for the preparation of a training dataset. Our method works in the query-by-example scenario where the user selects an exemplary occurrence of the query word. Afterwards, the entire collection of document images is searched according to visual similarity to the query. The proposed method requires only minimal assumptions about the visual appearance of text. This is achieved by processing document images as a whole without requiring a given segmentation of the images on word level or on line level. Therefore, the method is also segmentation-free. Word size variabilities can be handled by representing the sequential structure of text with a statistical sequence model. In order to make the computationally costly application of the sequence model feasible in practice, regions are retrieved according to approximate similarity with an efficient model decoding algorithm. Re-ranking these regions according to the visual similarity obtained with the sequence model leads to highly accurate word spotting results. The method is evaluated on five benchmark datasets. In the segmentation-free query-by-example scenario where no annotated training data is available, the method outperforms all other methods that have been evaluated on any of these five benchmarks.
Leonard Rothacker, Fabian Wolf, Gernot A. Fink
Int. J. Pattern Recognit. Artif. Intell.3
2020 Annotation-Free Learning of Deep Representations for Word Spotting Using Synthetic Data and Self Labeling
Fabian Wolf, Gernot A. Fink
DAS2
2020 Towards Query-by-eXpression Retrieval of Cuneiform Signs
abstract
Intermediate representations based on binary attributes are widely used for image retrieval tasks. Especially for the task of word spotting, attribute representations continually achieves state-of-the-art results. In contrast to Latin scripts, the cuneiform writing system is still a less well-known domain for the document analysis community. This writing system consists of wedge impressions, with the constellations and relative positions of wedges representing sign classes. While for Latin scripts characters are intuitive modeled as visual attributes, cuneiform signs do not reveal an evident approach to model visual attributes for cuneiform signs. However, in this work we introduce an attribute representation based on the so-called Gottstein-System. The idea is decompose signs according to wedge typology and enable a logical expression for sign classes, while sharing these expressions across visually similar cuneiform signs. We adapt this idea of describing wedge expressions in order to introduce to the Query-by-eXpression (QbX) retrieval scenario. Compared to queries based on sign IDs, our approach is capable of representing queries for an open-set retrieval scenario. Furthermore, we extend the Gottstein-System by expressions representing visual hints like wedge crossings or position relations. This way, promising results are achieved, as shown in our experiments.
Eugen Rusakov, Turna Somel, Gernot A. Fink, Gerfrid G. W. Müller
ICFHR3
2020 Identifying and Tackling Key Challenges in Semantic Word Spotting
abstract
Semantic word spotting is an extension of the traditional word spotting approach that uses not only visual but also semantic information to determine the similarity between a word image and a given query. Current approaches in this area achieve a semantic retrieval by embedding word images into a textually trained semantic space. The related literature presents remarkable results regarding established metrics indicating that the task of semantic word image retrieval is solved. A closer look at the results reveals, however, that this is only partially the case. In this work, we identify and solve current key challenges for semantic word spotting. We analyze the published works in this field towards these challenges and show why they do not solve them. For this purpose, we demonstrate that the used embedding space from current methods contains strong artifacts influencing the retrieval task. Furthermore, we evaluate a more suitable and established embedding approach from Natural Language Processing for semantic word spotting. We also explain the challenges of mapping word images into a semantic embedding space and evaluate different architectures for this task. Thereby, we present a new architecture that outperforms current approaches in this area. In addition, we show that commonly used metrics are not suitable for evaluating a semantic retrieval and present a new evaluation metric for this task.
Oliver Tüselmann, Fabian Wolf, Gernot A. Fink
ICFHR3
2020 Improving Handwritten Word Synthesis for Annotation-free Word Spotting
abstract
Annotation-free word spotting aims at retrieving relevant word images from a document collection without the need of a manually labeled training dataset. As annotated data is usually scarce in the application scenarios of a word spotting system, transfer learning and annotation-free methods became increasingly popular. One possibility to alleviate the annotation problem is to train on synthetically generated word images. Therefore, a common approach is to render word images from electronic fonts and to vary the synthesis parameters randomly. In this work, we show that an annotation-free word spotting method benefits from an adapted synthesis procedure. We investigate the influence of the choice of the underlying vocabulary and the combination of synthesis and data augmentation. Furthermore, we present a method to adapt the style of the synthesized word images to the target dataset. We evaluate the proposed changes to the synthesis procedure on three benchmark datasets and improve performances considerably.
Fabian Wolf, Kai Brandenbusch, Gernot A. Fink
ICFHR3
2020 Towards Tackling Multi-Label Imbalances in Remote Sensing Imagery
abstract
Recent advances in automated image analysis have lead to an increased number of proposed datasets in remote sensing applications. This permits the successful employment of data hungry state-of-the-art deep neural networks. However, the Earth is not covered equally by semantically meaningful classes. Thus, many land cover datasets suffer from a severe class imbalance. We show that by taking appropriate measures, the performance in the minority classes can be improved by up to 20 percent without affecting the performance in the majority classes strongly. Additionally, we investigate the use of an attribute encoding scheme to represent the inherent class hierarchies commonly observed in land cover analysis.
Dominik Koßmann, Thorsten Wilhelm, Gernot A. Fink
ICPR3
2020 From Human Pose to On-Body Devices for Human-Activity Recognition
abstract
Human Activity Recognition (HAR), using inertial measurements from on-body devices, has not seen a great advantage from deep architectures. This drawback is mainly due to the lack of annotated data, diversity of on-body device configurations, the class-unbalance problem, and non-standard human activity definitions. Approaches for improving the performance of such architectures, e.g., transfer learning, are therefore difficult to apply. This paper introduces a method for transfer learning from human-pose estimations as a source for improving HAR using inertial measurements obtained from on-body devices. We propose to fine-tune deep architectures, trained using sequences of human poses from a large dataset and their derivatives, for solving HAR on inertial measurements from on-body devices. Derivatives of human poses will be considered as a sort of synthetic data for HAR. We deploy two different temporal-convolutional architectures as classifiers. An evaluation of the method is carried out on three benchmark datasets improving the classification performance.
Fernando Moya Rueda, Gernot A. Fink
ICPR2
2019 Exploring Confidence Measures for Word Spotting in Heterogeneous Datasets
abstract
In recent years, convolutional neural networks (CNNs) took over the field of document analysis and they became the predominant model for word spotting. Especially attribute CNNs, which learn the mapping between a word image and an attribute representation, showed exceptional performances. The drawback of this approach is the overconfidence of neural networks when used out of their training distribution. In this paper, we explore different metrics for quantifying the confidence of a CNN in its predictions, specifically on the retrieval problem of word spotting. With these confidence measures, we limit the inability of a retrieval list to reject certain candidates. We investigate four different approaches that are either based on the network's attribute estimations or make use of a surrogate model. Our approach also aims at answering the question for which part of a dataset the retrieval system gives reliable results. We further show that there exists a direct relation between the proposed confidence measures and the quality of an estimated attribute representation.
Fabian Wolf, Philipp Oberdiek, Gernot A. Fink
ICDAR3
2019 Handwritten Arabic text recognition using multi-stage sub-core-shape HMMs
Irfan Ahmad 0001, Gernot A. Fink
Int. J. Document Anal. Recognit.2
2018 Learning Deep Representations for Word Spotting under Weak Supervision
abstract
Convolutional Neural Networks have made their mark in various fields of computer vision in recent years. They have achieved state-of-the-art performance in the field of document analysis as well. However, CNNs require a large amount of annotated training data and, hence, great manual effort. In our approach, we introduce a method to drastically reduce the manual annotation effort while retaining the high performance of a CNN for word spotting in handwritten documents. The model is learned with weak supervision using a combination of synthetically generated training data and a small subset of the training partition of the handwritten data set. We show that the network achieves results highly competitive to the state-of-the-art in word spotting with shorter training times and a fraction of the annotation effort.
Neha Gurjar, Sebastian Sudholt, Gernot A. Fink
DAS3
2018 Towards a Framework for Semi-Automated Annotation of Human Order Picking Activities Using Motion Capturing
abstract
Data creation for Human Activity Recognition (HAR) requires an immense human effort and contextual knowledge for manual annotation.This paper proposes a framework for semi-automated annotation of sequential data in the order picking process using a motion capturing system.Additionally, it introduces proper annotation labels by defining process steps, human activities and simple human movements in order picking scenarios.An attribute representation based on simple human movements meets the challenges set by the versatility of activities in warehousing.
Christopher Reining, Fernando Moya Rueda, Michael ten Hompel, Gernot A. Fink
FedCSIS4
2018 A Probabilistic Retrieval Model for Word Spotting Based on Direct Attribute Prediction
abstract
In recent years CNNs took over in various fields of computer vision. Adapted to document image analysis, they achieved state-of-the-art performance in word spotting by predicting word string embeddings. One prominent embedding splits a given string in temporal pyramidal regions of character occurrences, namely the Pyramidal Histogram of Characters (PHOC). This string embedding can be interpreted as a binary attribute representation. In this work we present a new approach for ranking retrieval lists originally proposed for zero-shot learning where attribute representations play an important role. Instead of a distance-based matching of the predicted string embedding, we compute the posterior probability of the attribute representation given a word image which can be interpreted as a posterior of the query. We can show that this probabilistic ranking improves word spotting performance, especially in the query-by-string scenario.
Eugen Rusakov, Leonard Rothacker, Hyunho Mo, Gernot A. Fink
ICFHR4
2018 Learning Attribute Representation for Human Activity Recognition
abstract
Attribute representations became relevant in image recognition and word spotting, providing support under the presence of unbalance and disjoint datasets. However, for human activity recognition using sequential data from on-body sensors, human-labeled attributes are lacking. This paper introduces a search for attributes that represent favorably signal segments for recognizing human activities. It presents three deep architectures, including temporal-convolutions and an IMU centered design, for predicting attributes. An empiric evaluation of random and learned attribute representations, and as well as the networks is carried out on two datasets, outperforming the state-of-the art.
Fernando Moya Rueda, Gernot A. Fink
ICPR2
2018 Special issue on deep learning for document analysis and recognition
Cheng-Lin Liu 0001, Gernot A. Fink, Venu Govindaraju
Int. J. Document Anal. Recognit.2
2018 Attribute CNNs for word spotting in handwritten documents
Sebastian Sudholt, Gernot A. Fink
Int. J. Document Anal. Recognit.2
2017 Word Hypotheses for Segmentation-Free Word Spotting in Historic Document Images
abstract
The generation of word hypotheses for segmentation-free word spotting on document level is usually subject to heuristic expert design. This involves strong assumptions about the visual appearance of text in the document images. In this paper we propose to generate hypotheses with text detectors. In order to do so, we present three detectors that are based on SIFT contrast scores, CNN region classification scores and attribute activation maps. The uncertainty in the detector scores is modeled with the extremal regions method. Retrieving word hypotheses is based on PHOC representations which we compute with the TPP-PHOCNet. We evaluate our method on the George Washington dataset and the ICFHR 2016 KWS competition benchmarks. In the evaluation we show that high word detection rates can be achieved. This is a prerequisite for high retrieval performance that is competitive with the state-of-the-art.
Leonard Rothacker, Sebastian Sudholt, Eugen Rusakov, Matthias Kasperidus, Gernot A. Fink
ICDAR5
2017 Evaluating Word String Embeddings and Loss Functions for CNN-Based Word Spotting
abstract
The recent past has seen CNNs take over the field of word spotting. The dominance of these neural networks is fueled by learning to predict a word string embedding for a given input image. While the PHOC (Pyramidal Histogram of Characters) is most prominently used, other embeddings such as the Discrete Cosine Transform of Words have been used as well. In this work, we investigate the use of different word string embeddings for word spotting. For this, we make use of the recently proposed PHOCNet and modify it to be able to not only learn binary representations. Our extensive evaluation shows that a large number of combinations of word string embeddings and loss functions achieve roughly the same results on different word spotting benchmarks. This leads us to the conclusion that no word string embedding is really superior to another and new embeddings should focus on incorporating more information than only character counts and positions.
Sebastian Sudholt, Gernot A. Fink
ICDAR2
2017 Query-by-Online Word Spotting Revisited: Using CNNs for Cross-Domain Retrieval
abstract
A word spotting system is in large parts characterized by the query modalities it is able to process. The most common modalities here are Query-by-Example and Query-by-String. However, recently a new query type has been proposed: In Query-by-Online-Trajectory (QbO) the query is presented as a set of online-handwritten trajectories. In this work we devise a cross-domain word spotting framework using CNNs which is able to accomplish the QbO task. In particular, we design two different QbO systems which we evaluate in a number of experiments. We are not only able to outperform the current state of the art in QbO word spotting but also show that a system using a single CNN for both online and offline data achieves superior results compared to a system that uses a CNN for each domain individually.
Sebastian Sudholt, Leonard Rothacker, Gernot A. Fink
ICDAR3
2017 Weakly-supervised localization of diabetic retinopathy lesions in retinal fundus images
abstract
Convolutional neural networks (CNNs) show impressive performance for image classification and detection, extending heavily to the medical image domain. Nevertheless, medical experts are skeptical in these predictions as the nonlinear multilayer structure resulting in a classification outcome is not directly graspable. Recently, approaches have been shown which help the user to understand the discriminative regions within an image which are decisive for the CNN to conclude to a certain class. Although these approaches could help to build trust in the CNNs predictions, they are only slightly shown to work with medical image data which often poses a challenge as the decision for a class relies on different lesion areas scattered around the entire image. Using the DiaretDB1 dataset, we show that on retina images different lesion areas fundamental for diabetic retinopathy are detected on an image level with high accuracy, comparable or exceeding supervised methods. On lesion level, we achieve few false positives with high sensitivity, though, the network is solely trained on image-level labels which do not include information about existing lesions. Classifying between diseased and healthy images, we achieve an AUC of 0.954 on the DiaretDB1.
Muhammad Waleed Gondal, Jan M. Köhler, René Grzeszick, Gernot A. Fink, Michael Hirsch 0001
ICIP4
2017 Optimistic and pessimistic neural networks for object recognition
abstract
In this paper the application of uncertainty modeling to convolutional neural networks is evaluated. A novel method for adjusting the network's predictions based on uncertainty information is introduced. This allows the network to be either optimistic or pessimistic in it's prediction scores. The proposed method builds on the idea of applying dropout at test time and sampling a predictive mean and variance from the network's output. Besides the methodological aspects, implementation details allowing for a fast evaluation are presented. In the evaluation on the ILSVRC2014 and VOC2011 datasets it will be shown that modeling uncertainty allows for improving the performance of a given model purely at test time without any further training steps.
René Grzeszick, Sebastian Sudholt, Gernot A. Fink
ICIP3
2017 Passive Online Geometry Calibration of Acoustic Sensor Networks
abstract
As we are surrounded by an increased number of mobile devices equipped with wireless links and multiple microphones, e.g., smartphones, tablets, laptops, and hearing aids, using them collaboratively for acoustic processing is a promising platform for emerging applications. These devices make up an acoustic sensor network comprised of nodes, i.e., distributed devices equipped with microphone arrays, communication unit, and processing unit. Algorithms for speaker separation and localization using such a network require a precise knowledge of the nodes' locations and orientations. To acquire this knowledge, a recently introduced approach proposed a combined direction of arrival and time difference of arrival (TDoA) target function for offline calibration with dedicated recordings. This letter proposes an extension of this approach to a novel online method with two new features: First, by employing an evolutionary algorithm on incremental measurements, it is online and fast enough for real-time application. Second, by using the sparse spike representation computed in a cochlear model for TDoA estimation, the amount of information shared between the nodes by transmission is reduced, while the accuracy is increased. The proposed approach is able to calibrate an acoustic senor network online during a meeting in a reverberant conference room.
Axel Plinge, Gernot A. Fink, Sharon Gannot
IEEE Signal Process. Lett.2
2017 Bag-of-Features Methods for Acoustic Event Detection and Classification
abstract
The detection and classification of acoustic events in various environments is an important task. Its applications range from multimedia analysis to surveillance of humans or even animal life. Several of these tasks require the capability of online processing. Besides many approaches that tackle the task of acoustic event detection, methods that are based on the well known bag-of-features principle also emerged into the field. Acoustic features are calculated for all frames in a given time window. Then, applying the bag-of-features concept, these features are quantized with respect to a learned codebook and a histogram representation is computed. Bag-of-features approaches are particularly interesting for online processing as they have a low computational cost. In this paper, the bag-of-features principle and various extensions are reviewed, including soft quantization, supervised codebook learning, and temporal modeling. Furthermore, Mel and Gammatone frequency cepstral coefficients that originate from psychoacoustic models are used as the underlying feature set for the bag-of-features. The possibility of fusing the results of multiple channels in order to improve the robustness is shown. Two databases are used for the experiments: The DCASE 2013 office live dataset and the ITC-IRST multichannel dataset.
René Grzeszick, Axel Plinge, Gernot A. Fink
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Word Spotting in Historical Document Collections with Online-Handwritten Queries
abstract
Pen-based systems are becoming more and more important due to the growing availability of touch sensitive devices in various forms and sizes. Their interfaces offer the possibility to directly interact with a system by natural handwriting. In contrast to other input modalities it is not required to switch to special modes, like software-keyboards. In this paper we propose a new method for querying digital archives of historical documents. Word images are retrieved with respect to search terms that users write on a pen-based system by hand. The captured trajectory is used as a query which we call query-by-online-trajectory word spotting. By using attribute embeddings for both online-trajectory and visual features, word images are retrieved based on their distance to the query in a common subspace. The system is therefore robust, as no explicit transcription for queries or word images is required. We evaluate our approach for writer-dependent as well as writer-independent scenarios, where we present highly accurate retrieval results in the former and compelling retrieval results in the latter case. Our performance is very competitive in comparison to related methods from the literature.
Christian Wieprecht, Leonard Rothacker, Gernot A. Fink
DAS3
2016 Class-Based Contextual Modeling for Handwritten Arabic Text Recognition
abstract
In this paper we will present our investigations related to contextual modeling for HMM-based handwritten Arabic text recognition. We will, first, discuss the justifications and the need for contextual modeling for handwritten Arabic text recognition. Next, we will discuss the issues related to contextual modeling for Arabic text recognition. Finally, we will present our novel class-based contextual modeling for HMM-based handwritten Arabic text recognition. Experiment results on word recognition tasks show improvements in word recognition rates when compared to using standard contextual HMMs. Moreover, the recognizers are significantly more compact as compared to the standard contextual HMM systems.
Irfan Ahmad 0001, Gernot A. Fink
ICFHR2
2016 Robust Output Modeling in Bag-of-Features HMMs for Handwriting Recognition
abstract
Bag-of-Features HMMs have been successfully applied to handwriting recognition and word spotting. In this paper we extend our previous work and present methods for modeling sequences of Bag-of-Features representations with Hidden Markov Models. We will discuss our previous approach that uses a pseudo-discrete model. Afterwards, we present a novel semi-continuous integration. The method is effective for probabilistic text clustering and is suitable for statistically modeling the characteristics of Bag-of-Features representations extracted from document images. Furthermore, its statistical expectation-maximization estimation can directly be integrated in Baum-Welch HMM training. In our experiments we present competitive results on the IfN/ENIT word recognition benchmark and state-of-the-art results for word spotting on the George Washington benchmark. Our evaluation gives insights into the properties of the models from the perspectives of modern as well as historic document analysis.
Leonard Rothacker, Gernot A. Fink
ICFHR2
2016 PHOCNet: A Deep Convolutional Neural Network for Word Spotting in Handwritten Documents
abstract
In recent years, deep convolutional neural networks have achieved state of the art performance in various computer vision tasks such as classification, detection or segmentation. Due to their outstanding performance, CNNs are more and more used in the field of document image analysis as well. In this work, we present a CNN architecture that is trained with the recently proposed PHOC representation. We show empirically that our CNN architecture is able to outperform state-of-the-art results for various word spotting benchmarks while exhibiting short training and test times.
Sebastian Sudholt, Gernot A. Fink
ICFHR2
2016 Motion Classification for Analyzing the Order Picking Process using Mobile Sensors - General Concepts, Case Studies and Empirical Evaluation
abstract
This contribution introduces a new concept to analyze the manual order picking process which is a key task in the field of logistics. The approach relies on a sensor-based motion classification already used in other domains like sports or medical science. Thereby, different sensor data, e. g. acceleration or rotation rate, are continuously recorded during the order picking process. With help of this data, the process can be analyzed to identify different motion classes, like walking or picking, and the time a subject spends in each class. Moreover, relevant motion classes within the order picking process are defined which were identified during field studies in two different companies. These classes are recognized by a classification system working with methods from the field of statistical pattern recognition. The classification is done with a supervised learning approach for which promising results can be shown.
Sascha Feldhorst, Mojtaba Masoudinejad, Michael ten Hompel, Gernot A. Fink
ICPRAM4
2016 An Iterative Partitioning-Based Method for Semi-Supervised Annotation Learning in Image Collections
abstract
Labeling images is tedious and a costly work that is required for many applications, for example, tagging, grouping and exploring of image collections. It is also necessary for training visual classifiers that recognize scenes or objects. It is therefore desirable to either reduce the human effort or infer additional knowledge by addressing this task with algorithms that allow for learning image annotations in a semi-supervised manner. In this paper, a semi-supervised annotation learning algorithm is introduced that is based on partitioning the data in a multi-view approach. The method is applied to large, diverse image collections of natural scene images. Experiments are performed on the 15 Scenes and SUN databases. It is shown that for sparsely labeled datasets the proposed annotation learning algorithm is able to infer additional knowledge from the unlabeled samples and therefore improve the performance of visual classifiers in comparison to supervised learning. Furthermore, the proposed algorithm outperforms other related semi-supervised learning approaches.
René Grzeszick, Gernot A. Fink
Int. J. Pattern Recognit. Artif. Intell.2
2016 Open-vocabulary recognition of machine-printed Arabic text using hidden Markov models
Irfan Ahmad 0001, Sabri A. Mahmoud, Gernot A. Fink
Pattern Recognit.3
2015 Training an Arabic handwriting recognizer without a handwritten training data set
abstract
Handwritten text recognition is an active research area in pattern recognition. One of the prerequisites of setting up a handwritten text recognizer is to train them using, mostly, large amounts of labeled training data. In the current paper we report our work on handwritten text recognition using no handwritten training set. We investigate different approaches including, computer generated text in different typefaces as training data, unsupervised adaptation, and using recognition hypothesis on the test sets as training data. Results from handwritten Arabic word recognition task show that the approach is promising with good recognition rates.
Irfan Ahmad 0001, Gernot A. Fink
ICDAR2
2015 Multi-stage HMM based Arabic text recognition with rescoring
abstract
In this paper, we present a multi-stage approach to handwritten Arabic text recognition using HMM where we separate the Arabic text image into core components and diacritics and recognize them separately using two separate HMM recognition systems. In the next stage, we combine the scores from both recognizers to make a final word hypothesis. This approach leads to huge reduction in the number of HMM models that need to be trained. Experiments conducted on a word recognition task using a publicly available benchmark database show the effectiveness of the technique. We achieve state-of-the-art results in addition to a compact model set for the recognition system.
Irfan Ahmad 0001, Gernot A. Fink
ICDAR2
2015 Segmentation-free query-by-string word spotting with Bag-of-Features HMMs
abstract
Word spotting allows to explore document images without requiring a full transcription. In the query-by-string scenario considered in this paper, it is possible to search arbitrary keywords while only limited prior information about the documents is required. We learn context-dependent character models from a training set that is small with respect to the number of models. This is possible due to the use of Bag-of-Features HMMs that are especially suited for estimating robust models from limited training material. In contrast to most query-by-string methods we consider a fully segmentation-free decoding framework that does not require any pre-segmentation on word or line level. Experiments on the well-known George Washington benchmark demonstrate the high accuracy of our method.
Leonard Rothacker, Gernot A. Fink
ICDAR2
2015 Learning local image descriptors for word spotting
abstract
The Bag-of-Features paradigm has enjoyed great success in computer vision as well as document image analysis applications. By far the most common approach here is to power the Bag-of-Features pipeline with SIFT descriptors which are then clustered into a visual vocabulary using Lloyd's algorithm. In contrast to using handcrafted descriptors, many researches have started to use descriptors that have been learned from data. While descriptor learning is common in other computer vision tasks, there has been little work on learning descriptors for document analysis purposes. In this work we propose a descriptor learning pipeline designed for word spotting. Evaluation results on the well known George Washington database demonstrate that word-spotting results can effectively be improved by learning specialized local image descriptors.
Sebastian Sudholt, Leonard Rothacker, Gernot A. Fink
ICDAR3
2015 Camera-Based Whiteboard Reading for Understanding Mind Maps
abstract
Mind maps, i.e. the spatial organization of ideas and concepts around a central topic and the visualization of their relations, represent a very powerful and thus popular means to support creative thinking and problem solving processes. Typically created on traditional whiteboards, they represent an important technique for collaborative brainstorming sessions. We describe a camera-based system to analyze hand-drawn mind maps written on a whiteboard. The goal of the presented system is to produce digital representations of such mind maps, which would enable digital asset management, i.e. storage and retrieval of manually created documents. Our system is based on image acquisition by means of a camera, followed by the segmentation of the particular whiteboard image focusing on the extraction of written context, i.e. the ideas captured by the mind map. The spatial arrangement of these ideas is recovered using layout analysis based on unsupervised clustering, which results in graph representations of mind maps. Finally, handwriting recognition derives textual transcripts of the ideas captured by the mind map. We demonstrate the capabilities of our mind map reading system by means of an experimental evaluation, where we analyze images of mind maps that have been drawn on whiteboards, without any further constraints other than the underlying topic. In addition to the promising recognition results, we also discuss training strategies, which effectively allow for system bootstrapping using out-of-domain sample data. The latter is important when addressing creative thinking processes where domain-related training data are difficult to obtain as they focus on novelty by definition.
Szilárd Vajda, Thomas Plötz, Gernot A. Fink
Int. J. Pattern Recognit. Artif. Intell.3
2014 Multi-speaker tracking using multiple distributed microphone arrays
abstract
Tracking multiple speakers with microphone arrays is one of the key tasks in smart environments. For good accuracy in reverberant environments, several arrays should be distributed in the room. The method presented is using distributed nodes with microphone arrays that compute local angular speech detections. In an integrating node, these are associated using the spectra and tracks for multiple concurrent speaker are computed. Euclidean coordinates are derived by triangulation, which is improved by a quality based weighting. The method is not only robust against reverberation, but also against transmission errors and jitter. Test with real recordings show that good precision for practical applications can be achieved.
Axel Plinge, Gernot A. Fink
ICASSP2
2014 A Bag-of-Features approach to acoustic event detection
abstract
The classification of acoustic events in indoor environments is an important task for many practical applications in smart environments. In this paper a novel approach for classifying acoustic events that is based on a Bag-of-Features approach is proposed. Mel and gammatone frequency cepstral coefficients that originate from psychoacoustic models are used as input features for the Bag-of representation. Rather than using a prior classification or segmentation step to eliminate silence and background noise, Bag-of-Features representations are learned for a background class. Supervised learning of codebooks and temporal coding are shown to improve the recognition rates. Three different databases are used for the experiments: the CLEAR sound event dataset, the D-CASE event dataset and a new set of smart room recordings.
Axel Plinge, René Grzeszick, Gernot A. Fink
ICASSP3
2014 Improvements in Sub-character HMM Model Based Arabic Text Recognition
abstract
Sub-character HMM models for Arabic text recognition allow sharing of common patterns between different position-dependent shape forms of an Arabic character as well as between different characters. The number of HMMs gets reduced considerably while still capturing the variations in shape patterns. This results in a compact, efficient, and robust recognizer with reduced model set. In the current paper we are presenting our recent improvements in sub-character HMM modeling for Arabic text recognition where we use special 'connector' and 'space' models. Additionally we investigated contextual sub-characters HMMs for text recognition. We also present multi-stream contextual sub-character HMMs where the features calculated from a sliding window frame form one stream and its derivative features are part of the second stream. We report state-of-the-art results on the IFN/ENIT (benchmark) database of handwritten Arabic text and the recognition rate of 85.12% on sets outperforms previously published results.
Irfan Ahmad 0001, Gernot A. Fink, Sabri A. Mahmoud
ICFHR2
2014 Grouping Historical Postcards Using Query-by-Example Word Spotting
abstract
Handwritten historical documents pose extremely challenging problems for automatic analysis. This is due to the high variability observed in handwritten script, the use of writing styles and script types unknown today, the frequently lacking orthographic standardization, and the degradation of the respective documents. Therefore, it is currently out of question to develop general purpose handwriting recognition systems for historical document collections. It is, however, possible to search relatively homogeneous document collections using word spotting techniques. In this paper we consider the analysis of a challenging collection of postcards from the period of World War I delivered by the German military postal service. More specifically, we consider the automatic grouping of mail pieces by spotting potentially identical addressees. As the annotation of such documents is extremely challenging even for trained experts, a manually developed ground truth annotation will, in general, not be available. Furthermore, a reliable segmentation on word level will hardly be possible. With our segmentation-free query-by-example word spotting method we investigate modifications addressing the better generalization to a multi-writer scenario and its application to degraded documents. Promising results could be achieved in this highly challenging scenario.
Gernot A. Fink, Leonard Rothacker, René Grzeszick
ICFHR1
2014 KHATT: An open Arabic offline handwritten text database
Sabri A. Mahmoud, Irfan Ahmad 0001, Wasfi G. Al-Khatib, Mohammad R. Alshayeb, Mohammad Tanvir Parvez, Volker Märgner, Gernot A. Fink
Pattern Recognit.7
2014 Semi-supervised learning for character recognition in historical archive documents
Jan Richarz, Szilárd Vajda, René Grzeszick, Gernot A. Fink
Pattern Recognit.4
2013 Novel Sub-character HMM Models for Arabic Text Recognition
abstract
Hidden Markov Model (HMM) is one of the most widely used classifier for text recognition. In this paper we are presenting novel sub-character HMM models for Arabic text recognition. Modeling at sub-character level allows sharing of common patterns between different contextual forms of Arabic characters as well as between different characters. The number of HMMs gets reduced considerably while still capturing the variations in shape patterns. This results in a compact and efficient recognizer with reduced model set and is expected to be more robust to the imbalance in data distribution. Experimental results using the sub-character model based recognition of handwritten Arabic text as well printed Arabic text are reported.
Irfan Ahmad 0001, Leonard Rothacker, Gernot A. Fink, Sabri A. Mahmoud
ICDAR3
2013 Statistical Modeling of the Relation between Characters and Diacritics in Lampung Script
abstract
Lampung Script is a non-cursive script where a rich set of diacritics is used to modify the syllable denoted by a character symbol. Consequently, the analysis of the relation between characters and diacritic marks associated with them plays an important role in the recognition process. As diacritics can appear in three different relative positions with respect to a character (top, bottom, and right) associating them correctly with a character is a challenging problem. In this paper we propose a novel approach for modeling the relations between characters and diacritics in handwritten Lampung documents. First, a document is segmented into characters and diacritic marks. Then every character defines a normalized coordinate system into which nearby diacritics can be mapped. The relation between a diacritic mark and its associated character can then be described by a statistical model. In a writer independent experimental evaluation we investigate models with different degrees of specialization with respect to their capability of predicting the correct character-to diacritic associations. We achieve significant error rate reductions with respect to a naive association model using a nearest-neighbor criterion.
Akmal Junaidi, René Grzeszick, Gernot A. Fink, Szilárd Vajda
ICDAR3
2013 Bag-of-Features HMMs for Segmentation-Free Word Spotting in Handwritten Documents
abstract
Recent HMM-based approaches to handwritten word spotting require large amounts of learning samples and mostly rely on a prior segmentation of the document. We propose to use Bag-of-Features HMMs in a patch-based segmentation-free framework that are estimated by a single sample. Bag-of-Features HMMs use statistics of local image feature representatives. Therefore they can be considered as a variant of discrete HMMs allowing to model the observation of a number of features at a point in time. The discrete nature enables us to estimate a query model with only a single example of the query provided by the user. This makes our method very flexible with respect to the availability of training data. Furthermore, we are able to outperform state-of-the-art results on the George Washington dataset.
Leonard Rothacker, Marçal Rusiñol, Gernot A. Fink
ICDAR3
2013 Bag-of-features representations using spatial visual vocabularies for object classification
abstract
This paper presents a novel method for combining local image features and spatial information for object classification tasks using the Bag-of-Features principle. The feature descriptor is extended by additional spatial information. Hence, similar feature descriptors do not only describe similar image patches, but similar patches in roughly the same region. Different spatial measures are evaluated on the Caltech 101 dataset showing the improvement by incorporating spatial information into the feature descriptor. Furthermore, the method achieves better classification rates than the comparable Spatial Pyramids with lower a dimensional representation.
René Grzeszick, Leonard Rothacker, Gernot A. Fink
ICIP3
2012 Towards Semi-supervised Transcription of Handwritten Historical Weather Reports
abstract
This paper addresses the automatic transcription of handwritten documents with a regular tabular structure. A method for extracting machine printed tables from images is proposed, using very little prior knowledge about the document layout. The detected table serves as query for retrieving and fitting a structural template, which is then used to extract handwritten text fields. A semi-supervised learning approach is applied to this fields, aiming at minimizing the human labeling effort for recognizer training. The effectiveness of the proposed approach is demonstrated experimentally on a set of historical weather reports. Compared to using all labels, competitive recognition performance is achieved by labeling only a small fraction of the data, keeping the required human effort very low.
Jan Richarz, Szilárd Vajda, Gernot A. Fink
Document Analysis Systems3
2012 KHATT: Arabic Offline Handwritten Text Database
abstract
In this paper, we report our comprehensive Arabic offline Handwritten Text database (KHATT) after completion of the collection of 1000 handwritten forms written by 1000 writers from different countries. It is composed of an image database containing images of the written text at 200, 300, and 600 dpi resolutions, a manually verified ground truth database that contains meta-data describing the written text at the page, paragraph, and line levels. A formal verification procedure is implemented to align the handwritten text with its ground truth at the form, paragraph and line levels. Tools to extract paragraphs from pages and segment paragraphs into lines are developed. Preliminary experiments on Arabic handwritten text recognition are conducted using sample data from the database and the results are reported. The database will be made freely available to researchers world-wide for research in various handwritten-related problems such as text recognition, writer identification and verification, etc.
Sabri A. Mahmoud, Irfan Ahmad 0001, Mohammad R. Alshayeb, Wasfi G. Al-Khatib, Mohammad Tanvir Parvez, Gernot A. Fink, Volker Märgner, Haikal El Abed
ICFHR6
2012 Annotating Handwritten Characters with Minimal Human Involvement in a Semi-supervised Learning Strategy
abstract
One obstacle in the automatic analysis of handwritten documents is the huge amount of labeled data typically needed for classifier training. This is especially true when the document scans are of bad quality and different writers and writing styles have to be covered. Consequently, the considerable human effort required in the process currently prohibits the automatic transcription of large document collections. In this paper, two semi-supervised multiview learning approaches are presented, reducing the manual burden by robustly deriving a large number of labels from relatively few manual annotations. The first is based on cluster-level annotation followed by a majority decision, whereas the second casts the labeling process as a retrieval task and derives labels by voting among ranked lists. Both methods are thoroughly evaluated in a handwritten character recognition scenario using realistic document data. It is demonstrated that competitive recognition performance can be maintained by labeling only a fraction of the data.
Jan Richarz, Szilárd Vajda, Gernot A. Fink
ICFHR3
2012 Bag-of-Features Representations for Offline Handwriting Recognition Applied to Arabic Script
abstract
Due to the great variabilities in human writing, unconstrained handwriting recognition is still considered an open research topic. Recent trends in computer vision, however, suggest that there is still potential for better recognition by improving feature representations. In this paper we focus on feature learning by estimating and applying a statistical bag-of-features model. These models are successfully used in image categorization and retrieval. The novelty here is the integration with a Hidden Markov Model (HMM) that we use for recognition. Our method is evaluated on the IFN/ENIT database consisting of images of handwritten Arabic town and village names.
Leonard Rothacker, Szilárd Vajda, Gernot A. Fink
ICFHR3
2012 Sift-based Camera Localization using Reference Objects for Application in Multi-camera Environments and Robotics
Hanno Jaspers, Boris Schauerte, Gernot A. Fink
ICPRAM (2)3
2011 A new method for combined face detection and identification using interest point descriptors
abstract
Although face recognition has been an active research area for decades, the general problem is still unsolved. While many methods addressing either face detection or face identification emerged, little attention has been paid to the combination or integration of detection and identification. In this paper we propose a combined method for face detection and identification using SIFT descriptors. This combined method includes an existing detection model and a new identification method based on object class invariants (OCIs), which is invariant to translation, scale, in-plane rotation and small 3D viewpoint changes. These models are combined using a bounding box around the OCI to filter face features for identification. We show the highly competitive performance of the newly developed identification method and the effectiveness of the proposed combination scheme.
Sebastian Stein 0003, Gernot A. Fink
FG2
2011 Multiple speaker tracking using a microphone array by combining auditory processing and a gaussian mixture cardinalized probability hypothesis density filter
abstract
Tracking speakers is an important application in smart environments. Acoustic tracking using microphone arrays is a challenging task due to two major reasons: On the one hand, multiple persons may speak simultaneously and thus the number of speakers varies over time; on the other hand, due to the nature of reverberated speech, the provided position hypotheses contain many gaps and clutter. In the proposed approach, the "glimpsing model" is realized by neurobiologically in spired calculation of robust but sparse position hypotheses in combination with a Gaussian mixture cardinalized probability hypothesis density filter. By iteratively applying the filter to the position hypotheses from multiple frequency bands, good results are achieved. Using a statistical speech model derived from recordings, a real-time capable implementation is used to track multiple speakers in a conference room with significant reverberation.
Axel Plinge, Daniel Hauschildt, Marius H. Hennecke, Gernot A. Fink
ICASSP4
2011 A Semi-supervised Ensemble Learning Approach for Character Labeling with Minimal Human Effort
abstract
One of the major issues in handwritten character recognition is the efficient creation of ground truth to train and test the different recognizers. The manual labeling of the data by a human expert is a tedious and costly procedure. In this paper we propose an efficient and low-cost semi-automatic labeling system for character datasets. First, the data is represented in different abstraction levels, which is clustered after in an unsupervised manner. The different clusters are labeled by the human experts and finally an unanimity voting is considered to decide if a label is accepted or not. The experimental results prove that labeling only less than 0.5% of the training data is sufficient to achieve 86.21% recognition rate for a brand new script (Lampung) and 94.81% for the MNIST benchmark dataset, considering only a K-nearest neighbor classifier for recognition.
Szilárd Vajda, Akmal Junaidi, Gernot A. Fink
ICDAR3
2011 Unsupervised Geometry Calibration of Acoustic Sensor Networks Using Source Correspondences
abstract
In this paper we propose a procedure for estimating the geometric configuration of an arbitrary acoustic sensor placement. It determines the position and the orientation of microphone arrays in 2D while locating a source by direction-of-arrival (DoA) estimation. Neither artificial calibration signals nor unnatural user activity are required. The problem of scale indeterminacy inherent to DoA-only observations is solved by adding time difference of arrival (TDOA) measurements. The geometry calibration method is numerically stable and delivers precise results in moderately reverberated rooms. Simulation results are confirmed by laboratory experiments.
Joerg Schmalenstroeer, Florian Jacob, Reinhold Häb-Umbach, Marius H. Hennecke, Gernot A. Fink
INTERSPEECH5
2010 Online Bangla Word Recognition Using Sub-Stroke Level Features and Hidden Markov Models
abstract
For automatic recognition of Bangla script, only a few studies are reported in the literature, which is in contrast to the role of Bangla as one of the world's major scripts. In this paper we present a new approach to online Bangla handwriting recognition and one of the first to consider cursively written words instead of isolated characters. Our method uses a sub-stroke level feature representation of the script and a writing model based on hidden Markov models. As for the latter an appropriate internal structure is crucial, we investigate different approaches to defining model structures for a highly compositional script like Bangla. In experimental evaluations of a writer independent Bangla word recognition task we show that the use of context-dependent sub-word units achieves quite promising results and significantly outperforms alternatively structured models.
Gernot A. Fink, Szilárd Vajda, Ujjwal Bhattacharya, Swapan K. Parui, Bidyut B. Chaudhuri
ICFHR1
2010 Strategies for Training Robust Neural Network Based Digit Recognizers on Unbalanced Data Sets
abstract
The performance of a neural network in a pattern recognition task may be influenced by several factors. One of these factors is related to the considerable difference between the number of examples belonging to each class to be recognized. The effect called imbalanced data can negatively influence the ability of a recognizer to learn the concept of the minority class. In this work we propose an under-sampling strategy based on selecting samples lying around the decision surface and an over-sampling strategy which uses kernel density estimation to populate the minority class. The experimental results on Roman and Bangla digit data using a neural network based recognizer confirm the effectiveness of the proposed solutions.
Szilárd Vajda, Gernot A. Fink
ICFHR2
2010 Exploring Pattern Selection Strategies for Fast Neural Network Training
abstract
Nowadays, the usage of neural network strategies in pattern recognition is a widely considered solution. In this paper we propose three different strategies to select more efficiently the patterns for a fast learning in such a neural framework by reducing the number of available training patterns. All the strategies rely on the idea of dealing just with samples close to the decision boundaries of the classifiers. The effectiveness (accuracy, speed) of these methods is confirmed through different experiments on the MNIST handwritten digit data [1], Bangla handwritten numerals [2] and the Shuttle data from the UCI machine learning repository [3].
Szilárd Vajda, Gernot A. Fink
ICPR2
2010 Saliency-based identification and recognition of pointed-at objects
abstract
When persons interact, non-verbal cues are used to direct the attention of persons towards objects of interest. Achieving joint attention this way is an important aspect of natural communication. Most importantly, it allows to couple verbal descriptions with the visual appearance of objects, if the referred-to object is non-verbally indicated. In this contribution, we present a system that utilizes bottom-up saliency and pointing gestures to efficiently identify pointed-at objects. Furthermore, the system focuses the visual attention by steering a pan-tilt-zoom camera towards the object of interest and thus provides a suitable model-view for SIFT-based recognition and learning. We demonstrate the practical applicability of the proposed system through experimental evaluation in different environments with multiple pointers and objects.
Boris Schauerte, Jan Richarz, Gernot A. Fink
IROS3
2009 Face Detection Using GPU-Based Convolutional Neural Networks
Fabian Naße, Christian Thurau, Gernot A. Fink
CAIP3
2009 Multi-modal and multi-camera attention in smart environments
abstract
This paper considers the problem of multi-modal saliency and attention. Saliency is a cue that is often used for directing attention of a computer vision system, e.g., in smart environments or for robots. Unlike the majority of recent publications on visual/audio saliency, we aim at a well grounded integration of several modalities. The proposed framework is based on fuzzy aggregations and offers a flexible, plausible, and efficient way for combining multi-modal saliency information. Besides incorporating different modalities, we extend classical 2D saliency maps to multi-camera and multi-modal 3D saliency spaces. For experimental validation we realized the proposed system within a smart environment. The evaluation took place for a demanding setup under real-life conditions, including focus of attention selection for multiple subjects and concurrently active modalities.
Boris Schauerte, Jan Richarz, Thomas Plötz, Christian Thurau, Gernot A. Fink
ICMI5
2009 A Multi-modal Attention System for Smart Environments
Boris Schauerte, Thomas Plötz, Gernot A. Fink
ICVS3
2009 Markov models for offline handwriting recognition: a survey
abstract
Since their first inception more than half a century ago, automatic reading systems have evolved substantially, thereby showing impressive performance on machine-printed text. The recognition of handwriting can, however, still be considered an open research problem due to its substantial variation in appearance. With the introduction of Markovian models to the field, a promising modeling and recognition paradigm was established for automatic offline handwriting recognition. However, so far, no standard procedures for building Markov-model-based recognizers could be established though trends toward unified approaches can be identified. It is therefore the goal of this survey to provide a comprehensive overview of the application of Markov models in the research field of offline handwriting recognition, covering both the widely used hidden Markov models and the less complex Markov-chain or n -gram models. First, we will introduce the typical architecture of a Markov-model-based offline handwriting recognition system and make the reader familiar with the essential theoretical concepts behind Markovian models. Then, we will give a thorough review of the solutions proposed in the literature for the open problems how to apply Markov-model-based approaches to automatic offline handwriting recognition.
Thomas Plötz, Gernot A. Fink
Int. J. Document Anal. Recognit.2
2008 SVM ensemble classification of NMR spectra based on different configurations of data processing techniques
abstract
The early detection of drug-induced organ toxicities is one of the major goals in safety pharmacology. Automating this process by classification of metabolic changes based on the analysis of1H nuclear magnetic resonance spectra improves this process. In this paper we propose an ensemble classification system based on support vector machines trained on diverse ldquoviewsrdquo on the data. These views are created by variation of preprocessing techniques and the final classification is achieved by voting on an optimized selection of all experts. Results of an experimental evaluation on a challenging data-set from industrial safety pharmacology show the effectiveness of the proposed approach w.r.t. the detection of drug-induced toxicity.
Kai Lienemann, Thomas Plötz, Gernot A. Fink
ICPR3
2008 Calibration-free camera hand-over for fast and reliable person tracking in multi-camera setups
abstract
Ensembles of multiple (active) cameras yield an important ingredient in modern tracking and surveillance applications. They overcome the limited fields-of-view of single cameras, however, require robust procedures for handing over tracking tasks from one camera to another. In this paper a calibration-free procedure is proposed that allows for fast and reliable camera hand-over in Ambient Intelligence (AmI) applications. The approach is based on online acquisition of scenario-specific target models and especially solves the problem of significant changes in object view during hand-over. Real-world results acquired in an AmI environment prove the effectiveness of our technique.
Birgit Möller 0001, Thomas Plötz, Gernot A. Fink
ICPR3
2008 Real-time detection and interpretation of 3D deictic gestures for interactionwith an intelligent environment
abstract
We present a system that enables pointing-based unconstrained interaction with a smart conference room using an arbitrary multicamera setup. For each individual camera stream, areas exhibiting strong motion are identified. In these areas, face and hand hypotheses are detected. The detections of multiple cameras are then combined to 3D hypotheses from which deictic gestures are identified and a pointing direction is derived. This is then used to identify objects in the scene. Since we use a combination of simple yet effective techniques, the system runs in real-time and is very responsive. We present evaluation results on realistic data that show the capabilities of the presented approach.
Jan Richarz, Thomas Plötz, Gernot A. Fink
ICPR3
2007 On the Use of Context-Dependent Modeling Units for HMM-Based Offline Handwriting Recognition
abstract
The use of context dependent modeling units in handwriting recognition has been considered by many authors as promising substantial performance improvements in systems based on Hidden-Markov models. Interestingly, in the literature only a few approaches limited to online recognition are documented to make use of this technology. Therefore, we investigated whether context dependent modeling also offers advantages for offline recognition systems. The moderate performance improvements we achieved on a challenging unconstrained handwriting recognition task suggest that context dependent modeling can not easily be exploited for offline recognition. In this paper we will present the principles behind context dependent modeling and discuss the reasons for its limited applicability in recognizing offline handwriting data.
Gernot A. Fink, Thomas Plötz
ICDAR1
2006 Pattern recognition methods for advanced stochastic protein sequence analysis using HMMs
Thomas Plötz, Gernot A. Fink
Pattern Recognit.2
2005 On Appearance-Based Feature Extraction Methods for Writer-Independent Handwritten Text Recognition
abstract
Most successful systems for the recognition of unconstrained handwriting currently rely on expert-crafted feature sets that compute local geometric properties from text images. However, by applying appearance based analysis techniques appropriate features could be derived from training data automatically. Therefore, in this paper, several different methods for computing appearance-based feature representations are investigated and compared to the performance of a state-of-the-art writer-independent recognition system based on geometric features. In extensive experiments, promising results were obtained on a challenging recognition task.
Gernot A. Fink, Thomas Plötz
ICDAR1
2005 Modality integration and dialog management for a robotic assistant
abstract
The communication with robotic assistants or companions is a challenging new domain for the use of dialog systems. In contrast to classical spoken language interfaces users interact with mobile robots mostly in a multi-modal way. In this paper we will present the integration of several modalities in the dialog system of BIRON — the Bielefeld Robot Companion. Besides speech as the main modality the system integrates deictic gestures and visual scene information in order to resolve object references in a task oriented dialog. We will present example interactions with BIRON and first qualitative results from the "home-tour" scenario defined within the COGNIRON project.
Ioannis Toptsis, Axel Haasch, Sonja Hwel, Jannik Fritsch, Gernot A. Fink
INTERSPEECH5
2005 Toward automatic video-based whiteboard reading
Markus Wienecke, Gernot A. Fink, Gerhard Sagerer
Int. J. Document Anal. Recognit.2
2004 A multi-modal dialog system for a mobile robot
abstract
A challenging domain for dialog systems is their use for the communication with robotic assistants. In contrast to the classical use of spoken language for information retrieval, on a mobile robot multi-modal dialogs and the dynamic interaction of the robot system with its environment have to be considered. \nIn this paper we will present the dialog system developed for BIRON — the Bielefeld Robot Companion. The system is able to handle multi-modal dialogs by augmenting semantic interpretation structures derived from speech with hypotheses for additional modalities as e.g. speech-accompanying gestures. The architecture of the system is modular with the dialog manager being the central component. In order to be aware of the dynamic behavior of the robot itself, the possible states of the robot control system are integrated into the dialog model.\nFor flexible use and easy configuration the communication between the individual modules as well as the declarative specification of the dialog model are encoded in XML.\nWe will present example interactions with BIRON from the ’ scenario defined within the COGNIRON project.
Ioannis Toptsis, Shuyin Li, Britta Wrede, Gernot A. Fink
INTERSPEECH4
2003 Towards Automatic Video-based Whiteboard Reading
abstract
As whiteboards have become a popular tool in meeting rooms, there has been a growing interest in making use of the whiteboard as a user interface for human computer interaction. Therefore, systems based on electronic whiteboards have been developed in order to serve as meeting assistants for e.g. collaborative working. However, as special pens and erasers are required, the natural interaction is restricted. In order to render this communication method more natural it was proposed to retain ordinary whiteboard and pens and to visually observe the writing process using a video camera by Stafford-Fraser and Robinson (1996). In this paper a prototype system for automatic video-based whiteboard reading is presented. The system is designed for recognizing unconstrained handwritten text and is further characterized by an incremental processing strategy in order to facilitate recognizing portions of text as soon as they have been written on the board. We present the methods employed for extracting text regions, pre-processing, feature extraction, and statistical modeling and recognition. Evaluation results on a writer independent unconstrained handwriting recognition task demonstrate the feasibility of the proposed approach.
Markus Wienecke, Gernot A. Fink, Gerhard Sagerer
ICDAR2
2003 Providing the basis for human-robot-interaction: a multi-modal attention system for a mobile robot
abstract
In order to enable the widespread use of robots in home and office environments, systems with natural interaction capabilities have to be developed. A prerequisite for natural interaction is the robot's ability to automatically recognize when and how long a person's attention is directed towards it for communication. As in open environments several persons can be present simultaneously, the detection of the communication partner is of particular importance. In this paper we present an attention system for a mobile robot which enables the robot to shift its attention to the person of interest and to maintain attention during interaction. Our approach is based on a method for multi-modal person tracking which uses a pan-tilt camera for face recognition, two microphones for sound source localization, and a laser range finder for leg detection. Shifting of attention is realized by turning the camera into the direction of the person which is currently speaking. From the orientation of the head it is decided whether the speaker addresses the robot. The performance of the proposed approach is demonstrated with an evaluation. In addition, qualitative results from the performance of the robot at the exhibition part of the ICVS'03 are provided.
Sebastian Lang 0002, Marcus Kleinehagenbrock, Sascha Hohenner, Jannik Fritsch, Gernot A. Fink, Gerhard Sagerer
ICMI5
2003 Data-driven pronunciation modeling for ASR using acoustic subword units
abstract
We describe a method to model pronunciation variation for ASR in a data-driven way, namely by use of automatically derived acoustic subword units. The inventory of units is designed so as to produce maximal separable pronunciation variants of words while at the same time only the most important variants for the particular application are trained. In doing so, the optimal number of variants per word is determined iteratively. All this is accomplished (almost) fully automatically by use of a state splitting algorithm and a variant distance measure. Compared to a baseline system using triphones as subword units and with minimal pronunciation variants, this method achieved a relative improvement of the word error rate by 10%.
Thurid Spiess, Britta Wrede, Gernot A. Fink, Franz Kummert
INTERSPEECH3
2002 Robust time-synchronous environmental adaptation for continuous speech recognition systems
abstract
In this paper we describe system architectures for robust MLLR based environmental adaptation of continuous speech recognition systems. Inspired by an existing broadcast news transcription system we refined the identification of acoustic scenarios by using a combined GMM/HMM method. Thus environmental adaptation regarding arbitrary acoustic scenarios beyond speaker changes becomes possible. For deploying acoustic adaptation in interactive applications, such as human machine interaction, a time-synchronous adaptation approach is proposed. For different corpora the evaluation of our approaches shows significant improvements in recognition accuracy while satisfying the constraint of time-synchronous processing.
Thomas Plötz, Gernot A. Fink
INTERSPEECH2
2002 Dynamic search-space pruning for time-constrained speech recognition
abstract
In automatic speech recognition complex state spaces are searched during the recognition process. By limiting these search spaces the computation time can be reduced, but unfortunately the recognition rate mostly decreases, too. However, especially for time-critical recognition tasks a search-space pruning is necessary. Therefore, we developed a dynamic mechanism to optimize the pruning parameters for time-constrained recognition tasks, e.g. speech recognition for robotic systems, in respect to word accuracy and computation time. With this mechanism an automatic speech recognition system can process speech signals with an approximately constant processing rate. Compared to a system without such a dynamic mechanism and the same time available for computation, the variance of the processing rate is decreased greatly without a significant loss of word accuracy. Furthermore, the extended system can be sped up to real-time processing, if desired or necessary.
Sascha Wendt, Gernot A. Fink, Franz Kummert
INTERSPEECH2
2002 Multi-modal human-machine communication for instructing robot grasping tasks
abstract
A major challenge for the realization of intelligent robots is to supply them with cognitive abilities in order to allow ordinary users to program them easily and intuitively. One approach to such programming is teaching work tasks by interactive demonstration. To make this effective and convenient for the user, the machine must be capable of establishing a common focus of attention and be able to use and integrate spoken instructions, visual perception, and non-verbal clues like gestural commands. We report progress in building a hybrid architecture that combines statistical methods, neural networks, and finite state machines into an integrated system for instructing grasping tasks by man-machine interaction. The system combines the GRAVIS-robot for visual attention and gestural instruction with an intelligent interface for speech recognition and linguistic interpretation, and a modality fusion module to allow multi-modal task-oriented man-machine communication with respect to dextrous robot manipulation of objects.
Patrick C. McGuire, Jannik Fritsch, Jochen J. Steil, Frank Röthling, Gernot A. Fink, Sven Wachsmuth, Gerhard Sagerer, Helge J. Ritter
IROS5
2002 Combining acoustic and articulatory feature information for robust speech recognition
Katrin Kirchhoff, Gernot A. Fink, Gerhard Sagerer
Speech Commun.2
2001 Video-Based On-line Handwriting Recognition
abstract
The use of handwriting provides a natural way of interacting with small portable computers. However, in order to capture handwritten text. online, special input devices are necessary. Therefore, M.E. Munich & P. Perona (1996) proposed to use visual input for pen-based computers. Writing can then be performed on ordinary paper, and pen trajectories are automatically extracted from image sequences recorded during the writing process. On the basis of this work, we developed a complete video-based online handwriting recognition system. We will present the techniques applied for pen tracking, pre-processing, feature extraction, and statistical modeling and recognition. Evaluation results on a writer-independent unconstrained handwriting recognition task demonstrate that the inherent limitations of the video-based approach can be compensated using robust modeling combined with adaptation techniques.
Gernot A. Fink, Markus Wienecke, Gerhard Sagerer
ICDAR1
2001 Forward masking for increased robustness in automatic speech recognition
abstract
In automatic speech recognition MFCC or LPCC are features commonly used today. However, their calculation considers only a few features of the auditory system. On the assumption that the human representation of speech is an optimal representation, considering more features of the auditory system might lead to a better performance of automatic speech recognition systems. In this paper a model proposed by Strope and Alwan (see references), which relies on the human acoustic perception and allows to consider the effect of forward masking, is incorporated after some modifications into an automatic speech recognition system with a MFCC-based front-end. The extended system is evaluated on recognition tasks, that are closer to real recognition than (connected) digit recognition commonly used in the literature. The evaluations show an increased robustness of the speech recognition system with forward masking on all recognition tasks, but especially on data recorded in noisy environments.
Sascha Wendt, Gernot A. Fink, Franz Kummert
INTERSPEECH2
2001 An investigation of modelling aspects for ratedependent speech recognition
abstract
For the modelling of speech rate variation in speech recognition many approaches have been suggested. However, the training of speech-rate dependent models has by far received most of the attention. In order to investigate problematic aspects related with the classification of the speech data which represents one of the major problems of these approaches extensive experiments were carried out on a German corpus of read speech. The results indicate that while the kind of the model-driven speech-rate measure is only of minor importance a data-driven classification of the speech data significantly improves the performance of rate-dependent models. Further results suggest a detailed modelling of speech rate based on more general models. This means that it might be possible to model speech rate adaptation by means of a transformation based on a continuous measure.
Britta Wrede, Gernot A. Fink, Gerhard Sagerer
INTERSPEECH2
2000 Conversational speech recognition using acoustic and articulatory input
abstract
The combination of multiple speech recognizers based on different signal representations is increasingly attracting interest in the speech community. In previous work we presented a hybrid speech recognition system based on the combination of acoustic and articulatory information which achieved significant word error rate reductions under highly noisy conditions on a small-vocabulary numbers recognition task. In this study we extend this approach to large-vocabulary conversational speech recognition using the Gaussian mixture acoustic modeling paradigm. We demonstrate that the articulatory input representation we propose contains information which is complementary to that provided by standard MFCC features, and that their combination can significantly reduce the word error rate on conversational speech. Various combination strategies (feature-level, state-level and word-level combination) are compared and evaluated.
Katrin Kirchhoff, Gernot A. Fink, Gerhard Sagerer
ICASSP2
2000 A hybrid speech recognizer combining HMMs and polynomial classification
abstract
In this paper, we present a hybrid speech recognizer combining Hidden Markov Models (HMMs) and a polynomial classifier. In our approach the emission probabilities are not modeled as a mixture of Gaussians but are calculated by the polynomial classifier. However, we do not apply the classifier directly to the feature vector but we make use of the density values of L Gaussians clustering the feature space. That means we model the emission probability as a polynomial of Gaussian distributions of n-th degree. As most of these density values are approximately zero for a single feature vector the calculation of a polynomial can be done very efficiently. The usefulness of this hybrid approach was successfully tested on a large conversational speech recognition task. 1.
Franz Kummert, Gernot A. Fink, Gerhard Sagerer
INTERSPEECH2
2000 Grapheme based speech recognition for large vocabularies
abstract
Common speech recognition systems use phonetically motivated subword units. To utilize words in these systems, one has to translate the available graphemic word representation into a phonetic one. To reduce this manual effort we propose to build grapheme based recognition systems. They can be used as speech interfaces for devices that can provide a graphemic representation of words like city names of navigation systems. Results of experiments on a 10,000 word lexicon of German cities are presented. 1.
Christoph Schillo, Gernot A. Fink, Franz Kummert
INTERSPEECH2
2000 Influence of duration on static and dynamic properties of German vowels in spontaneous speech
abstract
Changes in speech rate severely affect the performance of continuous speech recognition systems. In order to better understand the underlying effects of speech rate changes an analysis was carried out on the influence of duration on the spectral properties of vowels in a large German corpus of spontaneous speech. The results show a strong centralisation effect of the vowel formant frequencies due to shorter duration while the formant movements are only slightly affected. The data suggest that the movement velocity is not changed in vowels with a limited duration. As the means of the on- and offset frequencies also remain stable only the middle part of the vowels are affected by the centralisation effect. These results are discussed in the light of the modelling of varying speech rate in automatic speech recognition systems. 1.
Britta Wrede, Gernot A. Fink, Gerhard Sagerer
INTERSPEECH2
1999 A comparative study of HMM-based approaches for the automatic recognition of perceptually relevant aspects of spontaneous German speech melody
Christel Brindöpke, Gernot A. Fink, Franz Kummert
EUROSPEECH2
1998 Hybrid object recognition in image sequences
abstract
We present a hybrid approach attaching probabilistic formalisms, as artificial neural networks or hidden Markov models, to concepts of a semantic network for a robust and efficient detection of objects. Additionally, an efficient processing strategy for image sequences is outlined which propagates the structural results of the semantic network as an expectation for the next image. This method allows one to produce linked results over time supporting the recognition of events and actions.
Franz Kummert, Gernot A. Fink, Gerhard Sagerer, Elke Braun
ICPR2
1998 A HMM-based recognition system for perceptive relevant pitch movements of spontaneous German speech
abstract
This paper presents an HMM-based recognition system for perceptive relevant pitch movements of spontaneous German speech. The pitch movements are defined according to the perceptively and phonetically motivated IPO-approach to intonation. For recognition we use a hybrid approach combining polynomial classification with Hidden Markov Modelling. The recognition is based only on the speech signal, its fundamental frequency and eleven derived features. We evaluate the system on a speaker independant recognition task. 1 Introduction In current speech recognition systems, usually, no prosodic information is used. However, it is a wellknown assumption that prosody can contribute useful information to enhance speech recognition and understanding processes [6]. While for speech recognition, there is no doubt that recognition processes should be resulting in a sequence of word hypotheses for prosody recognition the chosen units depend on several competing linguistic theories. Although their ad...
Christel Brindöpke, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
ICSLP2
1996 A robust dialogue system for making an appointment
abstract
A complete dialogue system within the task domain of making an appointment is presented.It is based on a semantic network representation of linguistic knowledge and a word recognition system that communicates with the interpretation component bidirectionally.System robustness is achieved using a special metascore that evaluates the advance of the linguistic interpretation.
Hans Brandt-Pook, Gernot A. Fink, Bernd Hildebrandt, Franz Kummert, Gerhard Sagerer
ICSLP2
1996 Incremental generation of word graphs
Gerhard Sagerer, Heike Rautenstrauch, Gernot A. Fink, Bernd Hildebrandt, A. Jusek, Franz Kummert
ICSLP3
1995 A flexible formal language for the orthographic transcription of spontaneous spoken dialogues
abstract
Orthographic transcriptions of speech are important in most fields of research concerned with spoken language. For spontaneous speech they have to be created manually, resulting potentially in inconsistent or erroneous transcriptions. We propose a new flexible and easy-to-use formal language for the orthographic transcription of spontaneous speech. All relevant phenomena introduced by spontaneous spoken dialogues are covered. The transcription serves as a meta-language from which various representations for different research purposes can be generated automatically. 1. INTRODUCTION For linguistic research as well as for studies in the fields of speech recognition and understanding orthographic transcriptions of utterances play an important role. If read aloud speech is investigated those are instantly available. However, if we are concerned with spontaneous speech orthographic transcriptions have to be created by experts in time consuming and error prone work. This is due to the synta...
Gernot A. Fink, Michaela Johanntokrax, Brigitte Schaffranietz
EUROSPEECH1
1995 Detection of unknown words and its evaluation
abstract
Especially in recognition of spontaneous speech it is necessary to be able to cope with the occurance of unknown words. We present an approach to unknown word detection integrated with speech recognition. Though several schemes exist for the evaluation of detection algorithms all suffer from some deficiencies. Therefore, we also present a new evaluation measure called detection accuracy which is similar to the widely accepted word accuracy. We applied this measure to the evaluation of our approach on a large spontaneous speech recognition task from the German Verbmobil project.
A. Jusek, Gernot A. Fink, Franz Kummert, Heike Rautenstrauch, Gerhard Sagerer
EUROSPEECH2
1995 Generation of language models using the results of image analysis
abstract
We present a new approach towards using contextual information to enhance speech recognition and understanding. Dynamically inferred knowledge about the context is used in addition to the static linguistic and domain specific knowledge. Based on the results of image analysis of a given scene language models for constituents of possible utterances concerning that scene are generated.
Uta Naeve, Gudrun Socher, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
EUROSPEECH3
1994 A close high-level interaction scheme for recognition and interpretation of speech
abstract
The vast majority of speech understanding systems suffers from a bottleneck between the recognition and the interpretation components. Normally, only a relatively small set of word hypotheses is passed from the recognizer and no flow of information in the opposite direction is even possible. We propose an interaction scheme that tries to overcome many of the disadvantages of traditional systems. It makes use of the possibility to process abstract constituents in our word recognizer and pass them back as complex hypotheses. Predictions that define the complex analysis goal of the recognition can be derived dynamically during the interpretation of an utterance. A left-to-right processing in both recognition and interpretation makes an incremental analysis possible.
Gernot A. Fink, Franz Kummert, Gerhard Sagerer
ICSLP1
1994 Understanding of time constituents in spoken language dialogues
abstract
The analysis and interpretation of time constituents is a rather complex enterprise, since diverse time constituents are distributable in variable positions within an utterance. The first step in order to manage this complexity and variability is the syntactic analysis at phrase structure level. Within a utterance each time constituent is analyzed independently and tested for its syntactic coherence. A semantic interpretation of the time constituent has to follow. The second step consists of the analysis and interpretation at sentence structure level. The time interpretations need to be tested for consistency and merged into a single representation. Here it is usually possible to resolve ambiguities. As a last step, the interpretations of time constituents have to be merged at dialogue level. Although the system asks the user for verification of the time computed, users do seldom just reply ’ or ’. Mostly they add new information about time, and sometimes users correct the system’s interpretation without explicit negation. Thus, the merging of time constituents at dialogue level becomes a rather complex issue
Bernd Hildebrandt, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
ICSLP2
1994 A Speech Understanding and Dialog System with a Homogeneous Linguistic Knowledge Base
abstract
This article presents the speech understanding and dialog system EVAR. All levels of linguistic knowledge are used both to control the analysis process and for the interpretation of an utterance. All kinds of knowledge are integrated in a homogeneous knowledge base. The control algorithm used for the analysis is defined within the representation scheme and does not depend on the application. One of the aims of EVAR is to develop a system structure where linguistic and nonlinguistic expectations could be used not only for the interpretation but also as predictions for the recognition process.>
Marion Mast, Franz Kummert, Ute Ehrlich, Gernot A. Fink, Thomas Kuhn 0002, Heinrich Niemann, Gerhard Sagerer
IEEE Trans. Pattern Anal. Mach. Intell.4
1993 Robust interpretation of speech
abstract
In all fields of pattern recognition there may arise situations where severe distortions of the sensor data or errors of processing prevent a successful analysis. This is especially true in speech recognition. Therefore, a speech understanding system can not rely on being able to interpret the whole or even a given fixed percentage of the input data. Rather a dynamic criterion has to be applied to decide when the analysis has produced the best results obtainable from the input data. We propose an appropriate criterion and show how the requirements for robust interpretation of speech can be met in a dialog system. Keywords: Speech Understanding, Robustness, Search 1. INTRODUCTION In speech understanding systems a main problem is how to deal with poor segmentation results. How can it be decided that no or at least only a partial interpretation of the input data is possible? Let us consider the following two dialogs where the user's utterance (U), the word recognition result (R), and the...
Gernot A. Fink, Franz Kummert, Gerhard Sagerer, Bernd Seestaedt
EUROSPEECH1
1993 Speech recognition using semantic hidden Markov networks
abstract
Semantic Hidden Markov Networks (SHMNs) were first introduced in [2] as a new technique of interfacing between linguistic analysis and word recognition in speech understanding systems. The main difference between SHMNs and the use of traditional language models is that SHMNs always refer to a linguistic concept and impose the linguistic structure as closely as possible on its acoustic counterpart --- a hierarchically structured HMM. Normally the result of decoding a HMM is merely the sequence of best fitting elementary acoustic concepts, e.g. phonemes or words. Taking into account the structure of the recognition task a structured instance can be computed. This complex acoustic instance can easily be transformed into a linguistic instance by a recursive computation but without any searching. In this paper we present an algorithm for generating linguistic instances from word recognition results based on SHMNs. Additionally, we present recognition results obtained when evaluating a set o...
Gernot A. Fink, Franz Kummert, Gerhard Sagerer, Ernst Günter Schukat-Talamazzini
EUROSPEECH1
1993 Modeling of time constituents for speech understanding
abstract
The analysis and interpretation of time constituents is important for most applications of speech understanding systems. Problems can be caused by the varying distribution of constituents. A basic set of time constituents were found in a corpus of domain specific (train schedule) utterances. A distributed representation of surface structure models and an incremental semantic analysis is used to manage the complexity. The knowledge base of the speech understanding system that provides the framework for the analysis and interpretation of time constituents uses the semantic network language ERNEST. Keywords: Speech Understanding, Time Constituents 1. INTRODUCTION In a speech understanding system there are several domains of analysis and interpretation. Firstly, the system has to recognize single words in a torrent of speech sounds. Secondly, the system combines words to constituents, i.e. it performs a syntactic analysis. It also has to reconstruct the meaning of the utterance in questio...
Bernd Hildebrandt, Gernot A. Fink, Franz Kummert, Gerhard Sagerer
EUROSPEECH2
1992 Semantic hidden Markov networks
abstract
Although much effort has been put into speech understanding systems there still exists a rather wide gap between acoustic recognition and linguistic interpretation. We propose a formalism for an extremely close interaction of acoustic recognition and higher level analysis. Instead of a strict horizontal interface at the level of hypothesized word sequences or lattices, a vertical interface to the acoustic component is used that can be accessed from linguistic concepts of any degree of abstraction. As the linguistic knowledge is represented in the formalism of Semantic Networks and acoustic recognition is based on Hidden Markov Models the close interaction between the two components was termed Semantic Hidden Markov Networks. 1 INTRODUCTION Because of the high degree of uncertainty in the recognition of spoken language it is very important to exploit any possible predictions and restrictions to guide acoustic analysis as well as linguistic interpretation. Within traditional systems the...
Gernot A. Fink, Franz Kummert, Gerhard Sagerer, Ernst Günter Schukat-Talamazzini, Heinrich Niemann
ICSLP1