Robert Sablatnig

dblp:52/3025 · DBLP profile ↗
← Back
82ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0003-4195-1593ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 31 · 8 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Semantically Stable Image Composition Analysis via Saliency and Gradient Vector Flow Fusion
abstract
Abstract The reliable computational assessment of photographic composition requires features that are discriminative of spatial layout yet robust to semantic content. This paper proposes a low-level representation grounded in the assumption that composition can be understood as the flow of visual attention across geometric structure. We introduce VFCNet, which fuses saliency and edge information into a gradient vector flow (GVF) field. The model computes dual-stream GVF representations, integrates them via attention, and extracts multi-scale flow features with a DINOv3 backbone. VFCNet achieves state-of-the-art performance on the PICD benchmark (CDA-1: 0.683, CDA-2: 0.629), improving by 33.1% and 36.1% over the previous best method. We also show that a simple classifier on self-supervised DINOv3 features substantially outperforms more sophisticated, composition-specialized models. Code is available at https://github.com/ADadras/VFCNet .
Armin Dadras 0003, Robert Sablatnig, Franziska Proksa, Markus Seidl
ICPR (5)2
2026 SeaClips: A Video Dataset for Maritime Object Detection
abstract
Maritime computer vision is a requirement for autonomous surface vehicles and can improve maritime safety if a high level of robustness is achieved. As deep learning dominates the computer vision community, domain-specific datasets are required to obtain well-generalizing and reliable models. However, maritime datasets, especially those containing videos and temporally dense annotations, are still small compared to other domains, such as autonomous driving or generic computer vision datasets. This paper introduces SeaClips, a new maritime video dataset containing 74 videos with an average duration of 14 seconds and with 31k frames in total. Videos were recorded under varying conditions, with three cameras mounted on shore and on boats. SeaClips provides frame-by-frame annotations, encompassing 129k bounding boxes of seven categories, containing vessel and non-vessel classes. SeaClips contributes to a broader coverage of maritime scenarios, leading to more robust computer vision models. Baseline results on the dataset are established by evaluating six image-based models and three models using temporal context, ranging from lightweight models to heavy transformer architectures. It is found that the different scales and shapes at which objects appear in SeaClips pose a challenge to state-of-the-art detectors. SeaClips is accessible for research on maritime obstacle detection at: https://huggingface.co/datasets/SEA-AI/SeaClips.
Franziska Denk, Christian Rankl, Shaban Almouahed, David Moser, Robert Sablatnig
WACV5
2025 Towards the Influence of Text Quantity on Writer Retrieval
Marco Peer, Robert Sablatnig, Florian Kleber
ICDAR (2)2
2025 Few-Shot Segmentation of Historical Maps via Linear Probing of Vision Foundation Models
Rafael Sterzinger, Marco Peer, Robert Sablatnig
ICDAR (3)3
2025 Interactive Object Detection for Tiny Objects in Large Remotely Sensed Images
abstract
This paper highlights the potential of a Human-In-the-Loop (HIL) in interactive object detection methods. Although automation in computer vision is advancing rapidly, certain critical tasks, such as detecting UneXploded Ordnance (UXO), space/marine debris, or the generation of new datasets, require 100% recall and near-perfect precision. These tasks are often performed manually since automatic methods do not achieve the necessary accuracy. However, interactive object detection frameworks can potentially enhance annotation speed while maintaining the recall and accuracy of manual annotation. We propose IRTDETR, an interactive and real-time object detection method for very large imagery to address this. Using either point or bounding box annotations provided by a HIL, it globally relates the full image with the annotator inputs via a cross-attention-like mechanism, employs an attention loss to maximize the classification score based on similarity, and reuses portions of the network outputs during iterative refinements to conserve resources. We conduct experiments on five different datasets (Tiny-DOTA, CHAI, AITOD, SarDET, and COCO) to verify the efficacy of our approach. Our method surpasses existing interactive annotation approaches, achieving a higher mean Average Precision (mAP) with the same number of clicks. Additionally, we validate the annotation efficiency of our method in a user study, demonstrating it is$2.46\times$quicker and asks for only 72% of the task load (NASA-TLX) compared to fully manual annotation. The code is available under https://github.com/mburges-cvl/WACV_IAODF.
Marvin Burges, Sebastian Zambanini, Robert Sablatnig
WACV3
2024 Classical Photometric Stereo in Point Lighting Environments: Error Analysis and Mitigation
abstract
While classical photometric stereo assumes parallel lighting and thus allows for efficient solutions, this assumption is rarely met in practice. In acquisition setups using point-like light sources such as LEDs, direction and intensity of incident light depend on the unknown surface position, which leads to a non-linear problem. Although numerous approaches explicitly accounting for non-parallel lighting are proposed, photometric stereo with the parallel lighting model is widely used in point lighting scenarios. This is done for reasons of computational efficiency and ease of system calibration, and justified with the observation that lighting becomes approximately parallel when light sources are placed in a sufficient distance from the object. However, as no account of the relation between light source distance and reconstruction errors is found in literature, it is not clear which light source distance would be ’sufficient’ for a given application. In this work, we propose an upper bound for mean errors resulting from using parallel-lighting models in point light setups, depending on the distance of light sources and the size of the object imaged. This bound is based on analytical considerations and previous results of photometric stereo error analysis, and validated vie a Monte Carlo simulation. These results are then used to justify an error mitigation strategy via local solutions that is applicable in use cases where an approximate depth is known a priori. The theoretical propositions are demonstrated on real-world data.
Simon Brenner, Robert Sablatnig
3DV2
2024 Maximizing Data Efficiency of HTR Models by Synthetic Text
Markus Muth, Marco Peer, Florian Kleber, Robert Sablatnig
DAS4
2024 SAGHOG: Self-supervised Autoencoder for Generating HOG Features for Writer Retrieval
Marco Peer, Florian Kleber, Robert Sablatnig
ICDAR (2)3
2024 Drawing the Line: Deep Segmentation for Extracting Art from Ancient Etruscan Mirrors
Rafael Sterzinger, Simon Brenner, Robert Sablatnig
ICDAR (3)3
2024 Advancing Handwritten Text Detection by Synthetic Text
Markus Muth, Marco Peer, Florian Kleber, Robert Sablatnig
ICPR (19)4
2024 KaiRacters: Character-Level-Based Writer Retrieval for Greek Papyri
Marco Peer, Robert Sablatnig, Olga Serbaeva Saraogi, Isabelle Marthot-Santaniello
ICPR (19)2
2024 Fusing Forces: Deep-Human-Guided Refinement of Segmentation Masks
Rafael Sterzinger, Christian Stippel, Robert Sablatnig
ICPR (12)3
2024 ECSIC: Epipolar Cross Attention for Stereo Image Compression
abstract
In this paper, we present ECSIC, a novel learned method for stereo image compression. Our proposed method compresses the left and right images in a joint manner by exploiting the mutual information between the images of the stereo image pair using a novel stereo cross attention (SCA) module and two stereo context modules. The SCA module performs cross-attention restricted to the corresponding epipolar lines of the two images and processes them in parallel. The stereo context modules improve the entropy estimation of the second encoded image by using the first image as a context. We conduct an extensive ablation study demonstrating the effectiveness of the proposed modules and a comprehensive quantitative and qualitative comparison with existing methods. ECSIC achieves state-of-the-art performance in stereo image compression on the two popular stereo image datasets Cityscapes and InStereo2k while allowing for fast encoding and decoding.
Matthias Wödlinger, Jan Kotera, Manuel Keglevic, Jan Xu, Robert Sablatnig
WACV5
2023 Towards Writer Retrieval for Historical Datasets
Marco Peer, Florian Kleber, Robert Sablatnig
ICDAR (1)3
2022 SASIC: Stereo Image Compression with Latent Shifts and Stereo Attention
abstract
We propose a learned method for stereo image compression that leverages the similarity of the left and right images in a stereo pair due to overlapping fields of view. The left image is compressed by a learned compression method based on an autoencoder with a hyperprior entropy model. The right image uses this information from the previously encoded left image in both the encoding and decoding stages. In particular, for the right image, we encode only the residual of its latent representation to the optimally shifted latent of the left image. On top of that, we also employ a stereo attention module to connect left and right images during decoding. The performance of the proposed method is evaluated on two benchmark stereo image datasets (Cityscapes and InStereo2K) and outperforms previous stereo image compression methods while being significantly smaller in model size.
Matthias Wödlinger, Jan Kotera, Jan Xu, Robert Sablatnig
CVPR4
2022 Writer Identification and Writer Retrieval Using Vision Transformer for Forensic Documents
Michael Koepf, Florian Kleber, Robert Sablatnig
DAS3
2022 Self-supervised Vision Transformers with Data Augmentation Strategies Using Morphological Operations for Writer Retrieval
Marco Peer, Florian Kleber, Robert Sablatnig
ICFHR3
2022 Writer Retrieval using Compact Convolutional Transformers and NetMVLAD
abstract
This paper presents a method for writer retrieval where embeddings of patches extracted at SIFT keypoint locations are learned by a Compact Convolutional Transformer (CCT), a modified attention-based transformer architecture including convolutions, followed by a NetMVLAD layer and Generalized Max Pooling (GMP) to obtain global page descriptors. We introduce the application of CCTs for writer retrieval and show that they outperform Convolutional Neural Networks (CNNs) used in current State-of-the-Art methods for writer retrieval, namely ResNet18, while at the same time only have one-third of the number of parameters. Additionally, we propose Net-MVLAD, an extension of NetVLAD with multiple vocabularies, to encode information with different vocabulary sizes improving the original NetVLAD. An evaluation of the performance of CCTs compared to ResNet18 is provided on the ICDAR2013 Competition on Writer Identification dataset (ICDAR2013) and CVL dataset. The effect of multiple vocabularies applied within the NetVLAD layer is shown. CCT7 pretrained on CIFAR-100 combined with NetMVLAD achieves 89.3% Mean Average Precision (mAP) on the ICDAR2013 dataset and 96.5% on the CVL dataset.
Marco Peer, Florian Kleber, Robert Sablatnig
ICPR3
2021 Estimating Human Legibility in Historic Manuscript Images - A Baseline
Simon Brenner, Lukas Schügerl, Robert Sablatnig
ICDAR (3)3
2020 Text Baseline Recognition Using a Recurrent Convolutional Neural Network
abstract
The detection of baselines of text is a necessary preprocessing step for many modern methods of automatic handwriting recognition. In this work, we present a two-stage system for the automatic detection of text baselines of handwritten text. In a first step, we perform pixel-wise segmentation on the document image to classify pixels as baselines, start points, end points and background. This segmentation is then used to extract the start points of lines. Starting from these points we extract the baseline using a recurrent convolutional neural network that directly outputs the baseline coordinates. This method allows the direct extraction of baseline coordinates as the output of a neural network without the use of any post-processing steps. We evaluate the model on the cBAD dataset from the ICDAR 2019 competition on baseline detection.
Matthias Wödlinger, Robert Sablatnig
ICPR2
2019 cBAD: ICDAR2019 Competition on Baseline Detection
abstract
Baseline detection is a simplified text-line extraction that typically serves as pre-processing for Automated Text Recognition. The cBAD competition benchmarks state-of-the-art baseline detection algorithms. It is the successor of cBAD 2017 with a larger dataset that contains more diverse document pages. The images together with the manually annotated groundtruth are made publicly available which allows other teams to benchmark and compare their methods. We could also evaluate the winning method of cBAD 2017 on the newly introduced dataset which now serves as baseline. This competition shows that the performance of automated baseline detection increased substantially since 2017.
Markus Diem, Florian Kleber, Robert Sablatnig, Basilios Gatos
ICDAR3
2019 CNN Based Binarization of MultiSpectral Document Images
abstract
This work is concerned with the binarization of ancient manuscripts that have been imaged with a MultiSpectral Imaging (MSI) system. We introduce a new dataset for this purpose that is composed of 130 multispectral images taken from two medieval manuscripts. We propose to apply an end-to-end Convolutional Neural Network (CNN) for the segmentation of the historical writings. The performance of the CNN based method is superior compared to two state-of-the-art methods that are especially designed for multispectral document images. The CNN based method is also evaluated on a previous and smaller database, where its performance is slightly worse than the two state-of-the-art techniques.
Fabian Hollaus, Simon Brenner, Robert Sablatnig
ICDAR3
2018 MultiSpectral Image Binarization using GMMs
abstract
MultiSpectral Imaging enhances the study of degraded historical documents. It allows for visualizing washed out or even invisible ink but also improves the automated analysis because of a denser spectral sampling. We present a new methodology for binarization of multispectral document images that groups spectral signatures of different sources by fitting two Gaussian Mixture Models (GMMs) with Expectation Maximization. Both GMMs assign cluster labels to the multispectral samples and the clustering results are combined for the identification of the handwriting regions. The method is evaluated on the ICDAR 2015 MS-TEx dataset. Results on this publicly available benchmarking set are encouraging.
Fabian Hollaus, Markus Diem, Robert Sablatnig
ICFHR3
2018 Learning Features for Writer Retrieval and Identification using Triplet CNNs
abstract
This paper presents a method for writer retrieval and identification using a feature descriptor learned by a Convolutional Neural Network. Instead of using a network for classification, we propose the use of a triplet network that learns a similarity measure for image patches. Patches of the handwriting are extracted and mapped into an embedding where this similarity measure is defined by the L2distance. The triplet network is trained by maximizing the interclass distance, while minimizing the intraclass distance in this embedding. The image patches are encoded using the learned feature descriptor. By applying the Vector of Locally Aggregated Descriptors encoding to these features, we generate a feature vector for each document image. A detailed parameter evaluation is given which shows that this method achieves a mean average precision of 86.1% on the ICDAR 2013 writer identification dataset, but future work has to be done to improve the performance on historic datasets. In addition, the strategy for clustering the feature space is investigated.
Manuel Keglevic, Stefan Fiel, Robert Sablatnig
ICFHR3
2018 Word Beam Search: A Connectionist Temporal Classification Decoding Algorithm
abstract
Recurrent Neural Networks (RNNs) are used for sequence recognition tasks such as Handwritten Text Recognition (HTR) or speech recognition. If trained with the Connectionist Temporal Classification (CTC) loss function, the output of such a RNN is a matrix containing character probabilities for each time-step. A CTC decoding algorithm maps these character probabilities to the final text. Token passing is such an algorithm and is able to constrain the recognized text to a sequence of dictionary words. However, the running time of token passing depends quadratically on the dictionary size and it is not able to decode arbitrary character strings like numbers. This paper proposes word beam search decoding, which is able to tackle these problems. It constrains words to those contained in a dictionary, allows arbitrary non-word character strings between words, optionally integrates a word-level language model and has a better running time than token passing. The proposed algorithm outperforms best path decoding, vanilla beam search decoding and token passing on the IAM and Bentham HTR datasets. An open-source implementation is provided.
Harald Scheidl, Stefan Fiel, Robert Sablatnig
ICFHR3
2017 Retrieval of striated toolmarks using convolutional neural networks
abstract
The authors propose TripNet as method for calculating similarities between striated toolmark images. The objective for this system is detecting and comparing characteristics of the tools while being invariant to varying parameters like angle of attack, substrate material, and lighting conditions. Instead of designing a handcrafted feature extractor customised for this task, the authors propose the use of a convolutional neural network. With the proposed system, one‐dimensional profiles extracted from images of striated toolmarks are mapped into an embedding. The system is trained by minimising a triplet loss function, so that a similarity measure is defined by the distance in this embedding. The performance is evaluated on the NFI Toolmark database containing 300 striated toolmarks of screwdrivers published by the Netherlands Forensic Institute. The system proposed is able to adapt to a large range of angles of attack, achieving a mean average precision of 0.95 for toolmark comparisons with differences in angle of attack of – . Furthermore, four different triplet selection approaches are proposed and their effect on the retrieval of toolmarks from a database of unseen tools is evaluated in detail.
Manuel Keglevic, Robert Sablatnig
IET Comput. Vis.2
2016 MSIO: MultiSpectral Document Image BinarizatIOn
abstract
MultiSpectral (MS) imaging enriches document digitization by increasing the spectral resolution. We present a methodology which detects a target ink in document images by taking into account this additional information. The proposed method performs a rough foreground estimation to localize possible ink regions. Then, the Adaptive Coherence Estimator (ACE), a target detection algorithm, transforms the MS input space into a single gray-scale image where values close to one indicate ink. A spatial segmentation using GrabCut on the target detection's output is computed to create the final binary image. To find a baseline performance, the method is evaluated on the three most recent Document Image Binarization COntests (DIBCO) despite the fact that they only provide RGB images. In addition, an evaluation on three publicly available MS datasets is carried out. The presented methodology achieved the highest performance at the MultiSpectral Text Extraction (MS-TEx) contest 2015.
Markus Diem, Fabian Hollaus, Robert Sablatnig
DAS3
2015 Writer Identification and Retrieval Using a Convolutional Neural Network
Stefan Fiel, Robert Sablatnig
CAIP (2)2
2015 Binarization of MultiSpectral Document Images
Fabian Hollaus, Markus Diem, Robert Sablatnig
CAIP (2)3
2015 Investigation of Ancient Manuscripts based on Multispectral Imaging
abstract
This work is concerned with the digitization and analysis of historical documents. The investigation of the documents has been conducted in three successive interdisciplinary projects. The team involved in the projects consists of philologists, chemists and computer scientists specialized in the field of digital image processing. The manuscripts investigated are partially degraded since they have been infected by mold, are corrupted by background clutter or contain faded-out or even erased writings. Since these degradations impede a transcription by scholars and worsen the performance of automated document image analysis techniques, the documents have been imaged with a portable multispectral imaging system. By using this non-invasive investigation technique, the contrast of the faded out characters can be increased, compared to ordinary white light illumination. Post-processing techniques, such as dimension reduction tools, can be used to gain a further legibility increase. The resulting images are used as a basis for further document analysis methods. These methods have been especially designed for the historical documents investigated and involve Optical Character Recognition and writer identification. This paper presents an overview on selected methods that have been developed in the projects.
Fabian Hollaus, Markus Diem, Stefan Fiel, Florian Kleber, Robert Sablatnig
DocEng5
2015 Regularized single-image super-resolution based on progressive gradient estimation
abstract
Gradient domain optimization is widely used in regularized image super-resolution, in which the gradient of high resolution (HR) is estimated for calculating the regularization energy. In this paper, a progressive gradient estimation (PGE) is proposed. In PGE, the gradient of the reconstructed HR image in the previous round of optimization is taken as the estimated gradient in the current round. Then, the estimated image gradient is progressively improved. When the estimated image gradient converges, a high quality HR image can be reconstructed. Experimental results show that the reconstructed HR images by PGE have good qualitative and quantitative performances.
Lejun Yu, Feng-Xiang Ge, Bo Sun 0006, Jun He 0009, Robert Sablatnig
ICIP6
2014 End-to-End Text Recognition Using Local Ternary Patterns, MSER and Deep Convolutional Nets
abstract
Text recognition in natural scene images is an application for several computer vision applications like licence plate recognition, automated translation of street signs, help for visually impaired people or image retrieval. In this work an end-to-end text recognition system is presented. For detection an AdaBoost ensemble with a modified Local Ternary Pattern (LTP) feature-set with a post-processing stage build upon Maximally Stable Extremely Region (MSER) is used. The text recognition is done using a deep Convolution Neural Network (CNN) trained with backpropagation. The system presented outperforms state of the art methods on the ICDAR 2003 dataset in the text-detection (F-Score: 74.2%), dictionary-driven cropped-word recognition (F-Score: 87.1%) and dictionary-driven end-to-end recognition (F-Score: 72.6%) tasks.
Michael Opitz, Markus Diem, Stefan Fiel, Florian Kleber, Robert Sablatnig
Document Analysis Systems5
2014 Ruling analysis and classification of torn documents
abstract
A ruling classification is presented in this paper. In contrast to state-of-the-art methods which focus on ruling line removal, ruling lines are analyzed for document clustering in the context of document snippet reassembling. First, a background patch is extracted from a snippet at a position which minimizes the inscribed content. A novel Fourier feature is then computed on the image patch. The classification into void, lined and checked is carried out using Support Vector Machines. Finally, an accurate line localization is performed by means of projection profiles and robust line fitting. The ruling classification achieves an F-score of 0.987 evaluated on a dataset comprising real world document snippets. In addition the line removal was evaluated on a synthetically generated dataset where an F-score of 0.931 is achieved. This dataset is made publicly available so as to allow for benchmarking.
Markus Diem, Florian Kleber, Robert Sablatnig
ACM Symposium on Document Engineering3
2014 ICFHR 2014 Competition on Handwritten Digit String Recognition in Challenging Datasets (HDSRC 2014)
abstract
This paper presents the results of the HDSRC 2014 competition on handwritten digit string recognition in challenging datasets organized in conjunction with ICFHR 2014. The general objective of this competition is to identify, evaluate and compare recent developments in Western Arabic digit string recognition with varying length. In addition, this competition introduces two new challenging datasets for benchmarking. We describe competition details including the datasets and evaluation measures used, and give a comparative performance analysis of six (6) participating methods along with a short description of the respective methodologies.
Markus Diem, Stefan Fiel, Florian Kleber, Robert Sablatnig, José M. Saavedra, David Contreras, Juan Manuel Barrios, Luiz Eduardo Soares de Oliveira
ICFHR4
2014 Recognizing Glagolitic Characters in Degraded Historical Documents
abstract
This paper presents a method for the recognition of Glagolitic characters in degraded historical documents. The Glagolitic character recognition is based on Dense SIFT for which image restoration is proposed as a pre-processing step in order to suppress background noise in degraded documents. Two different methods for image restoration are used which are Total Variation regularization and a new restoration method. Each method performs robustly against background noise while preserving character edges and strokes in the documents defected by stain, bleed through, and faded out ink. The experimental results achieved on three datasets show that by using image restoration as a pre-processing step to Dense SIFT generates better recognition rates for Glagolitic characters in degraded documents.
Sajid Saleem, Fabian Hollaus, Markus Diem, Robert Sablatnig
ICFHR4
2014 Improving OCR Accuracy by Applying Enhancement Techniques on Multispectral Images
abstract
This work is concerned with the legibility enhancement of ancient and degraded handwritings. The writings are partially barely visible under normal white light and hence they have been imaged with a MultiSpectral Imaging (MSI) system in order to increase their legibility. Dimension reduction techniques - like Principal Component Analysis (PCA) - can be used to further enhance the contrast of the faded-out characters. In this work the dimensionality of the multispectral scan is lowered, by applying Linear Discriminant Analysis (LDA). Since LDA is a supervised dimension reduction method, it is necessary to label a subset of the multispectral samples as belonging to the fore-or background. For this purpose, an approach is suggested that uses spatial information. The enhancement method is evaluated by Optical Character Recognition (OCR). By applying the enhancement method the OCR performance is increased in the case of degraded writings, compared to OCR results gained on unprocessed multispectral images and to OCR results achieved on images, which have been produced by applying unsupervised dimension reductions.
Fabian Hollaus, Markus Diem, Robert Sablatnig
ICPR3
2014 Robust Skew Estimation of Handwritten and Printed Documents Based on Grayvalue Images
abstract
Skew estimation is a preprocessing step in document image analysis to determine the global dominant orientation of a document's text lines. A skew angle can be introduced during scanning, or if a document is photographed. The correction of the skew angle is necessary for further image analysis, to avoid an influence to the performance of skew sensitive methods, e.g. Optical Character Recognition (OCR) or page segmentation. The performance of current skew estimation methods is shown at the ICDAR2013 Document Image Skew Estimation Contest (DISEC), which uses a benchmark dataset of binarized printed documents with varying layouts and languages like English, Chinese or Greek. The proposed method is based on a Focused Nearest Neighbour Clustering (FNNC) of interest points and the analysis of paragraphs/lines and achieved rank 5 at the contest. In this paper it is shown, that the use of gray value images can outperform the results restricted to binarized images, thus the proposed method avoids the binarization step which is still an open research topic in document image analysis. The robustness of the method is also shown on a dataset comprising historical documents and on low resolution images. The method is evaluated on the DISEC dataset and three additional datasets (historical documents, low resolution documents, and machine printed documents).
Florian Kleber, Markus Diem, Robert Sablatnig
ICPR3
2014 A Gradient Extension of Center Symmetric Local Binary Patterns for Robust RGB-NIR Image Matching
abstract
Scene acquisition using RGB and Near Infra-Red (NIR) filters generates useful visual information about scene contents. But it induces significant intensity and textural changes between RGB and NIR images of the same scene. It becomes a challenging problem to perform interest point based image matching under such intensity and textural changes. To cope with this problem, a novel method for the description of interest points is proposed. The method proposed is based on Center Symmetric-Local Binary Patterns (CS-LBP) which extracts distinct image features from intensity and gradient magnitude maps of the image patches centered at interest points. Those features are then used in the SIFT algorithm to compute robust descriptors against intensity and textural changes. The experimental results show that the method proposed improves the descriptor matching between RGB and NIR images and achieves better image matching results than CS-LBP and SIFT based methods for the description of interest points.
Sajid Saleem, Abdul Bais, Robert Sablatnig
ICPR3
2014 Systematic skin segmentation: merging spatial and non-spatial data
Rehanullah Khan, Allan Hanbury, Robert Sablatnig, Julian Stöttinger, Farman Ali Khan, Fakhri Alam Khan
Multim. Tools Appl.3
2014 A Robust SIFT Descriptor for Multispectral Images
abstract
This letter presents a novel method for the description of multispectral image keypoints. The method proposed is based on a modified SIFT algorithm. It uses normalized gradients as local image features for the description of keypoints in order to achieve robustness against non linear intensity changes between multispectral images. The experimental results show that the method proposed achieves a better matching performance and outperforms the SIFT algorithm.
Sajid Saleem, Robert Sablatnig
IEEE Signal Process. Lett.2
2013 Evaluation of Methods for Optical 3-D Scanning of Human Pinnas
abstract
In the context of computational acoustics, a detailed evaluation of various 3-D scanning methods for the purpose of capturing the geometry of the human pinna (the visible part of the ear) is presented. Over 80 full 3-D scans of the head and ears of three subjects were performed with six different 3-D scanning methods, including photogrammetry with nine different camera and texturing options. We provide a numerical comparison of the scanning performance in terms of accuracy and completeness, and show the effect of ambient occlusion on the accuracy. The numerical and practical issues for scanning the ears for computational acoustics are discussed.
Andreas Reichinger, Piotr Majdak, Robert Sablatnig, Stefan Maierhofer
3DV3
2013 ICDAR 2013 Competition on Handwritten Digit Recognition (HDRC 2013)
abstract
This paper presents the results of the HDRC 2013 competition for recognition of handwritten digits organized in conjunction with ICDAR 2013. The general objective of this competition is to identify, evaluate and compare recent developments in character recognition and to introduce a new challenging dataset for benchmarking. We describe competition details including dataset and evaluation measures used, and give a comparative performance analysis of the nine (9) submitted methods along with a short description of the respective methodologies.
Markus Diem, Stefan Fiel, Angelika Garz, Manuel Keglevic, Florian Kleber, Robert Sablatnig
ICDAR6
2013 Text Line Detection for Heterogeneous Documents
abstract
Text line detection is a pre-processing step for automated document analysis such as word spotting or OCR. It is additionally used for document structure analysis or layout analysis. Considering mixed layouts, degraded documents and handwritten documents, text line detection is still challenging. We present a novel approach that targets torn documents having varying layouts and writing. The proposed method is a bottom up approach that fuses words, to globally minimize their fusing distance. In order to improve processing time and further layout analysis, text lines are represented by oriented rectangles. Even though, the method was designed for modern handwritten and printed documents, tests on medieval manuscripts give promising results. Additionally, the text line detection was evaluated on the ICDAR 2009 and ICFHR 2010 Handwriting Segmentation Contest datasets.
Markus Diem, Florian Kleber, Robert Sablatnig
ICDAR3
2013 Writer Identification and Writer Retrieval Using the Fisher Vector on Visual Vocabularies
abstract
In this paper a method for writer identification and writer retrieval is presented. Writer identification is the task of identifying the writer of a document out of a database of known writers. In contrast to identification, writer retrieval is the task of finding documents in a database according to the similarity of handwritings. The approach presented in this paper uses local features for this task. First a vocabulary is calculated by clustering features using a Gaussian Mixture Model and applying the Fisher kernel. For each document image the features are calculated and the Fisher Vector is generated using the vocabulary. The distance of this vector is then used as similarity measurement for the handwriting and can be used for writer identification and writer retrieval. The proposed method is evaluated on two datasets, namely the ICDAR 2011 Writer Identification Contest dataset which consists of 208 documents from 26 writers, and the CVL Database which contains 1539 documents from 309 writers. Experiments show that the proposed methods performs slightly better than previously presented writer identification approaches.
Stefan Fiel, Robert Sablatnig
ICDAR2
2013 Enhancement of Multispectral Images of Degraded Documents by Employing Spatial Information
abstract
This work aims at enhancing ancient and degraded writings, which are captured by MultiSpectral Imaging systems. The manuscripts captured, contain faded out characters and are partly corrupted by mold and hardly legible. Several works have shown that such writings can be enhanced by applying unsupervised dimension reduction tools - like Principal Component Analysis (PCA) or Independent Component Analysis (ICA). In this work the Fisher Linear Discriminate Analysis (LDA) is applied in order to reduce the dimension of the multispectral scan and to enhance the degraded writings. Since Fisher LDA is a supervised dimension reduction tool, it is necessary to label a subset of multispectral data. For this purpose, a semi-automated label generation step is conducted, which is based on an automated detection of text lines. Thus, the approach is not only based on spectral information - like PCA and ICA - but also on spatial information. The method has been tested on two Slavonic manuscripts. A qualitative analysis shows, that the LDA based dimension reduction gains better performance, compared to unsupervised techniques.
Fabian Hollaus, Melanie Gau, Robert Sablatnig
ICDAR3
2013 CVL-DataBase: An Off-Line Database for Writer Retrieval, Writer Identification and Word Spotting
abstract
In this paper a public database for writer retrieval, writer identification and word spotting is presented. The CVL-Database consists of 7 different handwritten texts (1 German and 6 English Texts) and 311 different writers. For each text an RGB color image (300 dpi) comprising the handwritten text and the printed text sample are available as well as a cropped version (only handwritten). A unique ID identifies the writer, whereas the bounding boxes for each single word are stored in an XML file. An evaluation of the best algorithms of the ICDAR and ICHFR writer identification contest has been performed on the CVL-database.
Florian Kleber, Stefan Fiel, Markus Diem, Robert Sablatnig
ICDAR4
2013 Lifelogging keyframe selection using image quality measurements and physiological excitement features
abstract
Keyframe selection is the process of finding a representative frame in an image sequence. Although mostly known from video processing, keyframe selection faces new challenges in the lifelog domain. To obtain a keyframe that is close to a user-selected frame, we propose a keyframe selection method based on image quality measurements and excitement features. Image quality measurements such as contrast, color variance, sharpness, noise and saliency are used to filter high quality images. However, high quality images are not necessarily keyframes because humans also use emotions in the selection process. In this study, we employ a biosensor to measure the excitement of humans. In previous investigation, keyframe selection using only image quality measurements yielded an acceptance rate of 79.70%. Our proposed method achieves an acceptance rate of 84.45%.
Photchara Ratsamee, Yasushi Mae, Amornched Jinda-Apiraksa, Jana Machajdik, Kenichi Ohara, Masaru Kojima, Robert Sablatnig, Tatsuo Arai
IROS7
2012 Segmentation of Building Facade Domes
Gayane Shalunts, Yll Haxhimusa, Robert Sablatnig
CIARP3
2012 Skew Estimation of Sparsely Inscribed Document Fragments
abstract
Document analysis is done to analyze entire forms (e.g. intelligent form analysis, table detection) or to describe the layout/structure of a document for further processing. A pre-processing step of document analysis methods is a skew estimation of scanned or photographed documents. Current skew estimation methods require the existence of large text areas, are dependent on the text type and can be limited on a specific angle range. The proposed method is gradient based in combination with a Focused Nearest Neighbor Clustering of interest points and has no limitations regarding the detectable angle range. The upside/down decision is based on statistical analysis of ascenders and descenders. It can be applied to entire documents as well as to document fragments containing only a few words. Results show that the proposed skew estimation is comparable with state-of-the-art methods and outperforms them on a real dataset consisting of 658 snippets.
Markus Diem, Florian Kleber, Robert Sablatnig
Document Analysis Systems3
2012 Writer Retrieval and Writer Identification Using Local Features
abstract
Writer identification determines the writer of one document among a number of known writers where at least one sample is known. Writer retrieval searches all documents of one particular writer by creating a ranking of the similarity of the handwriting in a dataset. This paper presents a method for writer retrieval and writer identification using local features and therefore the proposed method is not dependent on a binarization step. First the local features of the image are calculated and with the help of a predefined codebook an occurrence histogram can be created. This histogram is compared to determine the identity of the writer or the similarity of other handwritten documents. The proposed method has been evaluated on two datasets, namely the IAM dataset which contains 650 writers and the Trigraph Slant dataset which contains 47 writers. Experiments have shown that it can keep up with previous writer identification approaches. Regarding writer retrieval it outperforms previous methods.
Stefan Fiel, Robert Sablatnig
Document Analysis Systems2
2012 Binarization-Free Text Line Segmentation for Historical Documents Based on Interest Point Clustering
abstract
Segmenting page images into text lines is a crucial pre-processing step for automated reading of historical documents. Challenging issues in this open research field are given \eg by paper or parchment background noise, ink bleed-through, artifacts due to aging, stains, and touching text lines. In this paper, we present a novel binarization-free line segmentation method that is robust to noise and copes with overlapping and touching text lines. First, interest points representing parts of characters are extracted from gray-scale images. Next, word clusters are identified in high-density regions and touching components such as ascenders and descenders are separated using seam carving. Finally, text lines are generated by concatenating neighboring word clusters, where neighborhood is defined by the prevailing orientation of the words in the document. An experimental evaluation on the Latin manuscript images of the Saint Gall database shows promising results for real-world applications in terms of both accuracy and efficiency.
Angelika Garz, Andreas Fischer 0002, Robert Sablatnig, Horst Bunke
Document Analysis Systems3
2011 Text Classification and Document Layout Analysis of Paper Fragments
abstract
In general document image analysis methods are pre-processing steps for Optical Character Recognition (OCR) systems. In contrast, the proposed method aims at clustering document snippets, so that an automated clustering of documents can be performed. Therefore, words are classified according to printed text, manuscripts, and noise. Where, the third class corrects falsely segmented background elements. Having classified text elements, a layout analysis is carried out which groups words into text lines and paragraphs. A back propagation of the class weights - assigned to each word in the first step - enables correcting wrong class labels. The proposed method shows promising results on a dataset consisting of document snippets with varying shapes, content writing and layout. In addition, the system is compared to page segmentation methods of the ICDAR 2009 Page Segmentation Competition.
Markus Diem, Florian Kleber, Robert Sablatnig
ICDAR3
2011 Layout Analysis for Historical Manuscripts Using Sift Features
abstract
We propose a layout analysis method for historical manuscripts that relies on the part-based identification of layout entities. A layout entity -- such as letters of the text, initials or headings -- is composed of a set of characteristic segments or structures, which is dissimilar for distinct classes in the manuscripts under consideration. This fact is exploited in order to segment a manuscript page into homogeneous regions. Historical documents traditionally involve challenges such as uneven writing support and varying shapes of characters, fluctuating text lines, changing scripts and writing styles, and variance in the layout itself. Hence, a part-based detection of layout entities is proposed using a multi-stage algorithm for the localization of the entities, based on interest points. Results show that the proposed method is able to locate initials, headings and text areas in ancient manuscripts containing stains, tears and partially faded-out ink sufficiently well.
Angelika Garz, Robert Sablatnig, Markus Diem
ICDAR2
2011 Scale Space Binarization Using Edge Information Weighted by a Foreground Estimation
abstract
The proposed binarization algorithm uses a scale space to avoid the estimation of script size dependent parameters. Due to the continous smoothing from finer to coarse scales, noise such as background clutter is suppressed since coarse scales characterize homogeneous regions of the image. Thus, coarser scales of the scale space can be used as a foreground estimation to apply a weigthing scheme robust against noise present in, for instance carbon copies or ancient and degraded documents. Additionally the information of filled regions is propagated through the scales. The use of integral images for the calculation of the mean, standard deviation and morphological operations allow for an efficient implementation of the method presented. The binarization of each scale is based on changes of the local intensity as proposed by Su et al.
Florian Kleber, Markus Diem, Robert Sablatnig
ICDAR3
2010 Document analysis applied to fragments: feature set for the reconstruction of torn documents
abstract
Document analysis is done to analyze entire forms (e.g. intelligent form analysis, table detection) or to describe the layout/structure of a document. In this paper document analysis is applied to snippets of torn documents to calculate features that can be used for reconstruction. The main intention is to handle snippets of varying size and different contents (e.g. handwritten or printed text). Documents can either be destroyed by the intention to make the printed content unavailable (e.g. business crime) or due to time induced degeneration of ancient documents (e.g. bad storage conditions). Current reconstruction methods for manually torn documents deal with the shape, or e.g. inpainting and texture synthesis techniques. In this paper the potential of document analysis techniques of snippets to support a reconstruction algorithm by considering additional features is shown. This implies a rotational analysis, a color analysis, a line detection, a paper type analysis (checked, lined, blank) and a classification of the text (printed or hand written). Preliminary results show that these features can be determined reliably on a real dataset consisting of 690 snippets.
Markus Diem, Florian Kleber, Robert Sablatnig
Document Analysis Systems3
2010 Higher order MRF for foreground-background separation in multi-spectral images of historical manuscripts
abstract
Multi-spectral imaging for the analysis and preservation of ancient documents has gained high attention in recent years. While readability enhancement is based on the multi-spectral image corpus, foreground-background separation still relies mainly on gray level or color images. In this paper we propose a foreground-background separation algorithm designed for multi-spectral images. The main contribution is the simultaneously utilization of spectral and spatial features. While spectral features incorporate the spectral components of the multi-spectral images, the spatial features are based on stroke properties. Higher order Markov Random Fields enables an efficient way to combine both features. To solve higher order energy functions, we introduce a new message update rule in the well known belief propagation algorithm based on a higher order potential function.
Martin Lettner, Robert Sablatnig
Document Analysis Systems2
2010 Are Characters Objects?
abstract
This paper presents a character recognition system that handles degraded manuscript documents like the ones discovered at the St. Catherine's Monastery. In contrast to state-of-the-art OCR systems, no early decision (image binarization) needs to be performed. Thus, an object recognition methodology is adapted for the recognition of ancient manuscripts. The proposed system is based on local descriptors which are clustered in order to localize characters. Finally, a class probability histogram is assigned to each character present in an image which allows for the character classification. The system achieves an F0.5 score of 0.77 on real world data that contains 13.5% highly degraded characters.
Markus Diem, Robert Sablatnig
ICFHR2
2010 Detecting Text Areas and Decorative Elements in Ancient Manuscripts
abstract
An approach for the detection of decorative elements - such as initials and headlines - and text regions, focused on ancient manuscripts, is presented. Due to their age, ancient manuscripts suffer from degradation and staining as well as ink is faded-out over the time. Identifying decorative elements and text regions allows indexing a manuscript and serves as input for Optical Character Recognition (OCR) as it localizes regions of interest within document pages. We propose a robust method inspired by state-of-the-art object recognition methodologies. Scale Invariant Feature Transform (SIFT) descriptors are chosen to detect the regions of interest, and the scale of the interest points is used for localization. The classification is based on the fact that local properties of the decorative elements are different to those of regular text. The results show that the method is able to locate regular text in ancient manuscripts. The detection rate of decorative elements is not as high as for regular text but already yields to promising results.
Angelika Garz, Markus Diem, Robert Sablatnig
ICFHR3
2010 Combining Spectral and Spatial Features for Robust Foreground-Background Separation
abstract
Foreground-background separation in multispectral images of damaged manuscripts can benefit from both, spectral and spatial information. Therefore, we incorporate a Markov Random Field which provides a powerful tool to combine both features simultaneously. Higher order models enable the inclusion of spatial constraints based on stroke characteristics. We apply belief propagation for inference and include the higher order potentials by upgrading the message update. The proposed segmentation method requires no training and is independent of script, size, and style of characters. We will demonstrate the robust performance on a set of degraded documents and on synthetic images.
Martin Lettner, Robert Sablatnig
ICPR2
2010 Automatic image-based assessment of lesion development during hemangioma follow-up examinations
Sebastian Zambanini, Robert Sablatnig, Harald Maier, Georg Langs
Artif. Intell. Medicine2
2009 Recognition of Degraded Handwritten Characters Using Local Features
abstract
The main problems of Optical Character Recognition (OCR) systems are solved if printed latin text is considered. Since OCR systems are based upon binary images, their results are poor if the text is degraded. In this paper a codex consisting of ancient manuscripts is investigated. Due to environmental effects the characters of the analyzed codex are washed out which leads to poor results gained by state of the art binarization methods. Hence, a segmentation free approach based on local descriptors is being developed. Regarding local information allows for recognizing characters that are only partially visible. In order to recognize a character the local descriptors are initially classified with a Support Vector Machine (SVM) and then identified by a voting scheme of neighboring local descriptors. State of the art local descriptor systems are evaluated in this paper in order to compare their performance for the recognition of degraded characters.
Markus Diem, Robert Sablatnig
ICDAR2
2009 A Survey of Techniques for Document and Archaeology Artefact Reconstruction
abstract
An automated assembling of shredded/torn documents (2D) or broken pottery (3D) will support philologists, archaeologists and forensic experts. An automated solution for this task can be divided into shape based matching techniques (apictorial) or techniques that analyze additionally the visual content of the fragments (pictorial). In the case of visual content techniques like texture based analysis are used. Depending on the application, shape matching techniques are suitable for entities of the puzzle problem with small numbers of pieces (e.g. up to 20). Also artefacts like broken and lost pieces or overlapping parts of fragments increase the error rate of shape based techniques since the matching of adjacent boundaries can fail. As a result additional features, e.g. color, document structure, have to be used. This paper presents an overview about current puzzle applications in Cultural Heritage, and introduces also the main problems in puzzle solving.
Florian Kleber, Robert Sablatnig
ICDAR2
2009 Spatial and Spectral Based Segmentation of Text in Multispectral Images of Ancient Documents
abstract
In this paper we propose a character segmentation method for multispectral images of ancient documents. Due to the low quality of the images the main idea of this study is to combine the multispectral behavior and contextual spatial information. Therefore we utilize a Markov random field model using the spectral information of the images and stroke properties to include spatial dependencies of the characters. Since the stroke properties and the Gaussian parameters for the imaging model are evaluated automatically the proposed segmentation method requires no training phase. We compared the method to state of the art character segmentation methods and demonstrate the effectiveness of combining spectral and spatial features for the segmentation of characters in multispectral images.
Martin Lettner, Robert Sablatnig
ICDAR2
2008 Contrast Enhancement in Multispectral Images by Emphasizing Text Regions
abstract
This paper deals with the enhancement of the readability in historic texts written on parchment. Due to mold, air, humidity, water, etc. parchment and text are partially damaged and consequently hard to read. In order to enhance the readability of the text, the manuscript pages are imaged in different spectral bands ranging from 360 to 1000 nm. The readability enhancement is based on a spectral and spatial analysis of the multivariate image data by multivariate spatial correlation. The main advantage of the method is that especially the text regions are enhanced which is provided by generating a mask image. This mask is based on the automatic reconstruction of the ruling scheme of the text pages. The method is tested on two medieval Slavonic manuscripts written on parchment.
Martin Lettner, Florian Kleber, Robert Sablatnig, Heinz Miklas
Document Analysis Systems3
2008 Ancient document analysis based on text line extraction
abstract
In order to preserve our cultural heritage and for automated document processing libraries and national archives have started digitizing historical documents. In the case of degraded manuscripts (e.g. by mold, humidity, bad storage conditions) the text or parts of it can disappear. The remaining parts of the text can be segmented and the ruling can be extrapolated with the a priori knowledge. Since the ruling defines the position of the text within a page, it can be used for layout analysis and as a basis for the enhancement of the readability. Furthermore, information about the scribe (hand) of the manuscript, its spatiotemporal origin can be gained by analyzing the ruling. This paper presents an algorithm for ruling estimation of Glagolitic texts based on text line extraction and is suitable for degraded manuscripts by extrapolating the baselines with the a priori knowledge of the ruling. The algorithm was tested on 30 pages of the Missale Sinaiticum and the evaluation was based on visual criteria.
Florian Kleber, Robert Sablatnig, Melanie Gau, Heinz Miklas
ICPR2
2008 Automated stroke ending analysis for drawing tool classification
abstract
This paper proposes a drawing tool recognition method based on features calculated from the shape of stroke endings. The application for this method is to help art historians to identify the drawing tool used for a drawing. Since the style of a drawing depends on the drawing tool used, drawing tool recognition is an important step toward a style analysis. A dominant feature of a drawn stroke is its ending. Several features regarding curvature, proportions etc. are calculated out of the shape of the endings. These features are then used to classify stroke endings with a SVM classifier.
Maria Christine Vill, Robert Sablatnig
ICPR2
2007 Identification of drawing tools by classification of textural and boundary features of strokes
Paul Kammerer, Martin Lettner, Ernestine Zolda, Robert Sablatnig
Pattern Recognit. Lett.4
2007 Rule based system for archaeological pottery classification
Martin Kampel, Robert Sablatnig
Pattern Recognit. Lett.2
2006 Landmark Based Global Self-localization of Mobile Soccer Robots
Abdul Bais, Robert Sablatnig
ACCV (2)2
2006 Special issue on 3D acquisition technology for cultural heritage
Luc Van Gool, Robert Sablatnig
Mach. Vis. Appl.2
2005 Investigation on traditional and modern ceramic documentation
abstract
Archaeology is at a point where it can benefit greatly from the application of computer vision methods, and in turn provides a large number of new, challenging and interesting conceptual problems and data for computer science. This is true in particular in the study of ceramics - the most abundant and widespread of all archaeological finds. The traditional way of documenting archaeological sherds is to draw the profile line, which is the intersection of a sherd along the axis of symmetry. A profilograph is a mechanical device, which can directly acquire and transfer a profile line by pin-pointing the profile on a sherd to a computer. We developed a fully automated vision system, which is able to compute the profile line out of the acquired 3D model of the fragment. In this paper we want to give a thorough comparison between the traditional manual approach, the profilograph and our system and present an improvement of the robustness of our approach by finding circular rills on the fragments. Practical experiments have been undertaken at the excavation Tel Dor in Israel.
Martin Kampel, Hubert Mara, Robert Sablatnig
ICIP (2)3
2003 Stroke Boundary Analysis for Identification of Drawing Tools
Paul Kammerer, Georg Langs, Robert Sablatnig, Ernestine Zolda
CIARP3
2003 An automated pottery archival and reconstruction system
abstract
Abstract Motivated by the current requirements of archaeologists, we are developing an automated archival system for archaeological classification and reconstruction of ceramics. Our system uses the profile of an archaeological fragment, which is the cross‐section of the fragment in the direction of the rotational axis of symmetry, to classify and reconstruct it virtually. Ceramic fragments are recorded automatically by a 3D measurement system based on structured (coded) light. The input data for the estimation of the profile is a set of points produced by the acquisition system. By registering the front and the back views of the fragment the profile is computed and measurements like diameter, area percentage of the complete vessel, height and width are derived automatically. We demonstrate the method and give results on synthetic and real data. Copyright © 2003 John Wiley & Sons, Ltd.
Martin Kampel, Robert Sablatnig
Comput. Animat. Virtual Worlds2
2002 Model-Based Registration of Front- and Backviews of Rotationally Symmetric Objects
Robert Sablatnig, Martin Kampel
Comput. Vis. Image Underst.1
2000 Color Classification of Archaeological Fragments
abstract
We are developing an automated classification and reconstruction system for archaeological fragments. The goal is to relate different fragments belonging to the same vessel based on shape, material and color, thus the color information is important in the pre-classification process. In this work a color specification technique is proposed, which exploits the fact that the spectral reflectance of materials like archaeological fragments vary slowly. We explain how the acquisition system is calibrated in order to get accurate colorimetric information with respect to archaeological requirements. Experimental results are presented for archaeological objects and for a set of test color patches.
Martin Kampel, Robert Sablatnig
ICPR2
2000 Estimating the Next Sensor Position Based on Surface Characteristics
abstract
In order to reconstruct the viewable surface of an object completely, multiple views of the same object have to be used and integrated into a common coordinate system. One of the major problems of the 3D surface reconstruction using a turntable, is the varying resolution in the direction to the camera, due to the varying distance of object points to the rotational axis of the turntable. To guarantee a uniform object resolution, we calculate the next angle dynamically, depending on the entropy of the surface part actually acquired. To minimize the loss of information and to guarantee a uniform surface resolution, we derive a relation between the entropy and the next viewing angle, based on the profile sections acquired in the last two steps of the acquisition.
Christian Liska, Robert Sablatnig
ICPR2
2000 Increasing flexibility for automatic visual inspection: the general analysis graph
Robert Sablatnig
Mach. Vis. Appl.1
1999 On Registering Front- and Backviews of Rotationally Symmetric Objects
Robert Sablatnig, Martin Kampel
CAIP1
1998 Hierarchical classification of paintings using face- and brush stroke models
abstract
It is often difficult to attribute works of art to a certain artist. In the case of paintings, radiological methods like X-ray and infra-red diagnosis, digital radiography, computer-tomography, etc. and color analyzes are employed to authenticate works of art. But all these methods do not relate certain characteristics of an art work to a specific artist-the artist's personal style. In order to study this personal style, we examine the "structural signature" based on brush strokes in particular in portrait miniatures. A computer-aided classification and recognition system for portrait miniatures is developed, which enables a semi-automatic classification based on brush strokes. A hierarchically structured classification scheme is introduced which separates the classification into three different levels of information: color, shape of region, and structure of brush strokes.
Robert Sablatnig, Paul Kammerer, Ernestine Zolda
ICPR1
1996 Flexible automatic visual inspection based on the separation of detection and analysis
abstract
Since the mid-1970s, a large number of visual inspection systems and algorithms for industrial inspection have been developed. To be acceptable in industry, vision systems must be inexpensive, within the speed of the production-line flow, and very accurate. Furthermore the system should be flexible enough to accommodate changes in products. This flexibility can only be achieved by a modular concept that allows a quick and inexpensive adaption of the inspection process to changes in production. This paper presents a concept for visual inspection where the detection of primitives is separated from the model-based analysis process. Existing pattern recognition software is re-used in the detection stage and therefore the use of any detection algorithm is possible without changing the analysis process. The visual inspection of analog display measuring instruments serves as a demonstration of this concept. Results concerning time, accuracy, and reliability for the specific inspection task are given at the end of the paper.
Robert Sablatnig
ICPR1
1994 Automatic reading of analog display instruments
abstract
A general design strategy based on experiences of a successful application project concerning the reading of analog devices in order to automate the reading process is presented. The problem space is defined and a description language is introduced which controls the interpretation of specific types of instruments and decouples recognition of primitiva and formation of an aspired result. Furthermore, the interaction between Hough transform or line detection, the user-controlled computation of the measurement and the extension to other measuring instruments is described.
Robert Sablatnig, Walter G. Kropatsch
ICPR (1)1
1994 Application constraints in the design of an automatic reading device for analog display instruments
abstract
An analysis system design based on experience with a successful application in the field of inspection and calibration of an analog display measuring instrument is presented in this paper. First the measuring instrument is divided into its primitiva, defining the a priori known parameter of the primitiva: shape, relative position and size. According to the shape of the primitiva pattern recognition algorithms are used to detect the primitiva in intensity images. These independent detection algorithms are then grouped into a detecting order with respect to efficiency. Following a discussion of the general design of the detecting algorithm, specific constraints of the application and the industrial environment are considered in order to refine the general design to an applicable and efficient device by modifying both hardware and software configuration depending on the given constraints. Finally, results of the implementation of the algorithm and the constructed image acquisition device are discussed.>
Robert Sablatnig, Walter G. Kropatsch
WACV1