Andreas Dengel 0001

dblp:d/AndreasDengel · also Andreas R. Dengel · DBLP profile ↗
← Back
330ranked-venue papers
16as first author
113since 2021 · last 2026
0000-0002-6100-8255ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 204 · 14 first-author · 79 since 2021Databases, data management, data science and information retrieval · 113 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 2 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 17 since 2021Human-computer interaction and ubiquitous computing · 16 · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Can Generative Adversarial Networks Compete against Diffusion Models at Generating Time Series?
Philipp Engler, Ludger van Elst, Sheraz Ahmed, Andreas Dengel 0001
ICAART (4)4
2026 A Study in Dataset Distillation for Image Super-Resolution
Tobias Dietz, Brian B. Moser, Tobias Christian Nauen, Federico Raue, Stanislav Frolov, Andreas Dengel 0001
ICPR (1)6
2026 HyperCore: Coreset Selection Under Noise via Hypersphere Models
Brian B. Moser, Arundhati S. Shanbhag, Tobias Nauen, Stanislav Frolov, Federico Raue, Joachim Folz, Andreas Dengel 0001
ICPR (1)7
2026 A Low-Resolution Image is Worth 1 ˟ 1 Words: Enabling Fine Image Super-Resolution with Transformers and TaylorShift
Sanath Budakegowdanadoddi Nagaraju, Brian B. Moser, Tobias Christian Nauen, Stanislav Frolov, Federico Raue, Andreas Dengel 0001
ICPR (1)6
2026 FREQuency ATTribution: benchmarking frequency-based occlusion for time series data
Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed
Appl. Intell.2
2026 Comic: explainable drug repurposing via contrastive masking for interpretable connections
abstract
Many diseases worldwide remain untreated due to the slow and expensive process of drug development. Repurposing existing FDA-approved drugs offers a faster solution, especially with the assistance of artificial intelligence. Despite advancements in AI-driven drug repurposing, current approaches either have lackluster performance or fail to highlight the intricate pathways through which drugs act on diseases. The clinical utility of AI-driven drug repurposing remains constrained by these limitations, particularly for rare and undertreated diseases where data is scarce. To address the need for a precise and explainable predictor, this paper introduces COMIC (COntrastive Masking with Interpretable Connections), a predictor that employs a multi channel architecture consisting of a feature masking branch, which identifies critical drug-disease interaction patterns by extracting the most informative features, and a path masking branch, which highlights relevant biological pathways through which drugs exert their therapeutic effects. Comprehensive evaluation of the COMIC predictor on the PrimeKG knowledge graph (comprising 17,080 diseases, and 4 M+ relationships) with nine distinct disease area splits demonstrated a 9.55% average performance improvement over the current state-of-the-art. The practical applicability of the proposed predictor is evaluated on a set of the most recent 30 FDA-approved repurposed drug disease pairs. The COMIC predictor successfully identified 21 of these pairs with high confidence scores. To facilitate real-time drug repurposing investigations, we have developed a publicly available web-based interface for the COMIC predictor ( https://sds-genetic-interaction-analysis.opendfki.de/drug_prediction/ ). This application takes disease names as input and returns a ranked list of potential repurposing candidates, along with predicted mechanistic pathways elucidating the drug-disease interactions.
Naafey Aamer, Muhammad Nabeel Asim, Andreas Dengel 0001
BMC Bioinform.3
2025 TKG-DM: Training-free Chroma Key Content Generation Diffusion Model
abstract
Diffusion models have enabled the generation of high-quality images with a strong focus on realism and textual fidelity. Yet, large-scale text-to-image models, such as Stable Diffusion, struggle to generate images where foreground objects are placed over a chroma key background, limiting their ability to separate foreground and background elements without fine-tuning. To address this limitation, we present a novel Training-Free Chroma Key Content Generation Diffusion Model (TKG-DM), which optimizes the initial random noise to produce images with foreground objects on a specifiable color background. Our proposed method is the first to explore the manipulation of the color aspects in initial noise for controlled background generation, enabling precise separation of foreground and background without fine-tuning. Extensive experiments demonstrate that our training-free method outperforms existing methods in both qualitative and quantitative evaluations, matching or surpassing fine-tuned models. Finally, we successfully extend it to other tasks (e.g., consistency models and text-to-video), highlighting its transformative potential across various generative applications where independent control of foreground and background is crucial.
Ryugo Morita, Stanislav Frolov, Brian B. Moser, Takahiro Shirakawa, Ko Watanabe 0001, Andreas Dengel 0001, Jinjia Zhou
CVPR6
2025 Informed Learning for Estimating Drought Stress at Fine-Scale Resolution Enables Accurate Yield Prediction
abstract
Water is essential for agricultural productivity. Assessing water shortages and reduced yield potential is a critical factor in decision-making for ensuring agricultural productivity and food security. Crop simulation models, which align with physical processes, offer intrinsic explainability but often perform poorly. Conversely, machine learning models for crop yield modeling are powerful and scalable, yet they commonly operate as black boxes and lack adherence to the physical principles of crop growth. This study bridges this gap by coupling the advantages of both worlds. We postulate that the crop yield is inherently defined by the water availability. Therefore, we formulate crop yield as a function of temporal water scarcity and predict both the crop drought stress and the sensitivity to water scarcity at fine-scale resolution. Sequentially modeling the crop yield response to water enables accurate yield prediction. To enforce physical consistency, a novel physics-informed loss function is proposed. We leverage multispectral satellite imagery, meteorological data, and fine-scale yield data. Further, to account for the uncertainty within the model, we build upon a deep ensemble approach. Our method surpasses state-of-the-art models like LSTM and Transformers in crop yield prediction with a coefficient of determination (R2-score) of up to 0.82 while offering high explainability. This method offers decision support for industry, policymakers, and farmers in building a more resilient agriculture in times of changing climate conditions. The code is publicly available at https://github.com/mmiranda-l/Yield-Loss.
Miro Miranda, Marcela Charfuelan, Matias Valdenegro-Toro, Andreas Dengel 0001
ECAI4
2025 PupilSense: A Novel Application for Webcam-Based Pupil Diameter Estimation
Vijul Shah, Ko Watanabe 0001, Brian B. Moser, Andreas Dengel 0001
ETRA4
2025 Investigating the Configurability of LLMs for the Generation of Knowledge Work Datasets
Desiree Heim, Christian Jilek, Adrian Ulges, Andreas Dengel 0001
ICAART (3)4
2025 SAT: Segment and Track Anything for Microscopy
abstract
Integrating cell segmentation with tracking is critical for achieving a detailed and dynamic understanding of cellular behavior. This integration facilitates the study and quantification of cell morphology, movement, and interactions, offering valuable insights into a wide range of biological processes and diseases. However, traditional methods rely on labor-intensive and costly annotations, such as full segmentation masks or bounding boxes for each cell. To address this limitation, we present SAT: Segment and Track Anything for Microscopy, a novel pipeline that leverages point annotations in the first frame to automate cell segmentation and tracking across all subsequent frames. By significantly reducing annotation time and effort, SAT enables efficient and scalable analysis, making it well-suited for large-scale studies. The pipeline was evaluated on two diverse datasets, achieving over 80% Multiple Object Tracking Accuracy (MOTA), demonstrating its robustness and effectiveness across various imaging modalities and cell types. These results highlight SAT’s potential to streamline biomedical research and enable deeper exploration of cellular behavior.
Nabeel Khalid, Mohammadmahdi Koochali, Khola Naseem, Maria Caroprese, Gillian Lovell, Daniel A. Porto, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed
ICAART (2)8
2025 Synthesizing Annotated Cell Microscopy Images with Generative Adversarial Networks
Duway Nicolas Lesmes-Leon, Miro Miranda, Maria Caroprese, Gillian Lovell, Andreas Dengel 0001, Sheraz Ahmed
ICAART (3)5
2025 YeastFormer: An End-to-End Instance Segmentation Approach for Yeast Cells in Microstructure Environment
Khola Naseem, Nabeel Khalid, Lea Bertgen, Johannes M. Herrmann, Andreas Dengel 0001, Sheraz Ahmed
ICAART (2)5
2025 How to Box Your Cells: An Introduction to Box Supervision for 2.5D Cell Instance Segmentation and a Study of Applications
Fabian Schmeisser, Maria Caroprese, Gillian Lovell, Andreas Dengel 0001, Sheraz Ahmed
ICAART (3)4
2025 Knowledge Graph Enrichments for Credit Account Prediction
Michael Schulze, Andreas Dengel 0001
ICAART (2)2
2025 Webcam-Based Pupil Diameter Prediction Benefits from Upscaling
Vijul Shah, Brian B. Moser, Ko Watanabe 0001, Andreas Dengel 0001
ICAART (2)4
2025 Time Series Generation for Augmenting Multi-channel Automotive Audio Data
Philipp Engler, Ludger van Elst, Peter Schichtel, Andreas Dengel 0001, Sheraz Ahmed
ICANN (3)4
2025 DocForgeNet: Dual Cross-Stream Fusion Network for Robust Forgery Detection in Scanned Documents
Nauman Riaz, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (4)3
2025 DP-DocLDM: Differentially Private Document Image Generation Using Latent Diffusion Models
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (4)3
2025 When 512×512 is Not Enough: Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
abstract
Large-scale, pre-trained Text-to-Image (T2I) diffusion models have gained significant popularity in image synthesis and have shown unexpected potential in image Super-Resolution (SR). However, they are usually trained with a resolution limit of 512×512, making scaling beyond this resolution an unresolved but necessary challenge. To address this limitation, we propose a novel approach that enables them to generate 2K, 4K, and even 8K images without any additional training. Our method leverages MultiDiffusion, which distributes the generation across multiple diffusion paths, and local degradation-aware prompt extraction, which guides the T2I model according to its low-resolution input. As a result, we unlock higher resolutions, allowing T2I diffusion to be applied to image SR tasks without limitation on resolution.
Brian B. Moser, Stanislav Frolov, Tobias Christian Nauen, Federico Raue, Andreas Dengel 0001
ICIP5
2025 An Analysis of Temporal Dropout in Earth Observation Time Series for Regression Tasks
Miro Miranda, Francisco Alejandro Mena, Andreas Dengel 0001
IDA3
2025 Distill the Best, Ignore the Rest: A Study in Latent Dataset Distillation on Core-Sets
abstract
Latent dataset distillation, which exploits pre-trained generative priors, has gained significant interest in recent years because it can be applied agnostic to any distillation algorithm and addresses two significant limitations of classical distillation algorithms: cross-architecture generalization and high-resolution synthesis. However, existing approaches typically distill from the entire dataset, potentially including non-beneficial samples. We introduce a novel "Prune First, Distill After" framework that systematically prunes datasets via loss-based sampling prior to latent distillation. By leveraging pruning before classical distillation techniques and generative priors, we create a representative coreset that leads to enhanced generalization for unseen architectures - a significant challenge of current distillation methods. More specifically, our proposed framework significantly boosts distilled quality, achieving up to a 5.2 percentage points accuracy increase even with substantial dataset pruning, i.e., removing 80% of the original dataset prior to distillation. Overall, our experimental results highlight the advantages of our easy-sample prioritization and cross-architecture robustness, paving the way for more effective and high-quality dataset distillation.
Brian B. Moser, Federico Raue, Tobias Christian Nauen, Stanislav Frolov, Andreas Dengel 0001
IJCNN5
2025 Unlocking Dataset Distillation with Diffusion Models
abstract
Dataset distillation seeks to condense datasets into smaller but highly representative synthetic samples. While diffusion models now lead all generative benchmarks, current distillation methods avoid them and rely instead on GANs or autoencoders, or, at best, sampling from a fixed diffusion prior. This trend arises because naive backpropagation through the long denoising chain leads to vanishing gradients, which prevents effective synthetic sample optimization. To address this limitation, we introduce Latent Dataset Distillation with Diffusion Models (LD3M), the first method to learn gradient-based distilled latents and class embeddings end-to-end through a pre-trained latent diffusion model. A linearly decaying skip connection, injected from the initial noisy state into every reverse step, preserves the gradient signal across dozens of timesteps without requiring diffusion weight fine-tuning. Across multiple ImageNet subsets at $128\times128$ and $256\times256$, LD3M improves downstream accuracy by up to 4.8 percentage points (1 IPC) and 4.2 points (10 IPC) over the prior state-of-the-art. The code for LD3M is provided at https://github.com/Brian-Moser/prune_and_distill.
Brian B. Moser, Federico Raue, Sebastian Palacio, Stanislav Frolov, Andreas Dengel 0001
NeurIPS5
2025 SpotDiffusion: A Fast Approach for Seamless Panorama Generation Over Time
abstract
Generating high-resolution images with generative models has recently been made widely accessible by leveraging diffusion models pre-trained on large-scale datasets. Various techniques, such as MultiDiffusion and SyncDiffusion, have further pushed image generation beyond training resolutions, i.e., from square images to panorama, by merging multiple overlapping diffusion paths or employing gradient descent to maintain perceptual coherence. However, these methods suffer from significant computational inefficiencies due to generating and averaging numerous predictions, which is required in practice to produce high-quality and seamless images. This work addresses this limitation and presents a novel approach that eliminates the need to generate and average numerous overlapping denoising predictions. Our method shifts non-overlapping denoising windows over time, ensuring that seams in one timestep are corrected in the next. This results in coherent, high-resolution images with fewer over-all steps. We demonstrate the effectiveness of our approach through qualitative and quantitative evaluations, comparing it with MultiDijfusion, SyncDiffusion, and StitchDijfusion. Our method offers several key benefits, including improved computational efficiency and faster inference times while producing comparable or better image quality.
Stanislav Frolov, Brian B. Moser, Andreas Dengel 0001
WACV3
2025 Dynamic Attention-Guided Diffusion for Image Super-Resolution
abstract
Diffusion models in image Super-Resolution (SR) treat all image regions uniformly, which risks compromising the overall image quality by potentially introducing artifacts during denoising of less-complex regions. To address this, we propose “You Only Diffuse Areas” (YODA), a dynamic attention-guided diffusion process for image SR. YODA selectively focuses on spatial regions defined by attention maps derived from the low-resolution images and the current de-noising time step. This time-dependent targeting enables a more efficient conversion to high-resolution outputs by focusing on areas that benefit the most from the iterative refinement process, i.e., detail-rich objects. We empirically validate YODA by extending leading diffusion-based methods SR3, DiffBIR, and SRDiff. Our experiments demonstrate new state-of-the-art performances in face and general SR tasks across PSNR, SSIM, and LPIPS metrics. As a side effect, we find that YODA reduces color shift issues and stabilizes training with small batches.
Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001
WACV5
2025 Which Transformer to Favor: A Comparative Analysis of Efficiency in Vision Transformers
abstract
Self-attention in Transformers comes with a high computational cost because of their quadratic computational complexity, but their effectiveness in addressing problems in language and vision has sparked extensive research aimed at enhancing their efficiency. However, diverse experimental conditions, spanning multiple input domains, prevent a fair comparison based solely on reported results, posing challenges for model selection. To address this gap in comparability, we perform a large-scale benchmark of more than 45 models for image classification, evaluating key efficiency aspects, including accuracy, speed, and memory usage. Our benchmark provides a standardized baseline for efficiency-oriented transformers. We analyze the results based on the Pareto front - the boundary of optimal models. Surprisingly, despite claims of other models being more efficient, ViT remains Pareto optimal across multiple metrics. We observe that hybrid attention-CNN models exhibit remarkable inference memory- and parameter-efficiency. Moreover, our benchmark shows that using a larger model in general is more efficient than using higher resolution images. Thanks to our holistic evaluation, we provide a centralized resource for practitioners and researchers, facilitating informed decisions when selecting or developing efficient transformers.11https://github.com/tobna/WhatTransformerToFavor
Tobias Christian Nauen, Sebastian Palacio, Federico Raue, Andreas Dengel 0001
WACV4
2025 VAEneu: a new avenue for VAE application on probabilistic forecasting
abstract
This paper introduces VAEneu, a novel autoregressive method for multistep ahead univariate probabilistic time series forecasting, designed to address the challenges of generating sharp and well-calibrated probabilistic forecasts without assuming a specific parametric form for the predictive distribution. VAEneu leverages the Conditional VAE framework and optimizes the likelihood of the predictive distribution using the Continuous Ranked Probability Score (CRPS), a strictly proper scoring rule, as the loss function. This approach enables the model to learn flexible, sharp, and well-calibrated predictive distributions without the need for a tractable likelihood function. In a comprehensive empirical study, VAEneu is rigorously benchmarked against 12 baseline models across 12 datasets, demonstrating superior performance in both forecasting accuracy and uncertainty quantification. VAEneu provides a valuable tool for quantifying future uncertainties, and our extensive empirical study lays the foundation for future comparative studies for univariate multistep ahead probabilistic forecasting.
Alireza Koochali, Ensiye Tahaei, Andreas Dengel 0001, Sheraz Ahmed
Appl. Intell.3
2025 KRNN: A hybrid data and knowledge oriented time series forecasting approach for health care applications
Muhammad Ali Chattha, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
Expert Syst. Appl.3
2025 PassionNet: An innovative framework for duplicate and conflicting requirements identification
abstract
Early detection of duplicate and conflicting requirements in software development lifecycle is crucial to achieve software project efficiency, quality, and market success. Primarily, duplicate detection requires identifying semantic equivalence, intent alignment, functional overlap, and domain-specific terminology variations between differently worded requirements. Whereas, conflict detection demands recognising logical contradictions, constraint violations, resource conflicts and temporal incompatibilities between requirements. To handle multi-dimensional demands of two different task types, researchers have developed 32 AI based duplicate and conflicting requirement detection predictors. However, despite the utility of sophisticated large language models (LLMs) and sampling techniques, existing approaches significantly lack in performance because they fail to comprehensively handle multi-dimensional demands of both tasks. To address these gaps, this paper presents a modular framework “PassionNet” which implements a novel strategy of integrating 10 different multi-dimensional similarity assessments with the contextual understanding of 8 unique language model variants. The framework enables three distinct pipeline types: language model-based pipelines that capture semantic intent, similarity knowledge-driven pipelines that detect lexical, structural and distributional patterns, and hybrid pipelines that combine both approaches to simultaneously assess all dimensions of requirement relationships. Our experimental evaluation of 760 pipelines across six public datasets demonstrates that hybrid pipelines outperform the other two approaches in terms of F1-score as compared to state-of-the-art methods. Specifically, the hybrid pipeline achieves an improvement in F1-score of approximately 4% on the WorldVista dataset, 5% on the UAV dataset and 3% on the Pure dataset as compared to the state-of-the-art models. Statistical validation through t-tests confirms the significance of these improvements (p < 0.1 with 10 permutations, approaching zero with 1000 permutations). The results provide empirical evidence that effective requirement analysis requires simultaneously assessing semantic, lexical, structural, and logical dimensions of requirements rather than focusing on isolated aspects. To facilitate software engineers, researchers and practitioners, PassionNet web application is deployed at https://sds_requirement_engineering.opendfki.de/
Summra Saleem, Muhammad Nabeel Asim, Andreas Dengel 0001
Expert Syst. Appl.3
2025 Missing data as augmentation in the Earth Observation domain: A multi-view learning approach
abstract
Multi-view learning (MVL) leverages multiple sources or views of data to enhance machine learning model performance and robustness. This approach has been successfully used in the Earth Observation (EO) domain, where views have a heterogeneous nature and can be affected by missing data. Despite the negative effect that missing data has on model predictions, the ML literature has used it as an augmentation technique to improve model generalization, like masking the input data. Inspired by this, we introduce novel methods for EO applications tailored to MVL with missing views. Our methods integrate the combination of a set to simulate all combinations of missing views as different training samples. Instead of replacing missing data with a numerical value, we use dynamic merge functions, like average, and more complex ones like Transformer. This allows the MVL model to entirely ignore the missing views, enhancing its predictive robustness. We experiment on four EO datasets with temporal and static views, including state-of-the-art methods from the EO domain. The results indicate that our methods improve model robustness under conditions of moderate missingness, and improve the predictive performance when all views are present. The proposed methods offer a single adaptive solution to operate effectively with any combination of available views.
Francisco Alejandro Mena, Diego Arenas, Andreas Dengel 0001
Neurocomputing3
2025 RPCP-PURI: A robust and precise computational predictor for Phishing Uniform Resource Identification
Tayyaba Asif, Faiza Mehmood, Syed Ahmed Mazhar Gillani, Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Waqar Mahmood, Andreas Dengel 0001
J. Inf. Secur. Appl.7
2025 Addressing data dependency in neural networks: introducing the Knowledge Enhanced Neural Network (KENN) for time series forecasting +
Muhammad Ali Chattha, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
Mach. Learn.3
2025 Multi-modal co-learning for Earth observation: enhancing single-modality models via modality collaboration
Francisco Alejandro Mena, Dino Ienco, Cássio Fraga Dantas, Roberto Interdonato, Andreas Dengel 0001
Mach. Learn.5
2025 Diffusion Models, Image Super-Resolution, and Everything: A Survey
abstract
Diffusion models (DMs) have disrupted the image super-resolution (SR) field and further closed the gap between image quality and human perceptual preferences. They are easy to train and can produce very high-quality samples that exceed the realism of those produced by previous generative methods. Despite their promising results, they also come with new challenges that need further research: high computational demands, comparability, lack of explainability, color shifts, and more. Unfortunately, entry into this field is overwhelming because of the abundance of publications. To address this, we provide a unified recount of the theoretical foundations underlying DMs applied to image SR and offer a detailed analysis that underscores the unique characteristics and methodologies within this domain, distinct from broader existing reviews in the field. This article articulates a cohesive understanding of DM principles and explores current research avenues, including alternative input domains, conditioning techniques, guidance mechanisms, corruption spaces, and zero-shot learning approaches. By offering a detailed examination of the evolution and current trends in image SR through the lens of DMs, this article sheds light on the existing challenges and charts potential future directions, aiming to inspire further innovation in this rapidly advancing area.
Brian B. Moser, Arundhati S. Shanbhag, Federico Raue, Stanislav Frolov, Sebastian Palacio, Andreas Dengel 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Feature Estimation of Global Language Processing in EEG Using Attention Maps
Dai Shimizu, Ko Watanabe 0001, Andreas Dengel 0001
ACCV (2)3
2024 Towards Cyber Mapping the German Financial System with Knowledge Graphs
Markus Schröder 0001, Jacqueline Krüger, Neda Foroutan, Philipp Horn, Christoph Fricke, Ezgi Delikanli, Heiko Maus, Andreas Dengel 0001
ESWC (1)8
2024 Eye Movement in a Controlled Dialogue Setting
abstract
Designing realistic eye movements for animated avatars poses a challenge, as gaze behavior is predominantly unconscious. Accurately modulating those movements is crucial to avoid the Uncanny Valley. The human gaze exhibits different characteristics in conversations, depending on speaking or listening. Albeit these distinctions are known, data for synthesizing eye movement models suitable for avatars is scarce. This research introduces a novel dataset involving human gaze behavior during remote screen conversations. The data are collected from 19 participants, offering 4 hours of gaze data labeled as Speaking and Listening. Our data analysis substantiates prior knowledge of gaze behavior while providing new insights through higher precision. Furthermore, we demonstrate the dataset’s suitability for machine learning algorithms by training a classifier, achieving 88.1% binary classification accuracy.
David Dembinsky, Ko Watanabe 0001, Andreas Dengel 0001, Shoya Ishimaru
ETRA3
2024 Medi-CAT: Contrastive Adversarial Training for Medical Image Classification
Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed
ICAART (3)2
2024 A Unique Training Strategy to Enhance Language Models Capabilities for Health Mention Detection from Social Media Content
Pervaiz Iqbal Khan, Muhammad Nabeel Asim, Andreas Dengel 0001, Sheraz Ahmed
ICAART (3)3
2024 Knowledge-Aware Object Detection in Traffic Scenes
Jean-Francois Jacques Nicolas Nies, Syed Tahseen Raza Rizvi, Mohsin Munir, Ludger van Elst, Andreas Dengel 0001
ICAART (3)5
2024 CellSpot: Deep Learning-Based Efficient Cell Center Detection in Microscopic Images
Nabeel Khalid, Maria Caroprese, Gillian Lovell, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed
ICANN (8)5
2024 A Study in Dataset Pruning for Image Super-Resolution
Brian B. Moser, Federico Raue, Andreas Dengel 0001
ICANN (2)3
2024 Point-Based Weakly Supervised 2.5D Cell Segmentation
Fabian Schmeisser, Andreas Dengel 0001, Sheraz Ahmed
ICANN (8)2
2024 Latent Diffusion for Guided Document Table Generation
Syed Jawwad Haider Hamdani, Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (5)4
2024 StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation
Nauman Riaz, Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (2)4
2024 DocXplain: A Novel Model-Agnostic Explainability Method for Document Image Classification
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (4)3
2024 CCATS: Moving Forward with Class-Conditional Time Series Generation
Philipp Engler, Alireza Koochali, Ludger van Elst, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (3)4
2024 Improving Text Representation for Disease Detection from Social Media via Self-augmentation and Contrastive Learning
Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (5)2
2024 TaylorShift: Shifting the Complexity of Self-attention from Squared to Linear (and Back) Using Taylor-Softmax
Tobias Christian Nauen, Sebastian Palacio, Andreas Dengel 0001
ICPR (6)3
2024 Generating Counterfactual Trajectories with Latent Diffusion Models for Concept Discovery
Payal Varshney, Adriano Lucieri, Christoph Peter Balada, Andreas Dengel 0001, Sheraz Ahmed
ICPR (12)4
2024 Impact Assessment of Missing Data in Model Predictions for Earth Observation Applications
abstract
Earth observation (EO) applications involving complex and heterogeneous data sources are commonly approached with machine learning models. However, there is a common assumption that data sources will be persistently available. Different situations could affect the availability of EO sources, like noise, clouds, or satellite mission failures. In this work, we assess the impact of missing temporal and static EO sources in trained models across four datasets with classification and regression tasks. We compare the predictive quality of different methods and find that some are naturally more robust to missing data. The Ensemble strategy, in particular, achieves a prediction robustness up to 100%. We evidence that missing scenarios are significantly more challenging in regression than classification tasks. Finally, we find that the optical view is the most critical view when it is missing individually.
Francisco Alejandro Mena, Diego Arenas, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS5
2024 Multi-Modal Fusion Methods with Local Neighborhood Information for Crop Yield Prediction at Field and Subfield Levels
abstract
Yield prediction at both field and subfield level poses a significant challenge, yet it holds paramount importance for decision-making and food security within the agricultural sector. Recent efforts, focused on integrating remote sensing data coupled with machine learning models, thereby creating globally scalable models for various crop types. This study underscores the effectiveness of Sentinel-2 and complementary data sources such as weather, soil, and terrain in enhancing machine learning-based yield prediction. We address the limitations of previous works and introduce a framework that incorporates local neighborhood information using convolutional neural networks and geographical coordinates. Additionally, we address the complexity of sensor fusion, showcasing both input fusion and feature fusion frameworks. We highlight that handling modalities with varying spatial and temporal resolutions requires adequate and advanced fusion mechanisms in crop yield prediction. Notably, this study reports an R2of 0.86 for soybean in Argentina using a feature fusion scheme with attention mechanism. The results are demonstrated on a large yield dataset for soybean, wheat, and rapeseed distributed across Argentina, Uruguay, and Germany.
Miro Miranda, Deepak Pathak, Marlon Nuske, Andreas Dengel 0001
IGARSS4
2024 Xai-Guided Enhancement of Vegetation Indices for Crop Mapping
abstract
Vegetation indices allow to efficiently monitor vegetation growth and agricultural activities. Previous generations of satellites were capturing a limited number of spectral bands, and a few expert-designed vegetation indices were sufficient to harness their potential. New generations of multi- and hyperspectral satellites can however capture additional bands, but are not yet efficiently exploited. In this work, we propose an explainable-AI-based method to select and design suitable vegetation indices. We first train a deep neural network using multispectral satellite data, then extract feature importance to identify the most influential bands. We subsequently select suitable existing vegetation indices or modify them to incorporate the identified bands and retrain our model. We validate our approach on a crop classification task. Our results indicate that models trained on individual indices achieve comparable results to the baseline model trained on all bands, while the combination of two indices surpasses the baseline in certain cases.
Hiba Najjar, Francisco Alejandro Mena, Marlon Nuske, Andreas Dengel 0001
IGARSS4
2024 Assessment of Sentinel-2 Spatial and Temporal Coverage Based on the Scene Classification Layer
abstract
Since the launch of the Sentinel-2 (S2) satellites, many ML models have used the data for diverse applications. The scene classification layer (SCL) inside the S2 product provides rich information for training, such as filtering images with high cloud coverage. However, there is more potential in this. We propose a technique to assess the clean optical coverage of a region, expressed by a SITS and calculated with the S2-based SCL data. With a manual threshold and specific labels in the SCL, the proposed technique assigns a percentage of spatial and temporal coverage across the time series and a high/low assessment. By evaluating the AI4EO challenge for Enhanced Agriculture, we show that the assessment is correlated to the predictive results of ML models. The classification results in a region with low spatial and temporal coverage is worse than in a region with high coverage. Finally, we applied the technique across all continents of the global dataset LandCoverNet.
Cristhian Sanchez, Francisco Alejandro Mena, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS5
2024 Waving Goodbye to Low-Res: A Diffusion-Wavelet Approach for Image Super-Resolution
abstract
Image Super-Resolution (SR) remains challenging, particularly in achieving high-quality details without extensive computational cost. Existing methods often struggle to balance the trade-off between image quality, especially in high-frequency details, and computational efficiency. In this paper, we present a novel Diffusion-Wavelet (DiWa) approach for bridging this gap. It leverages the strengths of diffusion models and discrete wavelet transformation. By enabling the diffusion model to operate in the frequency domain, our models effectively hallucinate highfrequency information for SR images on the wavelet spectrum, resulting in high-quality and detailed reconstructions in image space. Quantitatively, our method outperforms other state-ofthe-art diffusion-based SR methods, namely SR3 and SRDiff, regarding PSNR, SSIM, and LPIPS on both face (8x scaling) and general (4x scaling) SR benchmarks. Meanwhile, using the frequency domain allows us to use fewer parameters than the compared models: 92M parameters instead of 550M compared to SR3 and 9.3M instead of 12M compared to SRDiff. Additionally, DiWa outperforms other state-of-the-art generative methods on general SR datasets while saving inference time (ca. 250 %).
Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001
IJCNN5
2024 ObjBlur: A Curriculum Learning Approach With Progressive Object-Level Blurring for Improved Layout-to-Image Generation
Stanislav Frolov, Brian B. Moser, Sebastian Palacio, Andreas Dengel 0001
ACM Multimedia4
2024 Context-based Entity Recommendation for Knowledge Workers: Establishing a Benchmark on Real-life Data
abstract
In recent decades, Recommender Systems (RS) have undergone significant advancements, particularly in popular domains like movies, music, and product recommendations. Yet, progress has been notably slower in leveraging these systems for personal information management and knowledge assistance. In addition to challenges that complicate the adoption of RS in this domain (such as privacy concerns, heterogeneous recommendation items, and frequent context switching), a significant barrier to progress in this area has been the absence of a standardized benchmark for researchers to evaluate their approaches. In response to this gap, this paper presents a benchmark built upon a publicly available dataset of Real-Life Knowledge Work in Context (RLKWiC). This benchmark focuses on evaluating context-based entity recommendation, a use case for leveraging RS to support knowledge workers in their daily digital tasks. By providing this benchmark, it is aimed to facilitate and accelerate research efforts in enhancing personal knowledge assistance through RS.
Mahta Bakhshizadeh, Heiko Maus, Andreas Dengel 0001
RecSys3
2024 SphereCraft: A Dataset for Spherical Keypoint Detection, Matching and Camera Pose Estimation
abstract
This paper introduces SphereCraft, a dataset specifically designed for spherical keypoint detection, matching, and camera pose estimation. The dataset addresses the limitations of existing datasets by providing extracted keypoints from various detectors, along with their ground truth correspondences. Synthetic scenes with photo-realistic rendering and accurate 3D meshes are included, as well as real-world scenes acquired from different spherical cameras. SphereCraft enables the development and evaluation of algorithms targeting multiple camera viewpoints, advancing the state-of-the-art in computer vision tasks involving spherical images. Our dataset is available at https://dfki.github.io/spherecraftweb/.
Christiano Couto Gava, Yunmin Cho, Federico Raue, Sebastian Palacio, Alain Pagani, Andreas Dengel 0001
WACV6
2024 From private to public: benchmarking GANs in the context of private time series classification
Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed
Appl. Intell.2
2024 DocXclassifier: towards a robust and interpretable deep neural network for document image classification
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
Int. J. Document Anal. Recognit.3
2024 Towards privacy preserved document image classification: a comprehensive benchmark
Saifullah Saifullah, Dominique Mercier, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
Int. J. Document Anal. Recognit.4
2024 Explainable Multimodal Learning in Remote Sensing: Challenges and Future Directions
abstract
Earth observation applications effectively leverage deep learning models to harness the abundantly available remote sensing data. In order to use all the different modalities relevant to a specific task, the fusion of these data sources can be achieved using multi-modal learning techniques. This is especially helpful when the input dataset contains both images and tabular data, or when the temporal and spatial resolutions vary across the modalities of interest. Nevertheless, these fusion techniques typically increase in complexity as the disparities in the nature of the fused modalities increase. The resulting complex deep learning models suffer from a lack of explainability and transparency, which is crucial in many sensitive human-related applications. In this letter, we describe how the research community in remote sensing addresses the issue of model explainability in the context of multi-modal learning. We additionally review the practices used in other application fields and identify some of the most promising explainability methods tailored for multi-modal deep networks to be exploited in remote sensing applications.
Alexander Günther, Hiba Najjar, Andreas Dengel 0001
IEEE Geosci. Remote. Sens. Lett.3
2024 Passion-Net: a robust precise and explainable predictor for hate speech detection in Roman Urdu text
abstract
Abstract With an aim to eliminate or reduce the spread of hate content across social media platforms, the development of artificial intelligence supported computational predictors is an active area of research. However, diversity of languages hinders development of generic predictors that can precisely identify hate content. Several language-specific hate speech detection predictors have been developed for most common languages including English, Chinese and German. Specifically, for Urdu language a few predictors have been developed and these predictors lack in predictive performance. The paper in hand presents a precise and explainable deep learning predictor which makes use of advanced language modelling strategies for the extraction of semantic and discriminative patterns. Extracted patterns are utilized to train an attention-based novel classifier that is competent in precisely identifying hate content. Over coarse-grained benchmark dataset, the proposed predictor significantly outperforms state-of-the-art predictor by 8.7% in terms of accuracy, precision and F1-score. Similarly, over fine-grained dataset, in comparison with state-of-the-art predictor, it achieves performance gain of 10.6%, 17.6%, 18.6% and 17.6% in terms of accuracy, precision, recall and F1-score.
Faiza Mehmood, Hina Ghafoor, Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Waqar Mahmood, Andreas Dengel 0001
Neural Comput. Appl.6
2023 Localized Semantic Feature Mixers for Efficient Pedestrian Detection in Autonomous Driving
abstract
Autonomous driving systems rely heavily on the underlying perception module which needs to be both performant and efficient to allow precise decisions in realtime. Avoiding collisions with pedestrians is of topmost priority in any autonomous driving system. Therefore, pedestrian detection is one of the core parts of such systems' perception modules. Current state-of-the-art pedestrian detectors have two major issues. Firstly, they have long inference times which affect the efficiency of the whole perception module, and secondly, their performance in the case of small and heavily occluded pedestrians is poor: We propose Local-ized Semantic Feature Mixers (LSFM), a novel, anchor-free pedestrian detection architecture. It uses our novel Super Pixel Pyramid Pooling module instead of the, computation-ally costly, Feature Pyramid Networks for feature encoding. Moreover, our MLPMixer-based Dense Focal Detection Network is used as a light detection head, reducing computational effort and inference time compared to existing approaches. To boost the performance of the proposed architecture, we adapt and use mixup augmentation which improves the performance, especially in small and heavily occluded cases. We benchmark LSFM against the state-of-the-art on well-established traffic scene pedestrian datasets. The proposed LSFM achieves state-of-the-art performance in Caltech, City Persons, Euro City Persons, and TJU-Traffic-Pedestrian datasets while reducing the inference time on average by 55%. Further, LSFM beats the human baseline for the first time in the history of pedestrian detection. Finally, we conducted a cross-dataset evaluation which proved that our proposed LSFM generalizes well to unseen data.
Abdul Hannan Khan, Mohammed Shariq Nawaz, Andreas Dengel 0001
CVPR3
2023 Randout-KD: Finetuning Foundation Models for Text Classification via Random Noise and Knowledge Distillation
Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed
ICAART (3)2
2023 Knowledge Forcing: Fusing Knowledge-Driven Approaches with LSTM for Time Series Forecasting
Muhammad Ali Chattha, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
ICANN (6)3
2023 PACE: Point Annotation-Based Cell Segmentation for Efficient Microscopic Image Analysis
Nabeel Khalid, Tiago Comassetto Fróes, Maria Caroprese, Gillian Lovell, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed
ICANN (2)6
2023 DWA: Differential Wavelet Amplifier for Image Super-Resolution
Brian B. Moser, Stanislav Frolov, Federico Raue, Sebastian Palacio, Andreas Dengel 0001
ICANN (2)5
2023 ColDBin: Cold Diffusion for Document Image Binarization
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (5)3
2023 Instance Segmentation Based Graph Extraction for Handwritten Circuit Diagram Images
Johannes Bayer, Amit Kumar Roy, Andreas Dengel 0001
ICPRAM3
2023 Sequential Spatial Transformer Networks for Salient Object Classification
David Dembinsky, Fatemeh Azimi, Federico Raue, Jörn Hees, Sebastian Palacio, Andreas Dengel 0001
ICPRAM6
2023 Crop Yield Prediction: An Operational Approach to Crop Yield Modeling on Field and Subfield Level with Machine Learning Models
abstract
Accurate and reliable crop yield prediction is a complex task. The yield of a crop depends on a variety of factors whose accurate measurement and modeling is challenging. At the same time, reliable yield prediction is highly desirable for farmers to optimize crop production. In this paper, we introduce a modeling based on remote sensing data and Machine Learning models evaluated on a large-scale dataset to address the challenge of an operational crop yield estimation and forecasting on field and subfield level. With our approach, we aim towards a global yield modeling based on Machine Learning models which operates across crop types without the need for crop-specific modeling. We demonstrate that our approach learns to map in-field variability for all studied crop types. Overall, the predictions have an error (RRMSE) of around 15% and an R2value of 0.77 at field level.
Patrick Helber, Benjamin Bischke, Peter Habelitz, Cristhian Sanchez, Deepak Pathak, Miro Miranda, Hiba Najjar, Francisco Alejandro Mena, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS14
2023 A Comparative Assessment of Multi-View Fusion Learning For Crop Classification
abstract
With a rapidly increasing amount and diversity of remote sensing (RS) data sources, there is a strong need for multi-view learning modeling. This is a complex task when considering the differences in resolution, magnitude, and noise of RS data. The typical approach for merging multiple RS sources has been input-level fusion, but other - more advanced - fusion strategies may outperform this traditional approach. This work assesses different fusion strategies for crop classification in the CropHarvest dataset. The fusion methods proposed in this work outperform models based on individual views and previous fusion methods. We do not find one single fusion method that consistently outperforms all other approaches. Instead, we present a comparison of multi-view fusion methods for three different datasets and show that, depending on the test region, different methods obtain the best performance. Despite this, we suggest a preliminary criterion for the selection of fusion methods.
Francisco Alejandro Mena, Diego Arenas, Marlon Nuske, Andreas Dengel 0001
IGARSS4
2023 Effect Of Terrain Information On Multimodal Deep Learning For Flood Disaster Detection
abstract
The utilization of multimodal analysis techniques, combining satellite imagery with terrain information, has gained prominence in flood detection. This study focuses on the inclusion of elevation data into the Sen1floods11 dataset, an open dataset for flood damage detection, to investigate the influence of terrain information on flood detection tasks. Among the considered terrain information, the inclusion of elevation data resulted in a bias of overestimating the presence of water in relatively low-lying areas, without contributing to accuracy improvement. However, the utilization of slope information, derived from differentiating the elevation data, mitigated such bias and yielded a slight improvement in accuracy. This finding aligns with the utilization of derivatives in physical equations describing flood flow, suggesting the explicit incorporation of physics-based principles, such as the flow of water based on slope, to enhance model accuracy in future research endeavors.
Takashi Miyamoto, Marco Stricker, Jun Ogishima, Kevin Iselborn, Marlon Nuske, Andreas Dengel 0001
IGARSS6
2023 Feature Attribution Methods for Multivariate Time-Series Explainability in Remote Sensing
abstract
Numerous remote sensing applications rely on temporal satellite data, and Deep learning models are increasingly being used for such tasks. Nevertheless, these models operate as black boxes, lacking transparency and understandability. We address this gap by using explainable AI on an agricultural task. Specifically, we trained a recurrent neural network on individual pixels from multispectral time-series of Sentinel-2 satellite images to predict crop yield. We then applied nine feature attribution methods on a sample of the dataset and computed the spectral and temporal contributions to the final individual predictions. The aggregated results were evaluated qualitatively and quantitatively. Results suggest that LIME and Shapley sampling value methods performed best on the quantitative scores, followed by GradientShap. Most backpropagation-based techniques had highly inconsistent scores across the explained data points. Finally, to guide remote sensing practitioners in using Explainable AI on similar datasets, we further discuss some selection criteria to be considered.
Hiba Najjar, Patrick Helber, Benjamin Bischke, Peter Habelitz, Cristhian Sanchez, Francisco Alejandro Mena, Miro Miranda, Deepak Pathak, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS14
2023 Predicting Crop Yield with Machine Learning: An Extensive Analysis of Input Modalities and Models on a Field and Sub-Field Level
abstract
We introduce a simple yet effective early fusion method for crop yield prediction that handles multiple input modalities with different temporal and spatial resolutions. We use high-resolution crop yield maps as ground truth data to train crop and machine learning model agnostic methods at the sub-field level. We use Sentinel-2 satellite imagery as the primary modality for input data with other complementary modalities, including weather, soil, and DEM data. The proposed method uses input modalities available with global coverage, making the framework globally scalable. We explicitly highlight the importance of input modalities for crop yield prediction and emphasize that the best-performing combination of input modalities depends on region, crop, and chosen model.
Deepak Pathak, Miro Miranda, Francisco Alejandro Mena, Cristhian Sanchez, Patrick Helber, Benjamin Bischke, Peter Habelitz, Hiba Najjar, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Marcela Charfuelan, Marlon Nuske, Andreas Dengel 0001
IGARSS14
2023 Influence of Data Cleaning Techniques on Sub-Field Yield Predictions
abstract
Modern combine harvesters can collect geo-located real-time yield measurement while harvesting. This data can be used to train Machine Learning models that predict the yield at sub-field level based on remote sensing input data. The performance of these models is, however, highly dependent on the quality of the yield data. It is therefore important to develop automatic cleaning techniques to correct for common errors in combine harvester yield maps. In this work, we compare different combinations of data cleaning techniques by evaluating their impact on the yield-prediction model performance at field and sub-field level. Our findings indicate that basic cleaning techniques such as absolute thresholds are sufficient at the field level, whereas the performance at the sub-field level is enhanced through the utilization of more intricate statistical cleaning methods.
Cristhian Sanchez, Deepak Pathak, Miro Miranda, Marcela Charfuelan, Patrick Helber, Marlon Nuske, Benjamin Bischke, Peter Habelitz, Nafisur Rahman, Francisco Alejandro Mena, Hiba Najjar, Jayanth Siddamsetty, Diego Arenas, Michaela Vollmer, Andreas Dengel 0001
IGARSS15
2023 Fusing Digital Elevation Maps with Satellite Imagery for Flood Mapping
abstract
Floods are one of the most severe natural catastrophes and therefore emergency response operations are crucial in order to save lifes. These operations require information about flooded areas so that rescue missions can precisely and efficiently use their available resources. This requires a quick automated procedure which is able to identify these regions from remote sensing images. To achieve this goal we utilize machine learning and apply our method on the Sen1Floods11 dataset. Our main contribution lies in the fusion of Digital Elevation Maps (DEMs) with Satellite data. We investigate the effect of several different combinations of processing methods of DEMs, such as depression filling, deriving slope and curvature or flow metrics. In total 44 different experiments have been performed where our best performing combination outperformed the benchmark in terms of mean IoU. Lastly, we also publish our code for downloading and processing DEMS as well as running our experiments.
Marco Stricker, Takashi Miyamoto, Kevin Iselborn, Marlon Nuske, Andreas Dengel 0001
IGARSS5
2023 Cross-Domain Transformation for Outlier Detection on Tabular Datasets
abstract
The overwhelming success of Deep Learning approaches in recent years is often driven by the availability of large public datasets. However, in some domains like finance, creating and sharing realistic datasets is hindered by secrecy or privacy concerns. This can lead to a mismatch, where approaches that have proven to work well on public, research-oriented datasets end up underperforming when applied to real-world (private) datasets. In this work, we focus on the task of Outlier Detection (OD) and bridge the above gap by building an autoencoder based Deep Learning approach that can transform samples between two tabular datasets (e.g., a private and public one). The goal of our approach is that transformed samples become similar to the target dataset, while inliers remain inliers and outliers remain outliers. Among others, after successful transformation, this allows applying of proven methods on public datasets to internal datasets, even if they are of different dimensionality (rows and columns). To evaluate our approach, we introduce metrics to measure dataset similarity and the quality of transformed samples. Our experimental results show that combining public datasets with transformed samples of other datasets leads to higher dataset similarity while sustaining performance w.r.t. common OD algorithms.
Dayananda Herurkar, Timur Sattarov, Jörn Hees, Sebastian Palacio, Federico Raue, Andreas Dengel 0001
IJCNN6
2023 Intelligence Augmentation: Future Directions and Ethical Implications in HCI
Andrew W. Vargo, Benjamin Tag, Mathilde Hutin, Victoria Abou Khalil, Shoya Ishimaru, Olivier Augereau, Tilman Dingler, Motoi Iwata, Koichi Kise, Laurence Devillers, Andreas Dengel 0001
INTERACT (4)11
2023 Deep Learning Architectures for the Prediction of YY1-Mediated Chromatin Loops
Ahtisham Fazeel Abbasi, Muhammad Nabeel Asim, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed
ISBRA4
2023 Quantifying quality of class-conditional generative models in time series domain
abstract
Abstract Despite recent breakthroughs in the domain of implicit generative models, the task of evaluating these models remains a challenging task. With no single metric to assess overall performance, various existing metrics only offer partial information. This issue is further compounded for unintuitive data types such as time series, where manual inspection is infeasible. This deficiency hinders the confident application of modern implicit generative models on time series data. To alleviate this problem, we propose two new metrics, the InceptionTime Score (ITS) and the Fréchet InceptionTime Distance (FITD), to assess the quality of class-conditional generative models on time series data. We conduct extensive experiments on 80 different datasets to study the discriminative capabilities of proposed metrics alongside two existing evaluation metrics: Train on Synthetic Test on Real (TSTR) and Train on Real Test on Synthetic (TRTS). Our evaluations reveal that the proposed assessment evaluation metrics, i.e., ITS and FITD in combination with TSTR, can accurately assess class-conditional generative model performance and detect common issues in implicit generative models. Our findings suggest that the proposed evaluation framework can be a valuable tool for confidently applying modern implicit generative models in time series analysis.
Alireza Koochali, Maria Walch, Sankrutyayan Thota, Peter Schichtel, Andreas Dengel 0001, Sheraz Ahmed
Appl. Intell.5
2023 DNA-MP: a generalized DNA modifications predictor for multiple species based on powerful sequence encoding method
abstract
Accurate prediction of deoxyribonucleic acid (DNA) modifications is essential to explore and discern the process of cell differentiation, gene expression and epigenetic regulation. Several computational approaches have been proposed for particular type-specific DNA modification prediction. Two recent generalized computational predictors are capable of detecting three different types of DNA modifications; however, type-specific and generalized modifications predictors produce limited performance across multiple species mainly due to the use of ineffective sequence encoding methods. The paper in hand presents a generalized computational approach "DNA-MP" that is competent to more precisely predict three different DNA modifications across multiple species. Proposed DNA-MP approach makes use of a powerful encoding method "position specific nucleotides occurrence based 117 on modification and non-modification class densities normalized difference" (POCD-ND) to generate the statistical representations of DNA sequences and a deep forest classifier for modifications prediction. POCD-ND encoder generates statistical representations by extracting position specific distributional information of nucleotides in the DNA sequences. We perform a comprehensive intrinsic and extrinsic evaluation of the proposed encoder and compare its performance with 32 most widely used encoding methods on $17$ benchmark DNA modifications prediction datasets of $12$ different species using $10$ different machine learning classifiers. Overall, with all classifiers, the proposed POCD-ND encoder outperforms existing $32$ different encoders. Furthermore, combinedly over 5-fold cross validation benchmark datasets and independent test sets, proposed DNA-MP predictor outperforms state-of-the-art type-specific and generalized modifications predictors by an average accuracy of 7% across 4mc datasets, 1.35% across 5hmc datasets and 10% for 6ma datasets. To facilitate the scientific community, the DNA-MP web application is available at https://sds_genetic_analysis.opendfki.de/DNA_Modifications/.
Muhammad Nabeel Asim, Muhammad Ali Ibrahim, Ahtisham Fazeel, Andreas Dengel 0001, Sheraz Ahmed
Briefings Bioinform.4
2023 Analyzing the potential of active learning for document image classification
abstract
Abstract Deep learning has been extensively researched in the field of document analysis and has shown excellent performance across a wide range of document-related tasks. As a result, a great deal of emphasis is now being placed on its practical deployment and integration into modern industrial document processing pipelines. It is well known, however, that deep learning models are data-hungry and often require huge volumes of annotated data in order to achieve competitive performances. And since data annotation is a costly and labor-intensive process, it remains one of the major hurdles to their practical deployment. This study investigates the possibility of using active learning to reduce the costs of data annotation in the context of document image classification, which is one of the core components of modern document processing pipelines. The results of this study demonstrate that by utilizing active learning (AL), deep document classification models can achieve competitive performances to the models trained on fully annotated datasets and, in some cases, even surpass them by annotating only 15–40% of the total training dataset. Furthermore, this study demonstrates that modern AL strategies significantly outperform random querying, and in many cases achieve comparable performance to the models trained on fully annotated datasets even in the presence of practical deployment issues such as data imbalance, and annotation noise, and thus, offer tremendous benefits in real-world deployment of deep document classification models. The code to reproduce our experiments is publicly available at https://github.com/saifullah3396/doc_al .
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
Int. J. Document Anal. Recognit.3
2023 Hitchhiker's Guide to Super-Resolution: Introduction and Recent Advances
abstract
With the advent of Deep Learning (DL), Super-Resolution (SR) has also become a thriving research area. However, despite promising results, the field still faces challenges that require further research, e.g., allowing flexible upsampling, more effective loss functions, and better evaluation metrics. We review the domain of SR in light of recent advances and examine state-of-the-art models such as diffusion (DDPM) and transformer-based SR models. We critically discuss contemporary strategies used in SR and identify promising yet unexplored research directions. We complement previous surveys by incorporating the latest developments in the field, such as uncertainty-driven losses, wavelet networks, neural architecture search, novel normalization methods, and the latest evaluation techniques. We also include several visualizations for the models and methods throughout each chapter to facilitate a global understanding of the trends in the field. This review ultimately aims at helping researchers to push the boundaries of DL applied to SR.
Brian B. Moser, Federico Raue, Stanislav Frolov, Sebastian Palacio, Jörn Hees, Andreas Dengel 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 EnML: Multi-label Ensemble Learning for Urdu Text Classification
abstract
Exponential growth of electronic data requires advanced multi-label classification approaches for the development of natural language processing (NLP) applications such as recommendation systems, drug reaction detection, hate speech detection, and opinion recognition/mining. To date, several machine and deep learning–based multi-label classification methodologies have been proposed for English, French, German, Chinese, Arabic, and other developed languages. Urdu is the 11th largest language in the world and has no computer-aided multi-label textual news classification approach. Unlike other languages, Urdu is lacking multi-label text classification datasets that can be used to benchmark the performance of existing machine and deep learning methodologies. With an aim to accelerate and expedite research for the development of Urdu multi-label text classification–based applications, this article provides multiple contributions as follows: First, it provides a manually annotated multi-label textual news classification dataset for the Urdu language. Second, it benchmarks the performance of traditional machine learning approaches particularly by adapting three data transformation approaches along with three top-performing machine learning classifiers and four algorithm adaptation-based approaches. Third, it benchmarks performance of 16 existing deep learning approaches and the four most widely used language models. Finally, it provides an ensemble approach that reaps the benefits of three different deep learning architectures to precisely predict different classes associated with a particular Urdu textual document. Experimental results reveal that proposed ensemble approach performance values (87% accuracy, 92% F1-score, and 8% hamming loss) are significantly higher than adapted machine and deep learning–based approaches.
Faiza Mehmood, Rehab Shahzadi, Hina Ghafoor, Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Waqar Mahmood, Andreas Dengel 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.7
2023 Performance Comparison of Transformer-Based Models on Twitter Health Mention Classification
abstract
Health mention classification classifies a given piece of text as a health mention or not. However, figurative usage of disease words makes the classification task challenging. To address this challenge, consideration of emojis and surrounding words of the disease names in the text can be helpful. Transformer-based methods are better at capturing the meaning of a word based on its surrounding words compared to traditional methods. However, there are numerous transformer-based methods available and pretrained on natural language processing (NLP) data that are inherently different from Twitter data. Moreover, the size of these models varies in terms of the number of parameters. Hence, it is challenging to decide and choose one of these methods for fine-tuning it on the downstream tasks such as tweet classification. In this work, we experiment with nine widely used transformer methods and compare their performance on the personal health mention classification of tweet data. Furthermore, we analyze the impact of model size on the classification task and provide a brief interpretation of the classification decision made by the best performing classifier. Experimental results show that RoBERTa outperforms all other models by achieving an F1 score of 93%, while two other models perform similarly by achieving an F1 score of 92.5%.
Pervaiz Iqbal Khan, Muhammad Imran Razzak, Andreas Dengel 0001, Sheraz Ahmed
IEEE Trans. Comput. Soc. Syst.3
2022 Search and Learn: Improving Semantic Coverage for Data-to-Text Generation
abstract
Data-to-text generation systems aim to generate text descriptions based on input data (often represented in the tabular form). A typical system uses huge training samples for learning the correspondence between tables and texts. However, large training sets are expensive to obtain, limiting the applicability of these approaches in real-world scenarios. In this work, we focus on few-shot data-to-text generation. We observe that, while fine-tuned pretrained language models may generate plausible sentences, they suffer from the low semantic coverage problem in the few-shot setting. In other words, important input slots tend to be missing in the generated text. To this end, we propose a search-and-learning approach that leverages pretrained language models but inserts the missing slots to improve the semantic coverage. We further finetune our system based on the search results to smooth out the search noise, yielding better-quality text and improving inference efficiency to a large extent. Experiments show that our model achieves high performance on E2E and WikiBio datasets. Especially, we cover 98.35% of input slots on E2E, largely alleviating the low coverage problem.
Shailza Jolly, Zi Xuan Zhang, Andreas Dengel 0001, Lili Mou
AAAI3
2022 Functional Component Descriptions for Electrical Circuits based on Semantic Technology Reasoning
Johannes Bayer, Mina Karami Zadeh, Markus Schröder 0001, Andreas Dengel 0001
DATA4
2022 Time to Focus: A Comprehensive Benchmark using Time Series Attribution Methods
Dominique Mercier, Jwalin Bhatt, Andreas Dengel 0001, Sheraz Ahmed
ICAART (2)3
2022 DT2I: Dense Text-to-Image Generation from Region Descriptions
Stanislav Frolov, Prateek Bansal, Jörn Hees, Andreas Dengel 0001
ICANN (2)4
2022 A Novel Approach to Train Diverse Types of Language Models for Health Mention Classification of Tweets
Pervaiz Iqbal Khan, Muhammad Imran Razzak, Andreas Dengel 0001, Sheraz Ahmed
ICANN (2)3
2022 Audioclip: Extending Clip to Image, Text and Audio
abstract
The rapidly evolving field of sound classification has greatly benefited from the methods of other domains. Today, the trend is to fuse domain-specific tasks and approaches together, which provides the community with new outstanding models.We present AudioCLIP – an extension of the CLIP model that handles audio in addition to text and images. Utilizing the AudioSet dataset, our proposed model incorporates the ESResNeXt audio-model into the CLIP framework, thus enabling it to perform multimodal classification and keeping CLIP’s zero-shot capabilities.AudioCLIP achieves new state-of-the-art results in the Environmental Sound Classification (ESC) task and out-performs others by reaching accuracies of 97.15 % on ESC-50 and 90.07 % on UrbanSound8K. Further, it sets new baselines in the zero-shot ESC-task on the same datasets (69.40 % and 68.78 %, respectively).We also asses the influence of different training setups on the final performance of the proposed model. For the sake of reproducibility, our code is published.
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICASSP4
2022 F2DNet: Fast Focal Detection Network for Pedestrian Detection
abstract
Two-stage detectors are state-of-the-art in object detection as well as pedestrian detection. However, the current two-stage detectors are inefficient as they do bounding box regression in multiple steps i.e. in region proposal networks and bounding box heads. Also, the anchor-based region proposal networks are computationally expensive to train. We propose F2DNet, a novel two-stage detection architecture which eliminates redundancy of current two-stage detectors by replacing the region proposal network with our focal detection network and bounding box head with our fast suppression head. We benchmark F2DNet on top pedestrian detection datasets, thoroughly compare it against the existing state-of-the-art detectors and conduct cross dataset evaluation to test the generalizability of our model to unseen data. Our F2DNet achieves 8.7%, 2.2%, and 6.1% MR2on City Persons, Caltech Pedestrian, and Euro City Person datasets respectively when trained on a single dataset and reaches 20.4% and 26.2% MR2in heavy occlusion setting of Caltech Pedestrian and City Persons datasets when using progressive fine-tunning. Furthermore, F2DNet have significantly lesser inference time compared to the current state-of-the-art. Code and trained models will be available at https://github.com/AbdulHannanKhan/F2DNet.
Abdul Hannan Khan, Mohsin Munir, Ludger van Elst, Andreas Dengel 0001
ICPR4
2022 Are Deep Models Robust against Real Distortions? A Case Study on Document Image Classification
abstract
As deep learning models in the context of document image classification are reaching diminishing returns with nearperfect recognition scores, their robustness characteristics are poorly understood. In order to evaluate the robustness of existing state-of-the-art document image classifiers against different types of distortions that are commonly encountered in the real world, we present two separate benchmark datasets, namely RVL-CDIPD and Tobacco3482-D. The proposed benchmarks are generated by augmenting the well-known pre-existing document image classification datasets (RVL-CDIP and Tobacco3482) with 21 different types of distortions including varying severity levels. We leverage the proposed benchmark datasets to analyze the robustness characteristics of existing document image classification systems. Our analysis reveals that despite higher accuracy models exhibiting relatively higher robustness, they still severely underperform on some specific distortions, with classification accuracies dropping from ~90% to as low as ~40% in some cases. Interestingly, some of these high accuracy models perform even worse than the baseline AlexNet model in the presence of distortions, with the relative decline in their accuracy sometimes reaching as high as 300-450%. We envision these benchmarks to serve as a strong signal of progress in document image classification tasks, beyond the saturated accuracy metrics. The datasets and code to reproduce them is publicly available: https://github.com/saifullah3396/docrobustness.
Saifullah Saifullah, Shoaib Ahmed Siddiqui, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed
ICPR4
2022 Self-supervised Test-time Adaptation on Video Data
abstract
In typical computer vision problems revolving around video data, pre-trained models are simply evaluated at test time, without adaptation. This general approach clearly cannot capture the shifts that will likely arise between the distributions from which training and test data have been sampled. Adapting a pre-trained model to a new video en-countered at test time could be essential to avoid the potentially catastrophic effects of such shifts. However, given the inherent impossibility of labeling data only available at test-time, traditional "fine-tuning" techniques cannot be lever-aged in this highly practical scenario. This paper explores whether the recent progress in test-time adaptation in the image domain and self-supervised learning can be lever-aged to adapt a model to previously unseen and unlabelled videos presenting both mild (but arbitrary) and severe covariate shifts. In our experiments, we show that test-time adaptation approaches applied to self-supervised methods are always beneficial, but also that the extent of their effectiveness largely depends on the specific combination of the algorithms used for adaptation and self-supervision, and also on the type of covariate shift taking place.
Fatemeh Azimi, Sebastian Palacio, Federico Raue, Jörn Hees, Luca Bertinetto, Andreas Dengel 0001
WACV6
2022 CircNet: an encoder-decoder-based convolution neural network (CNN) for circular RNA identification
Marco Stricker, Muhammad Nabeel Asim, Andreas Dengel 0001, Sheraz Ahmed
Neural Comput. Appl.3
2022 Evaluating Privacy-Preserving Machine Learning in Critical Infrastructures: A Case Study on Time-Series Classification
abstract
With the advent of machine learning in applications of critical infrastructure such as healthcare and energy, privacy is a growing concern in the minds of stakeholders. It is pivotal to ensure that neither the model nor the data can be used to extract sensitive information used by attackers against individuals or to harm whole societies through the exploitation of critical infrastructure. The applicability of machine learning in these domains is mostly limited due to a lack of trust regarding the transparency and the privacy constraints. Various safety-critical use cases (mostly relying on time-series data) are currently underrepresented in privacy-related considerations. By evaluating several privacy-preserving methods regarding their applicability on time-series data, we validated the inefficacy of encryption for deep learning, the strong dataset dependence of differential privacy, and the broad applicability of federated methods.
Dominique Mercier, Adriano Lucieri, Mohsin Munir, Andreas Dengel 0001, Sheraz Ahmed
IEEE Trans. Ind. Informatics4
2021 P2P-O: A Purchase-To-Pay Ontology for Enabling Semantic Invoices
Michael Schulze, Markus Schröder 0001, Christian Jilek, Torsten Albers, Heiko Maus, Andreas Dengel 0001
ESWC6
2021 ImpactCite: An XLNet-based Solution Enabling Qualitative Citation Impact Analysis Utilizing Sentiment and Intent
abstract
Citations play a vital role in understanding the impact of scientific literature. Generally, citations are analyzed quantitatively whereas qualitative analysis of citations can reveal deeper insights into the impact of a scientific artifact in the community. Therefore, citation impact analysis (which includes sentiment and intent classification) enables us to quantify the quality of the citations which can eventually assist us in the estimation of ranking and impact. The contribution of this paper is two-fold. First, we benchmark the well-known language models like BERT and ALBERT along with several popular networks for both tasks of sentiment and intent classification. Second, we provide ImpactCite, which is XLNet-based method for citation impact analysis. All evaluations are performed on a set of publicly available citation analysis datasets. Evaluation results reveal that ImpactCite achieves a new state-of-the-art performance for both citation intent and sentiment classification by outperforming the existing approaches by 3.44% and 1.33% in F1-score. Therefore, we emphasize ImpactCite (XLNet-based solution) for both tasks to better understand the impact of a citation. Additional efforts have been performed to come up with CSC-Clean corpus, which is a clean and reliable dataset for citation sentiment classification.
Dominique Mercier, Syed Tahseen Raza Rizvi, Vikas Rajashekar, Andreas Dengel 0001, Sheraz Ahmed
ICAART (2)4
2021 The Person Index Challenge: Extraction of Persons from Messy, Short Texts
abstract
When persons are mentioned in texts with their first name, last name and/or middle names, there can be a high variation which of their names are used, how their names are ordered and if their names are abbreviated. If multiple persons are mentioned consecutively in very different ways, especially short texts can be perceived as "messy". Once ambiguous names occur, associations to persons may not be inferred correctly. Despite these eventualities, in this paper we ask how well an unsupervised algorithm can build a person index from short texts. We define a person index as a structured table that distinctly catalogs individuals by their names. First, we give a formal definition of the problem and describe a procedure to generate ground truth data for future evaluations. To give a first solution to this challenge, a baseline approach is implemented. By using our proposed evaluation strategy, we test the performance of the baseline and suggest further improvements. For future research the source code is publicly available.
Markus Schröder 0001, Christian Jilek, Michael Schulze, Andreas Dengel 0001
ICAART (2)4
2021 Bridging the Technology Gap between Industry and Semantic Web: Generating Databases and Server Code from RDF
abstract
Despite great advances in the area of Semantic Web, industry rather seldom adopts Semantic Web technologies and their storage and query concepts. Instead, relational databases (RDB) are often deployed to store business-critical data, which are accessed via REST interfaces. Yet, some enterprises would greatly benefit from Semantic Web related datasets which are usually represented with the Resource Description Framework (RDF). To bridge this technology gap, we propose a fully automatic approach that generates suitable RDB models with REST APIs to access them. In our evaluation, generated databases from different RDF datasets are examined and compared. Our findings show that the databases sufficiently reflect their counterparts while the API is able to reproduce rather simple SPARQL queries. Potentials for improvements are identified, for example, the reduction of data redundancies in generated databases.
Markus Schröder 0001, Michael Schulze, Christian Jilek, Andreas Dengel 0001
ICAART (2)4
2021 Understanding and Mitigating the Impact of Model Compression for Document Image Classification
Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (1)2
2021 Analyzing the Potential of Zero-Shot Recognition for Document Image Classification
Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed
ICDAR (4)2
2021 L2S-MirLoc: A Lightweight Two Stage MiRNA Sub-Cellular Localization Prediction Framework
abstract
A comprehensive understanding of miRNA sub-cellular localization may leads towards better understanding of physiological processes and support the fixation of diverse irregularities present in a variety of organisms. To date, diverse computational methodologies have been proposed to automatically infer sub-cellular localization of miR-NAs solely using sequence information, however, existing approaches lack in performance. Considering the success of data transformation approaches in Natural Language Processing which primarily transform multi-label classification problem into multi-class classification problem, here, we introduce three different data transformation approaches namely binary relevance, label power set, and classifier chains. Using data transformation approaches, at 1ststage, multi-label miRNA sub-cellular localization problem is transformed into multi-class problem. Then, at 2ndstage, 3 different machine learning classifiers are used to estimate which classifier performs better with what data transformation approach for hand on task. Empirical evaluation on independent test set indicates that L2S-MirLoc selected combination based on binary relevance and deep random forest outperforms state-of-the-art performance values by significant margin.
Muhammad Nabeel Asim, Muhammad Ali Ibrahim, Christoph Zehe, Olivier Cloarec, Rickard Sjögren, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed
IJCNN7
2021 ESResNe(X)t-fbsp: Learning Robust Time-Frequency Transformation of Audio
abstract
Environmental Sound Classification (ESC) is a rapidly evolving field that recently demonstrated the advantages of application of visual domain techniques to the audio-related tasks. Previous studies indicate that the domain-specific modification of cross-domain approaches show a promise in pushing the whole area of ESC forward. In this paper, we present a new time-frequency transformation layer that is based on complex frequency B-spline (fbsp) wavelets. Being used with a high-performance audio classification model, the proposed fbsp-layer provides an accuracy improvement over the previously used Short-Time Fourier Transform (STFT) on standard datasets. We also investigate the influence of different pre-training strategies, including the joint use of two large-scale datasets for weight initialization: ImageNet and AudioSet. Our proposed model out-performs other approaches by achieving accuracies of 95.20 % on the ESC-50 and 89.14 % on the UrbanSound8K datasets. Additionally, we assess the increase of model robustness against additive white Gaussian noise and reduction of an effective sample rate introduced by the proposed layer and demonstrate that the fbsp-layer improves the model's ability to withstand signal perturbations, in comparison to STFT-based training. For the sake of reproducibility, our code is made available.
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel 0001
IJCNN4
2021 ANP-W2V: Effects of Composition Methods for Embedding Adjective-Noun Pairs
abstract
Adjective-Noun Pairs (ANPs) are often used for affective computing in textual and visual domains. Due to the training cost of more current models (e.g., transformers) many approaches still rely on Word2Vec to compute ANP embeddings by combining individual word embeddings. However, when combining an adjective and a noun into an ANP, there can potentially be more complex interactions which cannot be accounted for by using Word2Vec alone. To solve these challenges, we propose ANP- W2V, an approach that puts adjectives and nouns in different embedding spaces and hence outperforms the baselines based on Word2Vec. In this paper, we do a comprehensive comparison, where we systematically evaluate the role of six different fusion methods in four different tasks with different embedding sizes. The chosen tasks not only gauge the external and internal relationships in the ANPs but can also check whether ANP embeddings capture more complex interactions between adjectives and nouns.
Dayananda Herurkar, Philipp Blandfort, Federico Raue, Jörn Hees, Andreas Dengel 0001
IJCNN5
2021 DeepCeNS: An end-to-end Pipeline for Cell and Nucleus Segmentation in Microscopic Images
abstract
With the evolution of deep learning in the past decade, more biomedical related problems that seemed strenuous, are now feasible. The introduction of U-net and Mask R-CNN architectures has paved a way for many object detection and segmentation tasks in numerous applications ranging from security to biomedical applications. In the cell biology domain, light microscopy imaging provides a cheap and accessible source of raw data to study biological phenomena. By leveraging such data and deep learning techniques, human diseases can be easily diagnosed and the process of treatment development can be greatly expedited. In microscopic imaging, accurate segmentation of individual cells is a crucial step to allow better insight into cellular heterogeneity. To address the aforementioned challenges, DeepCeNS is proposed in this paper to detect and segment cells and nucleus in microscopic images. We have used EVICAN2 dataset which contains microscopic images from a variety of microscopes having numerous cell cultures, to evaluate the proposed pipeline. DeepCeNS outperforms EVICAN-MRCNN by a significant margin on the EVICAN2 dataset.
Nabeel Khalid, Mohsin Munir, Christoffer Edlund, Timothy R. Jackson, Johan Trygg, Rickard Sjögren, Andreas Dengel 0001, Sheraz Ahmed
IJCNN7
2021 PatchX: Explaining Deep Models by Intelligible Pattern Patches for Time-series Classification
abstract
Classification of time-series data is pivotal for a wide range of applications and comes with many challenges. Although the amount of publicly available datasets increases rapidly, deep neural models are only partially exploited in contrast to the traditional methods. These methods get preferred in safety-critical, financial, or medical fields because of their interpretable results. However, their performance and scalability are limited, and finding suitable explanations for time-series classification tasks is challenging due to the intrinsic nature of concealed concepts in the time-series data. Visual analysis of complete time-series comes with an extensive cognitive overload, as it is difficult to perceive and leads to confusion. Therefore, we believe that patch-wise processing of the data results in a more interpretable representation. To bridge this gap, and to reduce the cognitive overload for interpretation of time series data, we propose a novel hybrid approach that utilizes deep neural networks and traditional machine learning algorithms for an interpretable and scale-able time-series classification approach. Both quantitively and qualitatively PatchX shows superiority to its counterparts with an edge of interoperability.
Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed
IJCNN2
2021 Spread2RML: Constructing Knowledge Graphs by Predicting RML Mappings on Messy Spreadsheets
abstract
The RDF Mapping Language (RML) allows to map semi-structured data to RDF knowledge graphs. Besides CSV, JSON and XML, this also includes the mapping of spreadsheet tables. Since spreadsheets have a complex data model and can become rather messy, their mapping creation tends to be very time consuming. In order to reduce such efforts, this paper presents Spread2RML which predicts RML mappings on messy spreadsheets. This is done with an extensible set of RML object map templates which are applied for each column based on heuristics. In our evaluation, three datasets are used ranging from very messy synthetic data to spreadsheets from data.gov which are less messy. We obtained first promising results especially with regard to our approach being fully automatic and dealing with rather messy data.
Markus Schröder 0001, Christian Jilek, Andreas Dengel 0001
K-CAP3
2021 Correction to: Benchmarking performance of machine and deep learning-based methodologies for Urdu text document classification
Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Ali Ibrahim, Waqar Mahmood, Andreas Dengel 0001, Sheraz Ahmed
Neural Comput. Appl.5
2021 Benchmarking performance of machine and deep learning-based methodologies for Urdu text document classification
Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Ali Ibrahim, Waqar Mahmood, Andreas Dengel 0001, Sheraz Ahmed
Neural Comput. Appl.5
2021 Adversarial text-to-image synthesis: A review
abstract
With the advent of generative adversarial networks, synthesizing images from text descriptions has recently become an active research area. It is a flexible and intuitive way for conditional image generation with significant progress in the last years regarding visual realism, diversity, and semantic alignment. However, the field still faces several challenges that require further research efforts such as enabling the generation of high-resolution images with multiple objects, and developing suitable and reliable evaluation metrics that correlate with human judgement. In this review, we contextualize the state of the art of adversarial text-to-image synthesis models, their development since their inception five years ago, and propose a taxonomy based on the level of supervision. We critically examine current strategies to evaluate text-to-image synthesis models, highlight shortcomings, and identify new areas of research, ranging from the development of better datasets and evaluation metrics to possible improvements in architectural design and model training. This review complements previous surveys on generative adversarial networks with a focus on text-to-image synthesis which we believe will help researchers to further advance the field.
Stanislav Frolov, Tobias Hinz, Federico Raue, Jörn Hees, Andreas Dengel 0001
Neural Networks5
2020 From Automatic Keyword Detection to Ontology-Based Topic Modeling
Marc Beck, Syed Tahseen Raza Rizvi, Andreas Dengel 0001, Sheraz Ahmed
DAS3
2020 Classification of Visual Strategies in Physics Vector Field Problem-solving
Seyyed Saleh Mozafari Chanijani, Mohammad Al-Naser, Pascal Klein 0002, Stefan Küchemann, Jochen Kuhn, Thomas Widmann, Andreas Dengel 0001
ICAART (2)7
2020 DartsReNet: Exploring New RNN Cells in ReNet Architectures
Brian B. Moser, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (1)4
2020 Improved and Visually Enhanced Case-Based Retrieval of Room Configurations for Assistance in Architectural Design Education
Viktor Eisenstadt, Christoph Langenhan, Klaus-Dieter Althoff, Andreas Dengel 0001
ICCBR4
2020 Enhancer-DSNet: A Supervisedly Prepared Enriched Sequence Representation for the Identification of Enhancers and Their Strength
Muhammad Nabeel Asim, Muhammad Ali Ibrahim, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (3)4
2020 Improving Personal Health Mention Detection on Twitter Using Permutation Based Word Representation Learning
Pervaiz Iqbal Khan, Muhammad Imran Razzak, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (1)3
2020 Explaining AI-Based Decision Support Systems Using Concept Localization Maps
abstract
Human-centric explainability of AI-based Decision Support Systems (DSS) using visual input modalities is directly related to reliability and practicality of such algorithms. An otherwise accurate and robust DSS might not enjoy trust of experts in critical application areas if it is not able to provide reasonable justification of its predictions. This paper introduces Concept Localization Maps (CLMs), which is a novel approach towards explainable image classifiers employed as DSS. CLMs extend Concept Activation Vectors (CAVs) by locating significant regions corresponding to a learned concept in the latent space of a trained image classifier. They provide qualitative and quantitative assurance of a classifier's ability to learn and focus on similar concepts important for humans during image recognition. To better understand the effectiveness of the proposed method, we generated a new synthetic dataset called Simple Concept DataBase (SCDB) that includes annotations for 10 distinguishable concepts, and made it publicly available. We evaluated our proposed method on SCDB as well as a real-world dataset called CelebA. We achieved localization recall of above 80% for most relevant concepts and average recall above 60% for all concepts using SE-ResNeXt-50 on SCDB. Our results on both datasets show great promise of CLMs for easing acceptance of DSS in practice.
Adriano Lucieri, Muhammad Naseer Bajwa, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (4)3
2020 P2ExNet: Patch-Based Prototype Explanation Network
Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (3)2
2020 Benchmarking Adversarial Attacks and Defenses for Time-Series Data
Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (3)2
2020 Revisiting Sequence-to-Sequence Video Object Segmentation with Multi-Task Loss and Skip-Memory
abstract
Video Object Segmentation (VOS) is an active research area of the visual domain. One of its fundamental subtasks is semi-supervised / one-shot learning: given only the segmentation mask for the first frame, the task is to provide pixel-accurate masks for the object over the rest of the sequence. Despite much progress in the last years, we noticed that many of the existing approaches lose objects in longer sequences, especially when the object is small or briefly occluded. In this work, we build upon a sequence-to-sequence approach that employs an encoder-decoder architecture together with a memory module for exploiting the sequential data. We further improve this approach by proposing a model that manipulates multiscale spatio-temporal information using memory-equipped skip connections. Furthermore, we incorporate an auxiliary task based on distance classification which greatly enhances the quality of edges in segmentation masks. We compare our approach to the state of the art and show considerable improvement in the contour accuracy metric and the overall segmentation accuracy. Our source code and the pre-trained weights are publicly available11https://github.com/fatemehazimi990/RS2S.
Fatemeh Azimi, Benjamin Bischke, Sebastian Palacio, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICPR6
2020 ESResNet: Environmental Sound Classification Based on Visual Domain Models
abstract
Environmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years. However, many of the existing approaches achieve high accuracy by relying on domain-specific features and architectures, making it harder to benefit from advances in other fields (e.g., the image domain). Additionally, some of the past successes have been attributed to a discrepancy of how results are evaluated (i.e., on unofficial splits of the UrbanSound8K (US8K) dataset), distorting the overall progression of the field. The contribution of this paper is twofold. First, we present a model that is inherently compatible with mono and stereo sound inputs. Our model is based on simple log-power Short-Time Fourier Transform (STFT) spectrograms and combines them with several well-known approaches from the image domain (i.e., ResNet, Siamese-like networks and attention). We investigate the influence of cross-domain pre-training, architectural changes, and evaluate our model on standard datasets. We find that our model out-performs all previously known approaches in a fair comparison by achieving accuracies of 97.0 % (ESC-10), 91.5 % (ESC-50) and 84.2 % / 85.4 % (US8K mono / stereo). Second, we provide a comprehensive overview of the actual state of the field, by differentiating several previously reported results on the US8K dataset between official or unofficial splits. For better reproducibility, our code (including any re- implementations) is made available.
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICPR4
2020 P ≈ NP, at least in Visual Question Answering
abstract
In recent years, progress in the Visual Question Answering (VQA) field has largely been driven by public challenges and large datasets. One of the most widely-used of these is the VQA 2.0 dataset, consisting of polar (“yes/no”) and non-polar questions. Looking at the question distribution over all answers, we find that the answers “yes” and “no” account for 38% of the questions (19% per class), while the remaining 62% are spread over the remaining 3127 answers (0.02% per class). While several sources of biases have been investigated in the field, the effects of such an over-representation of polar questions remain unclear. In this paper, we measure the potential confounding factors when polar and non-polar samples are used jointly to train a baseline VQA classifier, and compare it to an upper bound where the over-representation of polar questions is excluded from the training. Further, we perform cross-over experiments to analyze how well the feature spaces of polar and non-polar samples align. Contrary to expectations, we find no evidence of counterproductive effects in the joint training of unbalanced classes. In fact, by exploring the intermediate feature space of visual-text embeddings, we find that the feature space of polar questions already encodes sufficient structure to answer many non-polar questions. Our results indicate that the polar (P) and the non-polar (NP) feature spaces are strongly aligned, hence the expression P ≈ NP.
Shailza Jolly, Sebastian Palacio, Joachim Folz, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICPR6
2020 Contextual Classification Using Self-Supervised Auxiliary Models for Deep Neural Networks
abstract
The following topics are dealt with: learning (artificial intelligence); feature extraction; neural nets; image classification; object detection; image segmentation; image representation; video signal processing; pattern classification; computer vision.
Sebastian Palacio, Philipp Engler, Jörn Hees, Andreas Dengel 0001
ICPR4
2020 K-mer Neural Embedding Performance Analysis Using Amino Acid Codons
abstract
Exponential growth of genome-wide assays of gene expressions and their public access open new horizons for machine learning methodologies to effectively perform genetic analysis. In this work, domain specific pre-train k-mer embeddings of DNA sequences are generated by utilising FastText approach. Sequence co-expression pattern information is embedded into 200 dimensional vectors by training Fasttext model on 317,151 samples of DNA sequences (with k-mers representation). We propose a novel idea to utilize the information of various codons present in amino acids for the evaluation of learned sequence vectors. We employ two diverse techniques to compare the performance of generated task-specific k-mer embeddings with state-of-the-art publicly available generic k-mer embeddings of genome. Firstly, we utilize a dimensionality reduction approach namely PCA to alleviate the dimensions of DNA sequences upto 50 features by preserving almost 85% of sequence features information. Afterwards, TSNE algorithms is used to visualize k-mer embeddings and to make sure whether different codons representing the same amino acid are more closer to each other than the ones representing different amino acids. Secondly, to assess the analogy of k-mer embeddings, generated domain specific k-mer embeddings are compared with state-of-the-art k-mer embeddings by estimating the cosine similarity among those codons vectors which represent same amino acid. Overall, we believe that task-specific distributed representation of k-mers would be useful for DNA methylation and Histone occupancy prediction tasks.
Muhammad Nabeel Asim, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
IJCNN3
2020 G1020: A Benchmark Retinal Fundus Image Dataset for Computer-Aided Glaucoma Detection
abstract
Scarcity of large publicly available retinal fundus image datasets for automated glaucoma detection has been the bottleneck for successful application of artificial intelligence towards practical Computer-Aided Diagnosis (CAD). A few small datasets that are available for research community usually suffer from impractical image capturing conditions and stringent inclusion criteria. These shortcomings in already limited choice of existing datasets make it challenging to mature a CAD system so that it can perform in real-world environment. In this paper we present a large publicly available retinal fundus image dataset for glaucoma classification called G1020. The dataset is curated by conforming to standard practices in routine ophthalmology and it is expected to serve as standard benchmark dataset for glaucoma detection. This database consists of 1020 high resolution colour fundus images and provides ground truth annotations for glaucoma diagnosis, optic disc and optic cup segmentation, vertical cup-to-disc ratio, size of neuroretinal rim in inferior, superior, nasal and temporal quadrants, and bounding box location for optic disc. We also report baseline results by conducting extensive experiments for automated glaucoma diagnosis and segmentation of optic disc and optic cup.
Muhammad Naseer Bajwa, Gur Amrit Pal Singh, Wolfgang Neumeier, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
IJCNN5
2020 Conceptual Explanations of Neural Network Prediction for Time Series
abstract
Deep neural networks are black boxes by construction. Explanation and interpretation methods therefore are pivotal for a trustworthy application. Existing methods are mostly based on heatmapping and focus on locally determining the relevant input parts triggering the network prediction. However, these methods struggle to uncover global causes. While this is a rare case in the image or NLP modality, it is of high relevance in the time series domain. This paper presents a novel framework, i.e. Conceptual Explanation, designed to evaluate the effect of abstract (local or global) input features on the model behavior. The method is model-agnostic and allows utilizing expert knowledge. On three time series datasets Conceptual Explanation demonstrates its ability to pinpoint the causes inherent to the data to trigger the correct model prediction.
Ferdinand Küsters, Peter Schichtel, Sheraz Ahmed, Andreas Dengel 0001
IJCNN4
2020 On Interpretability of Deep Learning based Skin Lesion Classifiers using Concept Activation Vectors
abstract
Deep learning based medical image classifiers have shown remarkable prowess in various application areas like ophthalmology, dermatology, pathology, and radiology. However, the acceptance of these Computer-Aided Diagnosis (CAD) systems in real clinical setups is severely limited primarily because their decision-making process remains largely obscure. This work aims at elucidating a deep learning based medical image classifier by verifying that the model learns and utilizes similar disease-related concepts as described and employed by dermatologists. We used a well-trained and high performing neural network developed by REasoning for COmplex Data (RECOD) Lab for classification of three skin tumours, i.e. Melanocytic Naevi, Melanoma and Seborrheic Keratosis and performed a detailed analysis on its latent space. Two well established and publicly available skin disease datasets, PH2and derm7pt, are used for experimentation. Human understandable concepts are mapped to RECOD image classification model with the help of Concept Activation Vectors (CAVs), introducing a novel training and significance testing paradigm for CAVs. Our results on an independent evaluation set clearly shows that the classifier learns and encodes human understandable concepts in its latent representation. Additionally, TCAV scores (Testing with CAVs) suggest that the neural network indeed makes use of disease-related concepts in the correct way when making predictions. We anticipate that this work can not only increase confidence of medical practitioners on CAD but also serve as a stepping stone for further development of CAV-based neural network interpretation methods.
Adriano Lucieri, Muhammad Naseer Bajwa, Stephan Alexander Braun, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
IJCNN5
2020 Interpreting Deep Models through the Lens of Data
abstract
Identification of input data points relevant for the classifier (i.e. serve as the support vector) has recently spurred the interest of researchers for both interpretability as well as dataset debugging. This paper presents an in-depth analysis of the methods which attempt to identify the influence of these data points on the resulting classifier. To quantify the quality of the influence, we curated a set of experiments where we debugged and pruned the dataset based on the influence information obtained from different methods. To do so, we provided the classifier with mislabeled examples that hampered the overall performance. Since the classifier is a combination of both the data and the model, therefore, it is essential to also analyze these influences for the interpretability of deep learning models. Analysis of the results shows that some interpretability methods can detect mislabels better than using a random approach, however, contrary to the claim of these methods, the sample selection based on the training loss showed a superior performance.
Dominique Mercier, Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed
IJCNN3
2020 DeepKAF: A Heterogeneous CBR & Deep Learning Approach for NLP Prototyping
abstract
With widespread modernization, digitization and transformations of most of industries, Artificial Intelligence (AI) has become the key enabler in that modernization journey. AI offers substantial capabilities to solve new problems and optimise existing solutions specialising on specific problems and learning from different domains. AI solutions can be either explainable or black box ones with the latter being urged to improve since they cannot trust. Case-based Reasoning (CBR) is an explainable AI approach where solutions are provided along with relevant explanations in terms of why a solution was selected. However, CBR, like most other explainable approaches, has several limitations in terms of scalability, large data volumes, domain complexity, that reduce its ability to scale any CBR system in industrial applications. In this paper, we provide a heterogeneous CBR framework - DeepKAF where we combine CBR paradigm with Deep Learning architectures to solve complicated Natural Language Processing (NLP) problems (eg. mixed language and grammatically incorrect text).DeepKAF is built based on continuous research in the area of Deep Learning and CBR. DeepKAF has been implemented and used across different domains, test use cases and research models as an ensemble deep learning and CBR Architecture.
Kareem Amin 0001, Stelios Kapetanakis, Nikolaos Polatidis, Klaus-Dieter Althoff, Andreas Dengel 0001
INISTA5
2020 Adversarial Defense based on Structure-to-Signal Autoencoders
abstract
Adversarial attacks have exposed the intricacies of the complex loss surfaces approximated by neural networks. In this paper, we present a defense strategy against gradient-based attacks, on the premise that input gradients need to expose information about the semantic manifold for attacks to be successful. We propose an architecture based on compressive autoencoders (AEs) with a two-stage training scheme, creating not only an architectural bottleneck but also a representational bottleneck. We show that the proposed mechanism yields robust results against a collection of gradient-based attacks under challenging white-box conditions. This defense is attack-agnostic and can, therefore, be used for arbitrary pre-trained models, while not compromising the original performance. These claims are supported by experiments conducted with state-of-the-art image classifiers (ResNet50 and Inception v3), on the full ImageNet validation set. Experiments, including counterfactual analysis, empirically show that the robustness stems from a shift in the distribution of input gradients, which mitigates the effect of tested adversarial attack methods. Gradients propagated through the proposed AEs represent less semantic information and instead point to low-level structural features.
Joachim Folz, Sebastian Palacio, Jörn Hees, Andreas Dengel 0001
WACV4
2019 A Reinforcement Learning Approach for Sequential Spatial Transformer Networks
Fatemeh Azimi, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (1)4
2019 DeepEX: Bridging the Gap Between Knowledge and Data Driven Techniques for Time Series Forecasting
Muhammad Ali Chattha, Shoaib Ahmed Siddiqui, Mohsin Munir, Muhammad Imran Malik, Ludger van Elst, Andreas Dengel 0001, Sheraz Ahmed
ICANN (2)6
2019 Conditional GANs for Image Captioning with Sentiments
Tushar Karayil, Asif Irfan, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (4)5
2019 Comparison Between U-Net and U-ReNet Models in OCR Tasks
Brian B. Moser, Federico Raue, Jörn Hees, Andreas Dengel 0001
ICANN (3)4
2019 A Robust Hybrid Approach for Textual Document Classification
abstract
Text document classification is an important task for diverse natural language processing based applications. Traditional machine learning approaches mainly focused on reducing dimensionality of textual data to perform classification. This although improved the overall classification accuracy, the classifiers still faced sparsity problem due to lack of better data representation techniques. Deep learning based text document classification, on the other hand, benefitted greatly from the invention of word embeddings that have solved the sparsity problem and researchers focus mainly remained on the development of deep architectures. Deeper architectures, however, learn some redundant features that limit the performance of deep learning based solutions. In this paper, we propose a two stage text document classification methodology which combines traditional feature engineering with automatic feature engineering (using deep learning). The proposed methodology comprises a filter based feature selection (FSE) algorithm followed by a deep convolutional neural network. This methodology is evaluated on the two most commonly used public datasets, i.e., 20 Newsgroups data and BBC news data. Evaluation results reveal that the proposed methodology outperforms the state-of-the-art of both the (traditional) machine learning and deep learning based text document classification methodologies with a significant margin of 7.7% on 20 Newsgroups and 6.6% on BBC news datasets.
Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
ICDAR4
2019 Two Stream Deep Network for Document Image Classification
abstract
This paper presents a novel two-stream approach for document image classification. The proposed approach leverages textual and visual modalities to classify document images into ten categories, including letter, memo, news article, etc. In order to alleviate dependency of textual stream on performance of underlying OCR (which is the case with general content based document image classifiers), we utilize a filter based feature-ranking algorithm. This algorithm ranks the features of each class based on their ability to discriminate document images and selects a set of top 'K' features that are retained for further processing. In parallel, the visual stream uses deep CNN models to extract structural features of document images.Finally, textual and visual streams are concatenated together using an average ensembling method. Experimental results reveal that the proposed approach outperforms the state-of-the-art system with a significant margin of 4.5% on publicly available Tobacco-3482 dataset.
Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Imran Malik, Khizar Razzaque, Andreas Dengel 0001, Sheraz Ahmed
ICDAR5
2019 Chemical Structure Recognition (CSR) System: Automatic Analysis of 2D Chemical Structures in Document Images
abstract
In this era of advanced technology and automation, information extraction has become a very common practice for the analysis of data. A technique known as Optical Character Recognition (OCR) is used for recognition of text. The purpose is to extract textual data for automatic information analysis or natural language processing of document images. However, in the field of cheminformatics where it is required to recognize 2D molecular structures as they are published in research journals or patent documents, OCR is not adequate for processing, as chemical compounds can be represented both in textual as well as in graphical format. The digital representation of an image based chemical structure allows not only patent analysis teams to provide customize insights but also cheminformatic research groups to enhance their molecular structure databases, which further can be used for querying structure as well as sub-structural patterns. Some tools have been made for extraction and processing of image-based molecular structures. Optical Structure Recognition Application (OSRA) being one of the tools that partially fulfill the task of recognizing chemical structural in document images into chemical formats (SMILES, SDF, or MOL). However, it has few problems such as poor character recognition, false structure extraction, and slow processing. In this paper, we have developed a prototype Chemical Structure Recognition (CSR) system using modern and advanced image processing open-source libraries, which allows us to extract structural information of a chemical structure embedded in the form of a digital raster image. The CSR system is capable of processing chemical information contained in chemical structure image and generates the SMILES or MOL representation. For performance evaluation, we have used two different data sets to measure the potential of the CSR system. It yields better results than OSRA that depict accurate recognition, fast extraction, and correctness of great significance.
Syed Saqib Bukhari, Zaryab Iftikhar, Andreas Dengel 0001
ICDAR3
2019 Analysis of Unsupervised Training Approaches for LSTM-Based OCR
abstract
In the context of historical documents, where labeled training data is especially expensive to acquire, the prospect of using unlabeled training data to improve and speed up the overall training is desirable. The most common way to use unlabeled data is unsupervised pretraining, which has been successfully applied to various CNN and RNN architectures in different domains. There is however not sufficient work for its application in the field of OCR. In this paper we investigate multiple architectures and how unlabeled data could be applied to them. We show that in combination with Connectionist Temporal Classification (CTC), a reconstruction objective has no apparent synergistic effect, with both objectives learning different representations. We therefore investigate the use of an LSTM-based Seq2Seq OCR architecture which shows promise regarding unsupervised pretraining.
Martin Jenckel, Syed Saqib Bukhari, Andreas Dengel 0001
ICDAR3
2019 DeepTabStR: Deep Learning based Table Structure Recognition
abstract
This paper presents a novel method for the analysis of tabular structures in document images using the potential of deformable convolutional networks. In order to assess the suitability of the model to the task of table structure recognition, most of the prior methods have been tested on the smaller ICDAR-13 table structure recognition dataset comprising of just 156 tables. We curated a new image-based table structure recognition dataset, TabStructDB2, comprising of 1081 tables densely labeled with row and column information. Instead of collecting new images for this purpose, we leveraged the famous Page-Object Detection dataset from ICDAR-17, and added structural information for all the tabular regions present in the dataset. This new publicly available dataset will enable the development of more sophisticated table structure recognition techniques in the future. We performed extensive evaluation on the two datasets (ICDAR-13 and TabStructDB) including cross-dataset testing in order to evaluate the efficacy of the proposed approach. We achieved state-of-the-art results with deformable models on ICDAR-13 with an average F-Measure of 92.98% (89.42% for rows and 96.55% for columns) and report baseline results on TabStructDB for guiding future research efforts with an F-Measure of 93.72% (91.26% for rows and 95.59% for columns). Despite promising results, structural analysis of tables with arbitrary layouts is still far from achievable at this point.
Shoaib Ahmed Siddiqui, Imran Ali Fateh, Syed Tahseen Raza Rizvi, Andreas Dengel 0001, Sheraz Ahmed
ICDAR4
2019 Rethinking Semantic Segmentation for Table Structure Recognition in Documents
abstract
Based on the recent advancements in the domain of semantic segmentation, Fully-Convolutional Networks (FCN) have been successfully applied for the task of table structure recognition in the past. We analyze the efficacy of semantic segmentation networks for this purpose and simplify the problem by proposing prediction tiling based on the consistency assumption which holds for tabular structures. For an image of dimensions H × W, we predict a single column for the rows (ŷrowϵ H) and a predict a single row for the columns (ŷrowϵ W). We use a dual-headed architecture where initial feature maps (from the encoder-decoder model) are shared while the last two layers generate class specific (row/column) predictions. This allows us to generate predictions using a single model for both rows and columns simultaneously, where previous methods relied on two separate models for inference. With the proposed method, we were able to achieve state-of-the-art results on ICDAR-13 image-based table structure recognition dataset with an average F-Measure of 92.39% (91.90% and 92.88% F-Measure for rows and columns respectively). With the proposed method, we were able to achieve state-of-the-art results on ICDAR-13. The obtained results advocate that constraining the problem space in the case of FCN by imposing valid constraints can lead to significant performance gains.
Shoaib Ahmed Siddiqui, Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed
ICDAR3
2019 Unsupervised OCR Model Evaluation Using GAN
abstract
Optical Character Recognition (OCR) has achieved its state-of-the-art performance with the use of Deep Learning for character recognition. Deep Learning techniques need large amount of data along with ground truth. Out of the available data, small portion of it has to be used for validation purpose as well. Preparing ground truth for historical documents is expensive and hence availability of data is of utmost concern. Jenckel et al. jenckel came up with an idea of using all the available data for training the OCR model and for the purpose of validation, they generated the input image from Softmax layer of the OCR model; using the decoder setup which can be used to compare with the original input image to validate the OCR model. In this paper, we have explored the possibilities of using Generative Adversial Networks (GANs) [6] for generating the image directly from the text obtained from OCR model instead of using the Softmax layer which is not always accessible for all the Deep Learning based OCR models. Using text directly to generate the input image back gives us the advantage to use this pipeline for any OCR models even whose Softmax layer is not accessible. In the results section, we have shown that the current state of using GANs for unsupervised OCR model evaluation.
Abhash Sinha, Martin Jenckel, Syed Saqib Bukhari, Andreas Dengel 0001
ICDAR4
2019 Multi-Task Learning for Segmentation of Building Footprints with Deep Neural Networks
abstract
The increased availability of high-resolution satellite imagery allows to sense very detailed structures on the surface of our planet. Access to such information opens up new directions in the analysis of remote sensing imagery. While deep neural networks have achieved significant advances in semantic segmentation of high-resolution images, most of the existing approaches tend to produce predictions with poor boundaries. In this paper, we address the problem of preserving semantic segmentation boundaries in high-resolution satellite imagery by introducing a novel multi-task loss. The loss leverages multiple output representations of the segmentation mask and biases the network to focus more on pixels near boundaries. We evaluate our approach on the large-scale Inria Aerial Image Labeling Dataset which contains high-resolution images. Our results show that we are able to outperform state-of-the-art methods by 9.8% on the Intersection over Union (IoU) metric without any additional post-processing steps. Source code and all models will be available under https://github.com/bbischke/MultiTaskBuildingSegmentation.
Benjamin Bischke, Patrick Helber, Joachim Folz, Damian Borth, Andreas Dengel 0001
ICIP5
2019 A Comparative Analysis of Traditional and Deep Learning-Based Anomaly Detection Methods for Streaming Data
abstract
With the Internet of Things (IoT) devices becoming an integral part of human life, the need for robust anomaly detection in streaming data has also been elevated. Dozens of distance-based, density-based, kernel-based, and cluster-based algorithms have been proposed in the area of anomaly detection. Recently, because of the robustness of the deep neural networks (DNN), different deep learning-based anomaly detection methods have also been proposed. With all these rapid developments, there exists a small number of comparative studies for anomaly detection methods. Even in those studies, the comparison is done only in typical anomaly detection settings without taking the streaming data into consideration. The presence of intrinsic time-series characteristics like trend, seasonality, and change-point makes it important to study the behavior of commonly used anomaly detection methods on streaming data. Moreover, the comparison of traditional methods with deep learning-based methods also brings exciting insights about the data which are generally overlooked by traditional methods. In this study, we compare 13 anomaly detection methods on two commonly used streaming data sets. We used four different evaluation metrics to evaluate the methods from different perspectives. Our analysis reveals that the deep learning-based anomaly detection methods are superior to traditional anomaly detection methods.
Mohsin Munir, Muhammad Ali Chattha, Andreas Dengel 0001, Sheraz Ahmed
ICMLA3
2019 Removal of Historical Document Degradations using Conditional GANs
abstract
One of the most crucial problem in document analysis and OCR pipeline is document binarization. Many traditional algorithms over the past few decades like Sauvola, Niblack, Otsu etc,. were used for binarization which gave insufficient results for historical texts with degradations. Recently many attempts have been made to solve binarization using deep learning approaches like Autoencoders, FCNs. However, these models do not generalize well to real world historical document images qualitatively. In this paper, we propose a model based on conditional GAN, well known for its high-resolution image synthesis. Here, the proposed model is used for image manipulation task which can remove different degradations in historical documents like stains, bleed-through and non-uniform shadings. The performance of the proposed model outperforms recent state-of-the-art models for document image binarization. We support our claims by benchmarking the proposed model on publicly available PHIBC 2012, DIBCO (2009-2017) and Palm Leaf datasets. The main objective of this paper is to illuminate the advantages of generative modeling and adversarial training for document image binarization in supervised setting which shows good generalization capabilities on different inter/intra class domain document images.
Veeru Dumpala, Sheela Raju Kurupathi, Syed Saqib Bukhari, Andreas Dengel 0001
ICPRAM4
2019 A Study of Various Text Augmentation Techniques for Relation Classification in Free Text
abstract
Item does not contain fulltext
Praveen Badimala, Chinmaya Mishra, Reddy Kumar Modam Venkataramana, Syed Saqib Bukhari, Andreas Dengel 0001
ICPRAM5
2019 Document Image Dewarping using Deep Learning
abstract
The distorted images have been a major problem for Optical Character Recognition (OCR). In order to perform OCR on distorted images, dewarping has become a principal preprocessing step. This paper presents a new document dewarping method that removes curl and geometric distortion of modern and historical documents. Finally, the proposed method is evaluated and compared to the existing Computer Vision based method. Most of the traditional dewarping algorithms are created based on the text line feature extraction and segmentation. However, textual content extraction and segmentation can be sophisticated. Hence, the new technique is proposed, which doesn’t need any complicated methods to process the text lines. The proposed method is based on Deep Learning and it can be applied on all type of text documents and also documents with images and graphics. Moreover, there is no preprocessing required to apply this method on warped images. In the proposed system, the document distortion problem is treated as an image-to-image translation. The new method is implemented using a very powerful pix2pixhd network by utilizing Conditional Generative Adversarial Networks (CGAN). The network is trained on UW3 dataset by supplying distorted document as an input and cleaned image as the target. The generated images from the proposed method are cleanly dewarped and they are of high-resolution. Furthermore, these images can be used to perform OCR.
Vijaya Kumar Bajjer Ramanna, Syed Saqib Bukhari, Andreas Dengel 0001
ICPRAM3
2019 A Robust Page Frame Detection Method for Complex Historical Document Images
abstract
Document layout analysis is the most important part of converting scanned page images into search-able full text. An intensive amount of research is going on in the field of structured and semi-structured documents (journal articles, books, magazines, invoices) but not much in historical documents. Historical document digitization is a more challenging task than regular structured documents due to poor image quality, damaged characters, big amount of textual and non-textual noise. In the scientific community, the extraneous symbols from the neighboring page are considered as textual noise, while the appearances of black borders, speckles, ruler, different types of image etc. along the border of the documents are considered as non-textual noise. Existing historical document analysis method cannot handle all of this noise which is a very strong reason of getting undesired texts as a result from the output of Optical Character Recognition (OCR) that needs to be removed afterward with a lot of extra afford. This paper presents a new perspective especially for the historical document image cleanup by detecting the page frame of the document. The goal of this method is to find actual contents area of the document and ignore noises along the page border. We use morphological transforms, the line segment detector, and geometric matching algorithm to find an ideal page frame of the document. After the implementation of page frame method, we also evaluate our approach over 16th-19th century printed historical documents. We have noticed in the result that OCR performance for the historical documents increased by 4.49% after applying our page frame detection method. In addition, we are able to increase the OCR accuracy around 6.69% for contemporary documents too.
Mohammad Mohsin Reza, Md. Ajraf Rakib, Syed Saqib Bukhari, Andreas Dengel 0001
ICPRAM4
2019 Location-Specific Embedding Learning for the Semantic Segmentation of Building Footprints on a Global Scale
abstract
In this paper, we analyze the feasability of learning a latent embedding space from aerial and satellite imagery in order to capture semantic properties of geographical locations. We show that deep neural network, trained with a triplet loss function, can be effectively used to obtain a location-specific embedding. Considering the problem of building footprint segmentation from aerial imagery of varying cities, we leverage these embeddings together with a clustering for the training of location-specific segmentation networks and the selection of the corresponding segmentation network during inference time. We evaluate our approach on the large-scale Inria Aerial Image Labeling Dataset which contains aerial images of globally distributed cities. Our approach achieves an outperformance against state-of-the-art approaches on the Intersection over Union metric for the building class over all cities and by more than 2% for specific cities.
Benjamin Bischke, Patrick Helber, Jörn Hees, Andreas Dengel 0001
IGARSS4
2019 Multi-Scale Machine Learning for the Classification of Building Property Values
abstract
In this paper, we describe a multi-scale machine learning approach to estimate socio-economic attributes of citizens based on the analysis of aerial images. To analyse the effectiveness of the proposed approach we predict building property value classes. The classification of these building property values is a proxy for the socio-economic status of the residents. The approach is based on the fusion of deep Convolutional Neural Networks (CNNs). We compare the proposed approach with non-image and single-scale CNN approaches and demonstrate the effectiveness in a case study using statistical data collected in the city of Amsterdam, Netherlands. We show that the proposed multi-scale approach outperforms the baseline methods.
Patrick Helber, Benjamin Bischke, Qiushi Guo, Jörn Hees, Andreas Dengel 0001
IGARSS5
2019 Towards a Sentinel-2 Based Human Settlement Layer
abstract
In this paper, we present how multi-spectral Sentinel-2 satellite images can be used in a machine learning approach based on an encoder-decoder semantic segmentation network to map human settlements. We show the effectiveness of the proposed CNN approach for the mapping of settlements in experiments with 785 European cities. The proposed approach to learn a settlement mapping with noisy ground truth data results in an effective settlement segmentation network with a mean intersection over union of 80.55% and a pixel accuracy of 87.40%.
Patrick Helber, Benjamin Bischke, Jörn Hees, Andreas Dengel 0001
IGARSS4
2019 Fusion Strategies for Learning User Embeddings with Neural Networks
abstract
Growing amounts of online user data motivate the need for automated processing techniques. In case of user ratings, one interesting option is to use neural networks for learning to predict ratings given an item and a user. While training for prediction, such an approach at the same time learns to map each user to a vector, a so-called user embedding. Such embeddings can for example be valuable for estimating user similarity. However, there are various ways how item and user information can be combined in neural networks, and it is unclear how the way of combining affects the resulting embeddings.In this paper, we run an experiment on movie ratings data, where we analyze the effect on embedding quality caused by several fusion strategies in neural networks. For evaluating embedding quality, we propose a novel measure, Pair-Distance Correlation, which quantifies the condition that similar users should have similar embedding vectors. We find that the fusion strategy affects results in terms of both prediction performance and embedding quality. Surprisingly, we find that prediction performance not necessarily reflects embedding quality. This suggests that if embeddings are of interest, the common tendency to select models based on their prediction ability should be reconsidered.
Philipp Blandfort, Tushar Karayil, Federico Raue, Jörn Hees, Andreas Dengel 0001
IJCNN5
2019 Inflection-Tolerant Ontology-Based Named Entity Recognition for Real-Time Applications
abstract
A growing number of applications users daily interact with have to operate in (near) real-time: chatbots, digital companions, knowledge work support systems - just to name a few. To perform the services desired by the user, these systems have to analyze user activity logs or explicit user input extremely fast. In particular, text content (e.g. in form of text snippets) needs to be processed in an information extraction task. Regarding the aforementioned temporal requirements, this has to be accomplished in just a few milliseconds, which limits the number of methods that can be applied. Practically, only very fast methods remain, which on the other hand deliver worse results than slower but more sophisticated Natural Language Processing (NLP) pipelines. In this paper, we investigate and propose methods for real-time capable Named Entity Recognition (NER). As a first improvement step, we address word variations induced by inflection, for example present in the German language. Our approach is ontology-based and makes use of several language information sources like Wiktionary. We evaluated it using the German Wikipedia (about 9.4B characters), for which the whole NER process took considerably less than an hour. Since precision and recall are higher than with comparably fast methods, we conclude that the quality gap between high speed methods and sophisticated NLP pipelines can be narrowed a bit more without losing real-time capable runtime performance.
Christian Jilek, Markus Schröder 0001, Rudolf Novik, Sven Schwarz, Heiko Maus, Andreas Dengel 0001
LDK6
2019 The Focus-Aspect-Value Model for Explainable Prediction of Subjective Visual Interpretation
abstract
Subjective visual interpretation is a challenging yet important topic in computer vision. Many approaches reduce this problem to the prediction of adjective- or attribute-labels from images. However,most of these do not take attribute semantics into account, or only process the image in a holistic manner. Furthermore, there is alack of relevant datasets with fine-grained subjective labels. In this paper, we propose the Focus-Aspect-Value (FAV) model to structure the process of capturing subjectivity in image processing,and introduce a novel dataset following this way of modeling. We run experiments on this dataset to compare several deep learning methods and find that incorporating context information based on tensor multiplication outperforms the default way of information fusion (concatenation).
Tushar Karayil, Philipp Blandfort, Jörn Hees, Andreas Dengel 0001
ICMR4
2019 Editorial for special issue on "Advanced Topics in Document Analysis and Recognition"
Cheng-Lin Liu 0001, Andreas Dengel 0001, Rafael Dueire Lins
Int. J. Document Anal. Recognit.2
2018 Overcoming Missing and Incomplete Modalities with Generative Adversarial Networks for Building Footprint Segmentation
abstract
The integration of information acquired with different modalities, spatial resolution and spectral bands has shown to improve predictive accuracies. Data fusion is therefore one of the key challenges in remote sensing. Most prior work focusing on multi-modal fusion, assumes that modalities are always available during inference. This assumption limits the applications of multi-modal models since in practice the data collection process is likely to generate data with missing, incomplete or corrupted modalities. In this paper, we show that Generative Adversarial Networks can be effectively used to overcome the problems that arise when modalities are missing or incomplete. Focusing on semantic segmentation of building footprints with missing modalities, our approach achieves an improvement of about 2% on the Intersection over Union (IoU) against the same network that relies only on the available modality.
Benjamin Bischke, Patrick Helber, Florian König, Damian Borth, Andreas Dengel 0001
CBMI5
2018 What Do Deep Networks Like to See?
abstract
We propose a novel way to measure and understand convolutional neural networks by quantifying the amount of input signal they let in. To do this, an autoencoder (AE) was fine-tuned on gradients from a pre-trained classifier with fixed parameters. We compared the reconstructed samples from AEs that were fine-tuned on a set of image classifiers (AlexNet, VGG16, ResNet-50, and Inception v3) and found substantial differences. The AE learns which aspects of the input space to preserve and which ones to ignore, based on the information encoded in the backpropagated gradients. Measuring the changes in accuracy when the signal of one classifier is used by a second one, a relation of total order emerges. This order depends directly on each classifier's input signal but it does not correlate with classification accuracy or network size. Further evidence of this phenomenon is provided by measuring the normalized mutual information between original images and auto-encoded reconstructions from different fine-tuned AEs. These findings break new ground in the area of neural network understanding, opening a new way to reason, debug, and interpret their results. We present four concrete examples in the literature where observations can now be explained in terms of the input signal that a model uses.
Sebastian Palacio, Joachim Folz, Jörn Hees, Federico Raue, Damian Borth, Andreas Dengel 0001
CVPR6
2018 anyAlign: An Intelligent and Interactive Text-Alignment Web-Application for Historical Document
abstract
Text alignment is an important step for analyzing historical archives. For analysis, it is important to align the text with their corresponding document images. It is a time and labor intensive work for many paleographers. In this paper, we have presented an end-to-end semi-automatic and interactive text alignment system for historical document. The presented system consists of five main sequential steps: binarization, automatic text-line extraction, interactive error correction in extracted text-line, automatic text alignment, and interactive error correction in aligned text. The anyOCR system [1] is used for the first three steps. Afterwards, text alignment is done automatically by the system using Oriented Fast and Rotated Brief (ORB) local image feature descriptors. The ORB features are matched by k-Nearest-Neighbor (KNN). Finally, the system provides an interactive user interface for rectifying wrong text alignment. The results are discussed in the evaluation section.
Syed Saqib Bukhari, Manabendra Saha, Praveen Badimala, Manesh Kumar Lohano, Andreas Dengel 0001
DAS5
2018 Layout Error Correction Using Deep Neural Networks
abstract
Layout analysis, mainly including binarization and text-line extraction, is one of the most important performance determining steps of an OCR system for complex medieval historical document images, which contain noise, distortions and irregular layouts. In this paper, we present a novel text-line error correction technique which include a VGG Net to classify non-text-line and adversarial network approach to obtain the layout bounding mask. The presented text-line error correction technique are applied to a collection of 15th century Latin documents, which achieved more than 75% accuracy for segmentation techniques.
Srie Raam Mohan, Syed Saqib Bukhari, Andreas Dengel 0001
DAS3
2018 OCR Error Correction: State-of-the-Art vs an NMT-based Approach
abstract
Although the performance of the state-of-the-art OCR systems is very high, they can still introduce errors due to various reasons, and when it comes to historical documents with old manuscripts the performance of such systems gets even worse. That is why Post-OCR error correction has been an open problem for many years. Many state-of-the-art approaches have been introduced through the recent years. This paper contributes to the field of Post-OCR Error Correction by introducing two novel deep learning approaches to improve the accuracy of OCR systems, and a post processing technique that can further enhance the quality of the output results. These approaches are based on Neural Machine Translation (NMT) and were motivated by the great success that deep learning introduced to the field of Natural Language Processing. Finally, we will compare the state-of-the-art approaches in Post-OCR Error Correction with the newly introduced systems and discuss the results.
Kareem Mokhtar, Syed Saqib Bukhari, Andreas Dengel 0001
DAS3
2018 Comparative Study between Traditional Machine Learning and Deep Learning Approaches for Text Classification
abstract
In this contemporaneous world, it is an obligation for any organization working with documents to end up with the insipid task of classifying truckload of documents, which is the nascent stage of venturing into the realm of information retrieval and data mining. But classification of such humongous documents into multiple classes, calls for a lot of time and labor. Hence a system which could classify these documents with acceptable accuracy would be of an unfathomable help in document engineering. We have created multiple classifiers for document classification and compared their accuracy on raw and processed data. We have garnered data used in a corporate organization as well as publicly available data for comparison. Data is processed by removing the stop-words and stemming is implemented to produce root words. Multiple traditional machine learning techniques like Naive Bayes, Logistic Regression, Support Vector Machine, Random forest Classifier and Multi-Layer Perceptron are used for classification of documents. Classifiers are applied on raw and processed data separately and their accuracy is noted. Along with this, Deep learning technique such as Convolution Neural Network is also used to classify the data and its accuracy is compared with that of traditional machine learning techniques. We are also exploring hierarchical classifiers for classification of classes and subclasses. The system classifies the data faster and with better accuracy than if done manually. The results are discussed in the results and evaluation section.
Cannannore Nidhi Kamath, Syed Saqib Bukhari, Andreas Dengel 0001
DocEng3
2018 iDocChip: A Configurable Hardware Architecture for Historical Document Image Processing: Percentile Based Binarization
abstract
End-to-end Optical Character Recognition (OCR) systems are heavily used to convert document images into machine-readable text. Commercial and open-source OCR systems (like Abbyy, OCRopus, Tesseract etc.) have traditionally been optimized for contemporary documents like books, letters, memos, and other end-user documents. However, these systems are difficult to use equally well for digitizing historical document images, which contain degradations like non-uniform shading, bleed-through, and irregular layout; such degradations usually do not exist in contemporary document images.
Vladimir Rybalkin, Syed Saqib Bukhari, Muhammad Mohsin Ghaffar, Aqib Ghafoor, Norbert Wehn, Andreas Dengel 0001
DocEng6
2018 Evaluating similarity measures for gaze patterns in the context of representational competence in physics education
abstract
The competent handling of representations is required for understanding physics' concepts, developing problem-solving skills, and achieving scientific expertise. Using eye-tracking methodology, we present the contributions of this paper as follows: We first investigated the preferences of students with the different levels of knowledge; experts, intermediates, and novices, in representational competence in the domain of physics problem-solving. It reveals that experts more likely prefer to use vector than other representations. Besides, a similar tendency of table representation usage was observed in all groups. Also, diagram representation has been used less than others. Secondly, we evaluated three similarity measures; Levenshtein distance, transition entropy, and Jensen-Shannon divergence. Conducting Recursive Feature Elimination technique suggests Jensen-Shannon divergence is the best discriminating feature among the three. However, investigation on mutual dependency of the features implies transition entropy mutually links between two other features where it has mutual information with Levenshtein distance (Maximal Information Coefficient = 0.44) and has a correlation with Jensen-Shannon divergence (r(18313) = 0.70, p < .001).
Seyyed Saleh Mozafari Chanijani, Pascal Klein 0002, Jouni Viiri, Sheraz Ahmed, Jochen Kuhn, Andreas Dengel 0001
ETRA6
2018 Hierarchical Model for Zero-shot Activity Recognition using Wearable Sensors
Mohammad Al-Naser, Hiroki Ohashi, Sheraz Ahmed, Katsuyuki Nakamura, Takayuki Akiyama, Takuto Sato, Phong Xuan Nguyen, Andreas Dengel 0001
ICAART (2)8
2018 Towards a Digital Personal Trainer for Health Clubs - Sport Exercise Recognition Using Personalized Models and Deep Learning
Sebastian Baumbach, Arun Bhatt, Sheraz Ahmed, Andreas Dengel 0001
ICAART (2)4
2018 SentiCite - An Approach for Publication Sentiment Analysis
abstract
With the rapid growth in the number of scientific publications, year after year, it is becoming increasingly difficult to identify quality authoritative work on a single topic. Though there is an availability of scientometric measures which promise to offer a solution to this problem, these measures are mostly quantitative and rely, for instance, only on the number of times an article is cited. With this approach, it becomes irrelevant if an article is cited 10 times in a positive, negative or neutral way. In this context, it is quite important to study the qualitative aspect of a citation to understand its significance. This paper presents a novel system for sentiment analysis of citations in scientific documents (SentiCite) and is also capable of detecting nature of citations by targeting the motivation behind a citation, e.g., reference to a dataset, reading reference. Furthermore, the paper also presents two datasets (SentiCiteDB and IntentCiteDB) containing about 2,600 citations with their ground truth for sentiment and nature of citation. SentiCite along with other state-of-the-art methods for sentiment analysis are evaluated on the presented datasets. Evaluation results reveal that SentiCite outperforms state-of-the-art methods for sentiment analysis in scientific publications by achieving a F1-measure of 0.71.
Dominique Mercier, Akansha Bhardwaj, Andreas Dengel 0001, Sheraz Ahmed
ICAART (2)3
2018 Ontology-based Information Extraction from Technical Documents
Syed Tahseen Raza Rizvi, Dominique Mercier, Stefan Agne, Steffen Erkel, Andreas Dengel 0001, Sheraz Ahmed
ICAART (2)5
2018 Saliency based Adjective Noun Pair Detection System
Marco Stricker, Syed Saqib Bukhari, Damian Borth, Andreas Dengel 0001
ICAART (2)4
2018 Answering with Cases: A CBR Approach to Deep Learning
Kareem Amin 0001, Stelios Kapetanakis, Klaus-Dieter Althoff, Andreas Dengel 0001, Miltos Petridis
ICCBR4
2018 An Investigative Analysis of Different LSTM Libraries for Supervised and Unsupervised Architectures of OCR Training
abstract
Optical Character Recognition (OCR) involves conversion of images of text into machine encoded editable text. Despite the wide research advancements in the field of OCR systems, the recognition capability of OCR systems on unseen or degraded historical documents is still questionable. The degradations in the document like torn pages, ink spread and blurred documents are major challenges especially in the old paper documents. Most of such degraded documents lack a generalized and reliable OCR system mainly because of the unavailability of ground-truth data and poor generalization capabilities of the OCR systems. Also manually transcribing the documents is cumbersome task which also require certain language-specific expertise. This paper presents a feasibility study of different OCR architectures together with different preprocessing stages for a reliable OCR on such challenging documents. To this end, we evaluate various OCR settings on a dataset containing highly degraded historical German typewriter documents. This paper investigates various key aspects of OCR training such as the impact of incorporation of different LSTM libraries, grayscale or binarized data for training and training data size used on the subject dataset. In addition, difference in the effect of using completely manually transcribed data as compared to semi-corrected ground-truth data for anyOCR architecture of unsupervised OCR training have been analyzed on a small dataset. The anyOCR framework has shown promising results as an efficient OCR system which was evident with its comparison with other OCR systems. The various factors analyzed provided a feasible strategy for approaching the problem and evaluating highly challenging historical documents.
Syed Saqib Bukhari, Sumam Francis, Cannannore Nidhi Kamath, Andreas Dengel 0001
ICFHR4
2018 Transcription Free LSTM OCR Model Evaluation
abstract
In recent years there has been significant progress in the field of Optical Character Recognition (OCR), mainly due to the use of various LSTM-based architectures. In the classic supervised training setup for LSTM-based OCR, the available image data and corresponding transcription is split into a training, a validation and a test set. Especially in the context of historical documents generating these transcriptions can be very costly, therefore minimizing the required transcribed data or maximizing the size of the training set to generate better models are desirable. We propose a novel method to evaluate LSTM OCR-models without requiring transcription ground truth data. For this we employ a second LSTM in an encoder-decoder setup to recreate the image data from the OCR output and evaluate the model based on its difference to the original input. We show that this approach performs similar to traditional transcription based evaluation on a historical document from the 16th century.
Martin Jenckel, Syed Saqib Bukhari, Andreas Dengel 0001
ICFHR3
2018 Interactive LSTM-Based Design Support in a Sketching Tool for the Architectural Domain - Floor Plan Generation and Auto Completion based on Recurrent Neural Networks
Johannes Bayer, Syed Saqib Bukhari, Andreas Dengel 0001
ICPRAM3
2018 High Performance Layout Analysis of Medieval European Document Images
Syed Saqib Bukhari, Anil Kumar Tiwari, Andreas Dengel 0001
ICPRAM4
2018 Impact of Training LSTM-RNN with Fuzzy Ground Truth
Martin Jenckel, Sourabh Sarvotham Parkala, Syed Saqib Bukhari, Andreas Dengel 0001
ICPRAM4
2018 Segmentation of Imbalanced Classes in Satellite Imagery using Adaptive Uncertainty Weighted Class Loss
abstract
We propose a novel loss function for the training of deep Convolutional Neural Networks (CNNs) focusing on land use and land cover classification in remote sensed data. In satellite imagery, object classes are often highly imbalanced leading to poor pixel-wise classification results when using standard training methods only. In this work, we introduce a loss function which leverages the per class uncertainty of the model during training together with median frequency balancing of the class pixels. We evaluate our result on aerial images of the state-of-the-art dataset Vaihingen. We obtain a significant improvement of the F1-Score and pixel accuracy against the standard cross entropy loss on the small car class. The overall Fl-Score using a single CNN achieves 89.35% resulting in an error reduction of 21.22% against the baseline.
Benjamin Bischke, Patrick Helber, Damian Borth, Andreas Dengel 0001
IGARSS4
2018 Introducing Eurosat: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification
abstract
In this paper, we address the challenge of land use and land cover classification using Sentinel-2 satellite images. The key contributions are as follows. We present a novel dataset based on Sentinel-2 satellite images covering 13 different spectral bands and consisting of 10 classes with in total 27,000 labeled images. We evaluate state-of-the-art deep Convolutional Neural Networks (CNNs) on this novel dataset with its different spectral bands. We also evaluate deep CNNs on existing remote sensing datasets and compare the obtained results. With the proposed novel dataset, we achieved an overall classification accuracy of 98.57%. The classification system resulting from the proposed research opens a gate towards various Earth observation applications. We demonstrate how the classification system can assist in improving geographical maps.
Patrick Helber, Benjamin Bischke, Andreas Dengel 0001, Damian Borth
IGARSS3
2018 Symbol Grounding Association in Multimodal Sequences with Missing Elements
abstract
In this paper, we extend a symbolic association framework for being able to handle missing elements in multimodal sequences. The general scope of the work is the symbolic associations of object-word mappings as it happens in language development in infants. In other words, two different representations of the same abstract concepts can associate in both directions. This scenario has been long interested in Artificial Intelligence, Psychology, and Neuroscience. In this work, we extend a recent approach for multimodal sequences (visual and audio) to also cope with missing elements in one or both modalities. Our method uses two parallel Long Short-Term Memories (LSTMs) with a learning rule based on EM-algorithm. It aligns both LSTM outputs via Dynamic Time Warping (DTW). We propose to include an extra step for the combination with the max operation for exploiting the common elements between both sequences. The motivation behind is that the combination acts as a condition selector for choosing the best representation from both LSTMs. We evaluated the proposed extension in the following scenarios: missing elements in one modality (visual or audio) and missing elements in both modalities (visual and sound). The performance of our extension reaches better results than the original model and similar results to individual LSTM trained in each modality.
Federico Raue, Andreas Dengel 0001, Thomas M. Breuel, Marcus Liwicki
J. Artif. Intell. Res.2
2017 Measuring the Performance of Push-ups - Qualitative Sport Activity Recognition
Sebastian Baumbach, Andreas Dengel 0001
ICAART (2)2
2017 Where is that Button Again?! - Towards a Universal GUI Search Engine
abstract
In feature-rich software a wide range of functionality is spread across various menus, dialog windows, toolbars etc. Remembering where to find each feature is usually very hard, especially if it is not regularly used. We therefore provide a GUI search engine which is universally applicable to a large number of applications. Besides giving an overview of related approaches, we describe three major problems we had to solve, which are analyzing the GUI, understanding the users’ query and executing a suitable solution to find a desired UI element. Based on a user study we evaluated our approach and showed that it is particularly useful if a not regularly used feature is searched for. We already identified much potential for further applications based on our approach.
Sven Hertling, Markus Schröder 0001, Christian Jilek, Andreas Dengel 0001
ICAART (2)4
2017 Which Saliency Detection Method is the Best to Estimate the Human Attention for Adjective Noun Concepts?
Marco Stricker, Syed Saqib Bukhari, Mohammad Al-Naser, Seyyed Saleh Mozafari Chanijani, Damian Borth, Andreas Dengel 0001
ICAART (2)6
2017 Classless Association Using Neural Networks
Federico Raue, Sebastian Palacio, Andreas Dengel 0001, Marcus Liwicki
ICANN (2)3
2017 Extending the Flexibility of Case-Based Design Support Tools: A Use Case in the Architectural Domain
abstract
This paper presents results of a user study into extending the functionality of an existing case-based search engine for similar architectural designs to a flexible process-oriented case-based support tool for the architectural conceptualization phase. Based on a research examining the target group’s (architects) thinking and working processes during the early conceptualization phase (especially during the search for similar architectural references), we identified common features for defining retrieval strategies for a more flexible case-based search for similar building designs within our system. Furthermore, we were also able to infer a definition for implementing these strategies into the early conceptualization process in architecture, that is, to outline a definition for this process as a wrapping structure for a user model. The study was conducted among the target group representatives (architects, architecture students and teaching personnel) by means of applying the paper prototyping method and Business Processing Model and Notation (BPMN). The results of this work are intended as a foundation for our upcoming research, but we also think it could be of wider interest for the case-based design research area.
Viktor Ayzenshtadt, Christoph Langenhan, Syed Saqib Bukhari, Klaus-Dieter Althoff, Frank Petzold, Andreas Dengel 0001
ICCBR6
2017 Academic Community Explorer (ACE) for Syntactic, Semantic and Pragmatic Document Analysis
abstract
This paper presents a novel Academic Community Explorer (ACE) which performs syntactic, semantic and pragmatic document analysis of scientific publications. Firstly, ACE uses syntactic structure to extract relevant information from a scientific document. Secondly, semantic analysis is performed to derive an article based co-authorship and citation network. Finally, ACE uses these document based networks to build a complete community network for pragmatic analysis. Furthermore, scientometric analysis is performed to extract the pragmatics by analyzing authors and publication community networks through micro and macro indicators. Two novel micro indicators Senti-Index, reflecting the sentiment present in citations and, Overlap index, reflecting community behavior have been introduced. This is a step in the direction of automatic qualitative assessment of scientific documents. In addition, ACE provides a rich visualization interface which helps in exploratory analysis of the community to identify hidden patterns, e.g, isolated small groups in the community which collaborate and cite each other frequently. A feasibility study is performed on the corpus of ICDAR publications from 1993-2015 to show the insights and benefits of the ACE framework. The results reveals that ICDAR is a highly collaborative community which has most likely arrived at its 'phase transition' stage with 70% of the community closely connected to each other.
Akansha Bhardwaj, Dominique Mercier, Hisham Hashmi, Sheraz Ahmed, Andreas Dengel 0001
ICDAR5
2017 anyOCR: An Open-Source OCR System for Historical Archives
abstract
Currently an intensive amount of research is going on in the field of digitizing historical archives for converting scanned document images into searchable full text. This paper presents the "anyOCR" system which mainly emphasize the techniques requires for digitizing a historical archive with high accuracy. It is an open-source system for the research community who can easily apply the anyOCR system for digitizing historical archives. The anyOCR system supports a complete document processing pipeline, which includes layout analysis, training OCR models and text line prediction, with an addition of intelligent and interactive layout and OCR error corrections web applications. The anyOCR system can also be used for contemporary document images containing diverse, simple to complex, layouts. This paper describes the current state of the anyOCR system, its architecture, as well as its major features. This paper also provides information about the availability, documentation, and tutorials of the anyOCR system.
Syed Saqib Bukhari, Ahmad Kadi, Mohammad Ayman Jouneh, Fahim Mahmood Mir, Andreas Dengel 0001
ICDAR5
2017 AirScript - Creating Documents in Air
abstract
This paper presents a novel approach, called AirScript, for creating, recognizing and visualizing documents in air. We present a novel algorithm, called 2-DifViz, that converts the hand movements in air (captured by a Myo-armband worn by a user) into a sequence of x, y coordinates on a 2D Cartesian plane, and visualizes them on a canvas. Existing sensor-based approaches either do not provide visual feedback or represent the recognized characters using prefixed templates. In contrast, AirScript stands out by giving freedom of movement to the user, as well as by providing a real-time visual feedback of the written characters, making the interaction natural. AirScript provides a recognition module to predict the content of the document created in air. To do so, we present a novel approach based on deep learning, which uses the sensor data and the visualizations created by 2-DifViz. The recognition module consists of a Convolutional Neural Network (CNN). and two Gated Recurrent Unit (GRU) Networks. The output from these three networks is fused to get the final prediction about the characters written in air. AirScript can be used in highly sophisticated environments like a smart classroom, a smart factory or a smart laboratory, where it would enable people to annotate pieces of texts wherever they want without any reference surface. We have evaluated AirScript against various well-known learning models (HMM, KNN, SVM, etc.) on the data of 12 participants. Evaluation results show that the recognition module of AirScript largely outperforms all of these models by achieving an accuracy of 91.7% in a person independent evaluation and a 96.7% accuracy in a person dependent evaluation.
Ayushman Dash, Amit Sahu, Rajveer Shringi, John Cristian Borges Gamboa, Muhammad Zeshan Afzal, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed
ICDAR7
2017 Table Recognition in Heterogeneous Documents Using Machine Learning
abstract
Tables are an easy way to represent information in a structural form. Table recognition is important for the extraction of such information from document images. Usually, modern OCR systems provide textual information coming from tables without recognizing actual table structure. However, recognition of table structure is important to get the contextual meaning of the contents. Table structure recognition in heterogeneous documents is challenging due to a variety of table layouts. It becomes harder where no physical rulings are present in a table. This work proposes a novel learning based methodology for the recognition of table contents in heterogeneous document images. Textual contents of documents are classified as table or non-table elements using a pre-trained neural network model. The output of the neural network is further enhanced by applying a contextual post processing on each element to correct the classifications errors if any. The system is trained using a subset of UNLV and UW3 document images and depicted more than 97% accuracy on a test set in detection of table and non-table elements.
Sheikh Faisal Rashid, Abdullah Akmal, Ali Adnan Aslam, Andreas Dengel 0001
ICDAR5
2017 Classification and Information Extraction for Complex and Nested Tabular Structures in Images
abstract
Understanding of technical documents, like manuals, is one of the most important steps in automatic reporting and/or troubleshooting of defects. The majority of the relevant information exists in tabular structure. There are some solutions for extracting tabular structures from text. However, it is still a big issue to extract tabular information from images and, on top of that, from complex and nested tables. This paper aims to propose classification and information extraction methods for complex tabular structures in document images. These are hybrid approaches using both image layout and OCRed text. The proposed methods outperform on a real-world technical documents dataset from a German railway company (Deutsche Bahn AG) as compared to other state-of-the-art approaches. As a result, the proposed approaches won the competition held by Deutsche Bahn AG in 2016 against other participating research groups and companies.
Amir Riad, Christian Sporer, Syed Saqib Bukhari, Andreas Dengel 0001
ICDAR4
2017 DeepDeSRT: Deep Learning for Detection and Structure Recognition of Tables in Document Images
abstract
This paper presents a novel end-to-end system for table understanding in document images called DeepDeSRT. In particular, the contribution of DeepDeSRT is two-fold. First, it presents a deep learning-based solution for table detection in document images. Secondly, it proposes a novel deep learning-based approach for table structure recognition, i.e. identifying rows, columns, and cell positions in the detected tables. In contrast to existing rule-based methods, which rely on heuristics or additional PDF metadata (like, for example, print instructions, character bounding boxes, or line segments), the presented system is data-driven and does not need any heuristics or metadata to detect as well as to recognize tabular structures in document images. Furthermore, in contrast to most existing table detection and structure recognition methods, which are applicable only to PDFs, DeepDeSRT processes document images, which makes it equally suitable for born-digital PDFs (as they can automatically be converted into images) as well as even harder problems, e.g. scanned documents. To gauge the performance of DeepDeSRT, the system is evaluated on the publicly available ICDAR 2013 table competition dataset containing 67 documents with 238 pages overall. Evaluation results reveal that DeepDeSRT outperforms state-of-the-art methods for table detection and structure recognition and achieves F1-measures of 96.77% and 91.44% for table detection and structure recognition, respectively. Additionally, DeepDeSRT is evaluated on a closed dataset from a real use case of a major European aviation company comprising documents which are highly unlike those in ICDAR 2013. Tested on a randomly selected sample from this dataset, DeepDeSRT achieves high detection accuracy for tables which demonstrates the sound generalization capabilities of our system.
Sebastian Schreiber 0001, Stefan Agne, Ivo Wolf, Andreas Dengel 0001, Sheraz Ahmed
ICDAR4
2017 DeepBIBX: Deep Learning for Image Based Bibliographic Data Extraction
Akansha Bhardwaj, Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (2)3
2017 Semantic Pattern-based Retrieval of Architectural Floor Plans with Case-based and Graph-based Searching Techniques and their Evaluation and Visualization
Qamer Uddin Sabri, Johannes Bayer, Viktor Ayzenshtadt, Syed Saqib Bukhari, Klaus-Dieter Althoff, Andreas Dengel 0001
ICPRAM6
2017 Grid-based outlier detection in large data sets for combine harvesters
abstract
Outlier detection is one of the most widely used technique to identify abnormal behavior in raw data. The sense of abnormal deviation mentioned here accounts not only for human made or system errors that naturally occur as part of the data but also as seldomly occuring events. In this paper, we propose a new algorithm called Grid Based Outlier Detection (GBOD) to find the hidden outliers in large data sets. In contrast to existing grid based methods which are limited to only some statistical based approaches, the GBOD algorithm is raised with two alternations to figure out different range of outliers depending on the interest of the user. First, the number of points in a local grid cell is used to decide whether a point is an outlier or not. In a second step, this approach is extended to method that assigns an outlier score to each data point. The simple design makes this algorithm extremely efficient for large data sets.
Ram Kumar Ganesan, Benjamin Bischke, Ansgar Bernardi, Alexander Maier, Heinrich Warkentin, Thilo Steckel, Andreas Dengel 0001
INDIN8
2016 High Performance OCR for Camera-Captured Blurred Documents with LSTM Networks
abstract
Documents are routinely captured by digital cameras in today's age owing to the availability of high quality cameras in smart phones. However, recognition of camera-captured documents is substantially more challenging as compared to traditional flat bed scanned documents due to the distortions introduced by the cameras. One of the major performancelimiting artifacts is the motion and out-of-focus blur that is often induced in the document during the capturing process. Existing approaches try to detect presence of blur in the document to inform the user for re-capturing the image. This paper reports, for the first time, an Optical Character Recognition (OCR) system that can directly recognize blurred documents on which the stateof-the-art OCR systems are unable to provide usable results. Our presented system is based on the Long Short-Term Memory (LSTM) networks and has shown promising character recognition results on both the motion-blurred and out-of-focus blurred images. One important feature of this work is that the LSTM networks have been applied directly to the gray-scale document images to avoid error-prone binarization of blurred documents. Experiments are conducted on publicly available SmartDoc-QA dataset that contains a wide variety of image blur degradations. Our presented system achieves 12.3% character error rate on the test documents, which is an over three-fold reduction in the error rate (38.9%) of the best-performing contemporary OCR system (ABBYY Fine Reader) on the same data.
Fallak Asad, Adnan Ul-Hasan, Faisal Shafait, Andreas Dengel 0001
DAS4
2016 What You See is What You Get? Automatic Image Verification for Online News Content
abstract
Consuming news over online media has witnessed rapid growth in recent years, especially with the increasing popularity of social media. However, the ease and speed with which users can access and share information online facilitated the dissemination of false or unverified information. One way of assessing the credibility of online news stories is by examining the attached images. These images could be fake, manipulated or not belonging to the context of the accompanying news story. Previous attempts to news verification provided the user with a set of related images for manual inspection. In this work, we present a semi-automatic approach to assist news-consumers in instantaneously assessing the credibility of information in hypertext news articles by means of meta-data and feature analysis of images in the articles. In the first phase, we use a hybrid approach including image and text clustering techniques for checking the authenticity of an image. In the second phase, we use a hierarchical feature analysis technique for checking the alteration in an image, where different sets of features, such as edges and SURF, are used. In contrast to recently reported manual news verification, our presented work shows a quantitative measurement on a custom dataset. Results revealed an accuracy of 72.7% for checking the authenticity of attached images with a dataset of 55 articles. Finding alterations in images resulted in an accuracy of 88% for a dataset of 50 images.
Sarah Elkasrawi, Andreas Dengel 0001, Ahmed Abdelsamad, Syed Saqib Bukhari
DAS2
2016 OCRoRACT: A Sequence Learning OCR System Trained on Isolated Characters
abstract
Digitizing historical documents is crucial in preserving the literary heritage. With the availability of low cost capturing devices, libraries and institutes all over the world have old literature preserved in the form of scanned documents. However, searching through these scanned images is still a tedious job as one is unable to search through them. Contemporary machine learning approaches have been applied successfully to recognize text in both printed and handwriting form, however, these approaches require a lot of transcribed training data in order to obtain satisfactory performance. Transcribing the documents manually is a laborious and costly task, requiring many man-hours and language-specific expertise. This paper presents a generic iterative training framework to address this issue. The proposed framework is not only applicable to historical documents, but for present-day documents as well, where manually transcribed training data is unavailable. Starting with the minimal information available, the proposed approach iteratively corrects the training and generalization errors. Specifically, we have used a segmentation-based OCR method to train on individual symbols and then use the semi-corrected recognized text lines as the ground-truth data for segmentation-free sequence learning, which learns to correct the errors in the ground-truth by incorporating context-aware processing. The proposed approach is applied to a collection of 15th century Latin documents. The iterative procedure using segmentation-free OCR was able to reduce the initial character error of about 23% (obtained from segmentation-based OCR) to less than 7% in few iterations.
Adnan Ul-Hasan, Syed Saqib Bukhari, Andreas Dengel 0001
DAS3
2016 An Evolutionary Algorithm to Learn SPARQL Queries for Source-Target-Pairs - Finding Patterns for Human Associations in DBpedia
Jörn Hees, Rouven Bauer, Joachim Folz, Damian Borth, Andreas Dengel 0001
EKAW5
2016 Thinking With Containers: A Multi-Agent Retrieval Approach for the Case-Based Semantic Search of Architectural Designs
Viktor Ayzenshtadt, Christoph Langenhan, Syed Saqib Bukhari, Klaus-Dieter Althoff, Frank Petzold, Andreas Dengel 0001
ICAART (1)6
2016 Human Activity Recognition - Using Sensor Data of Smartphones and Smartwatches
abstract
Unobtrusive and mobile activity monitoring using ubiquitous, cheap and widely available technology is the key requirement for human activity recognition supporting novel applications, such as health monitoring. With the recent progress in wearable technology, pervasive sensing and computing has become feasible. However, recognizing complex activities on light-weight devices is a challenging task. In this work, a platform to combine off-the-shelf sensors of smartphones and smartwatches for recognizing human activities in real-time is proposed. In order to achieve the best tradeoff between the system’s computational complexity and recognition accuracy, several evaluations were carried out to determine which classification algorithm and features to be used. Therefore, a data set from 16 participants was collected that includes normal daily activities and several fitness exercises. The analysis results showed that naive Bayes performs best in our experiment in both the accuracy and efficiency of classification, while the overall classification accuracy is 87% ± 2.4.
Bishoy Sefen, Sebastian Baumbach, Andreas Dengel 0001, Slim Abdennadher
ICAART (2)3
2016 Symbolic Association Using Parallel Multilayer Perceptron
Federico Raue, Sebastian Palacio, Thomas M. Breuel, Wonmin Byeon, Andreas Dengel 0001, Marcus Liwicki
ICANN (2)5
2016 Comparative Evaluation of Rule-Based and Case-Based Retrieval Coordination for Search of Architectural Building Designs
Viktor Ayzenshtadt, Christoph Langenhan, Johannes Roith, Syed Saqib Bukhari, Klaus-Dieter Althoff, Frank Petzold, Andreas Dengel 0001
ICCBR7
2016 KPTI: Katib's Pashto Text Imagebase and Deep Learning Benchmark
abstract
This paper presents the first Pashto text image database for scientific research and thereby the first dataset with complete handwritten and printed text line images which ultimately covers all alphabets of Arabic and Persian languages. Language like Pashto, written in a complex way by calligraphers, still requires a mature Optical Character Recognition (OCR), system. Although 50 million people use this language both for oral and written communication, there is no significant effort which is devoted to the recognition of Pashto Script. A real dataset of 17,015 images having Pashto text lines is introduced. The images are acquired via scanning from hand scribed Pashto books. Further, in this work, we evaluated the performance of deep learning based models like Bidirectional and Multi-Dimensional Long Short Term Memory (BLSTM and MDLSTM) networks for Pashto texts and provide a baseline character error rate of 9.22%.
Riaz Ahmad 0001, Muhammad Zeshan Afzal, Sheikh Faisal Rashid, Marcus Liwicki, Thomas M. Breuel, Andreas Dengel 0001
ICFHR6
2016 A Tesseract-based OCR framework for historical documents lacking ground-truth text
abstract
Computationally transcribing historical document images to digital text often requires an initial, labor intensive recording of ground-truths by language experts to provide the OCR system with training text. This paper presents a framework for the automatic generation of training data, provided only with labeled character images and a digital font, thus removing the need for manually generated text. In contrast to sample images and their real text as ground-truth, our approach is based on the random, rule-based generation of “meaningless” text in an image file and a ground-truth text file. Furthermore, we experimentally demonstrate that there is a correlation between the similarity of the character sample images of a subset to each other and the resulting model's classification performance. This allows us to calculate upper- and lower-bound performance subsets for model generation using only the sample images themselves. We show that using more training samples does not unequivocally improve model performance, allowing us to focus on the case of one sample per character during training. Training a Tesseract model only with samples that maximize a dissimilarity metric for each character in regards to all other mean character sample images yields a character recognition error of ca. 15% on our custom benchmark of 15thcentury Latin documents, as compared to ca. 27% error rate for training a model in a traditional Tesseract style using a synthetically generated training image from (manually marked) real text.
Brennan Nunamaker, Syed Saqib Bukhari, Damian Borth, Andreas Dengel 0001
ICIP4
2016 anyOCR: A sequence learning based OCR system for unlabeled historical documents
abstract
Institutes and libraries around the globe are preserving the literary heritage by digitizing historical documents. However, to make this data easily accessible the scanned documents need to be transformed into search-able text. State of the art OCR systems using Long-Short-Term-Memory networks (LSTM) have been applied successfully to recognize text in both printed and handwritten form. Besides the general challenges with historical documents, e.g. poor image quality, damaged characters, etc., especially unknown scripts and old fonds make it difficult to provide the large amount of transcribed training data required for these methods to perform well. Transcribing the documents manually is very costly in terms of man-hours and require language specific expertise. The unknown fonds and requirement for meaningful context also make the use of synthetic data unfeasible. We therefore propose an end-to-end framework anyOCR that cuts the required input from language experts to a minimum and is therefore easily extendable to other documents. Our approach combines the strengths of segmentation-based OCR methods utilizing clustering on individual characters and segmentation-free OCR methods utilizing a LSTM architecture. The proposed approach is applied to a collection of 15thcentury Latin documents. Combining the initial clustering with segmentation-free OCR was able to reduce the initial error of about 16% to less than 8%.
Martin Jenckel, Syed Saqib Bukhari, Andreas Dengel 0001
ICPR3
2016 Introducing Concept And Syntax Transition Networks for Image Captioning
abstract
The area of image captioning i.e. the automatic generation of short textual descriptions of images has experienced much progress recently. However, image captioning approaches often only focus on describing the content of the image without any emotional or sentimental dimension which is common in human captions. This paper presents an approach for image captioning designed specifically to incorporate emotions and feelings into the caption generation process. The presented approach consists of a Deep Convolutional Neural Network (CNN) for detecting Adjective Noun Pairs in the image and a novel graphical network architecture called "Concept And Syntax Transition (CAST)" network for generating sentences from these detected concepts.
Philipp Blandfort, Tushar Karayil, Damian Borth, Andreas Dengel 0001
ICMR4
2016 Contextual Enrichment of Remote-Sensed Events with Social Media Streams
abstract
The availability of satellite images for academic or commercial purpose is increasing rapidly due to efforts made by governmental agencies (NASA, ESA) to publish such data openly or commercial startups (PlanetLabs) to provide real-time satellite data. Beyond many commercial application, satellite data is helpful to create situation awareness in disaster recovery and emergency situations such as wildfires, earthquakes, or flooding. To fully utilize such data sources, we present a scalable system for the contextual enrichment of satellite images by crawling and analyzing multimedia content from social media. This information stream can provide vital information from the ground and help to complement remote sensing in situations. We use Twitter as main data source and analyze its textual, visual, temporal, geographical and social dimensions. Visualizations show different aspects of the event allowing high-level comprehension and provide deeper insights into the event as complemented by social media.
Benjamin Bischke, Damian Borth, Christian Schulze 0001, Andreas Dengel 0001
ACM Multimedia4
2016 Generating Affective Captions using Concept And Syntax Transition Networks
abstract
The area of image captioning i.e. the automatic generation of short textual descriptions of images has experienced much progress recently. However, image captioning approaches often only focus on describing the content of the image without any emotional or sentimental dimension which is common in human captions. This paper presents an approach for image captioning designed specifically to incorporate emotions and feelings into the caption generation process. The presented approach consists of a Deep Convolutional Neural Network (CNN) for detecting Adjective Noun Pairs in the image and a graphical network architecture called "Concept And Syntax Transition (CAST)" network for generating sentences from these detected concepts.
Tushar Karayil, Philipp Blandfort, Damian Borth, Andreas Dengel 0001
ACM Multimedia4
2016 Seed, an End-User Text Composition Tool for the Semantic Web
Bahaa Eldesouky, Menna Bakry, Heiko Maus, Andreas Dengel 0001
ISWC (1)4
2016 Preference based Filtering and Recommendations for Running Routes
Hassan Issa 0001, Amir Guirguis, Shary Beshara, Stefan Agne, Andreas Dengel 0001
WEBIST (2)5
2015 Quantifying reading habits: counting how many words you read
abstract
Reading is a very common learning activity, a lot of people perform it everyday even while standing in the subway or waiting in the doctors office. However, we know little about our everyday reading habits, quantifying them enables us to get more insights about better language skills, more effective learning and ultimately critical thinking. This paper presents a first contribution towards establishing a reading log, tracking how much reading you are doing at what time. We present an approach capable of estimating the words read by a user, evaluate it in an user independent approach over 3 experiments with 24 users over 5 different devices (e-ink reader, smartphone, tablet, paper, computer screen). We achieve an error rate as low as 5% (using a medical electrooculography system) or 15% (based on eye movements captured by optical eye tracking) over a total of 30 hours of recording. Our method works for both an optical eye tracking and an Electrooculography system. We provide first indications that the method works also on soon commercially available smart glasses.
Kai Kunze, Katsutoshi Masai, Masahiko Inami, Ömer Sacakli, Marcus Liwicki, Andreas Dengel 0001, Shoya Ishimaru, Koichi Kise
UbiComp6
2015 Deepdocclassifier: Document classification with deep Convolutional Neural Network
abstract
This paper presents a deep Convolutional Neural Network (CNN) based approach for document image classification. One of the main requirement of deep CNN architecture is that they need huge number of samples for training. To overcome this problem we adopt a deep CNN which is trained using big image dataset containing millions of samples i.e., ImageNet. The proposed work outperforms both the traditional structure similarity methods and the CNN based approaches proposed earlier. The accuracy of the proposed approach with merely 20 images per class outperforms the state-of-the-art by achieving classification accuracy of 68.25%. The best results on Tobbacoo-3428 dataset show that our proposed method outperforms the state-of-the-art method by a significant margin and achieved a median accuracy of 77.6% with 100 samples per class used for training and validation.
Muhammad Zeshan Afzal, Samuele Capobianco, Muhammad Imran Malik, Simone Marinai, Thomas M. Breuel, Andreas Dengel 0001, Marcus Liwicki
ICDAR6
2015 Recognizable units in Pashto language for OCR
abstract
Atomic segmentation of cursive scripts into constituent characters is one of the most challenging problems in pattern recognition. To avoid segmentation in cursive script, concrete shapes are considered as recognizable units. Therefore, the objective of this work is to find out the alternate recognizable units in Pashto cursive script. These alternatives are ligatures and primary ligatures. However, we need sound statistical analysis to find the appropriate numbers of ligatures and primary ligatures in Pashto script. In this work, a corpus of 2, 313, 736 Pashto words are extracted from a large scale diversified web sources, and total of 19, 268 unique ligatures have been identified in Pashto cursive script. Analysis shows that only 7000 ligatures represent 91% portion of overall corpus of the Pashto unique words. Similarly, about 7, 681 primary ligatures are also identified which represent the basic shapes of all the ligatures.
Riaz Ahmad 0001, Muhammad Zeshan Afzal, Sheikh Faisal Rashid, Marcus Liwicki, Andreas Dengel 0001, Thomas M. Breuel
ICDAR5
2015 Visual appearance based document classification methods: Performance evaluation and benchmarking
abstract
Most of the traditional document image classification techniques concentrate on document segmentation and OCR analysis, in spite of so many complexities and limitations involved. Recently, many of the document image classification problems are easily solved just by adapting standard computer vision approaches for natural image retrieval and classification, that are referred as visual appearance based document classification techniques. These approaches have reported better results as compared to the traditional approaches on proprietary datasets. However, so far these approaches are not compared with each other and, despite having potential, they are not evaluated on distorted camera-captured documents, which is one of the challenging requirements in our present commercial document analysis projects. In this paper, we present simple and effective descriptions of different visual appearance based document image classification techniques. We compare their performance on various standard and publicly available datasets, that are differ in degree of image degradations and content variations. We also demonstrate their advantages and limitations. Additionally, we make the implemented versions of these method publicly available to research community for usage and further testing on other domains.
Syed Saqib Bukhari, Andreas Dengel 0001
ICDAR2
2015 Supporting early contextualization of textual content in digital documents on the Web
abstract
The World Wide Web is arguably the most important source of digital documents nowadays. These documents mainly consist of unstructured and semi-structured data comprising a wealth of information at the disposal of the DAR (Document Analysis and Recognition) community. Contextualization plays an important role in understanding the content of those documents. In this paper, we present an approach to early contextualization of textual data in HTML documents. It combines automatic as well as semiautomatic annotation of named entities with user interaction to support contextualization of the content of digital documents as early as in the authoring stage of their life cycle. We also present the results of an online experimental evaluation involving 120 human test subjects. They show that our approach successfully managed to produce semantically annotated versions of unstructured textual content, which contain reliable contextual information, thus facilitating the task of later document analysis stages.
Bahaa Eldesouky, Menna Bakry, Heiko Maus, Andreas Dengel 0001
ICDAR4
2015 Sentiment analysis of texts by capturing underlying sentiment patterns
abstract
With the rapid proliferation of online-blogging and micro-blogging websites, millions of text posts are generated and made available online every day. Utilizing this rich data channel could facilitate educated purchasing of items, discovering trends and public tendencies regarding various products available in the market, discovering political inclination of societies prior to a national election, etc. Since the last decade, Sentiment Analysis (SA) has received increased attention from many researchers as a method for addressing topics, such as the aforementioned ones. This paper focuses on SA using sentiment features and patterns. We propose different sentiment polarity detection methods, two unsupervised methods and one supervised, which we compare with two baseline methods, a state-of-the-art Support Vector Machine (SVM) classifier trained on a unigram bag-of-words model, and an unsupervised SentiStrength [38] algorithm. In our experiments, we show that our polarity detection methods are highly effective and can outperform the aforementioned baselines in most of our conducted experiments.
Seyed Ali Bahrainian, Andreas Dengel 0001
Web Intell.2
2014 An Aggressive Feature Selection Technique for Rule-based Text Categorization
Salma Tayel, Stefan Agne, Andreas Dengel 0001, Slim Abdennadher
ICAART (1)3
2014 Automatic Signature Stability Analysis and Verification Using Local Features
abstract
The purpose of writing this paper is two-fold. First, it presents a novel signature stability analysis based on signature's local / part-based features. The Speeded Up Local features (SURF) are used for local analysis which give various clues about the potential areas from whom the features should be exclusively considered while performing signature verification. Second, based on the results of the local stability analysis we present a novel signature verification system and evaluate this system on the publicly available dataset of forensic signature verification competition, 4NSigComp2010, which contains genuine, forged, and disguised signatures. The proposed system achieved an EER of 15%, which is considerably very low when compared against all the participants of the said competition. Furthermore, we also compare the proposed system with some of the earlier reported systems on the said data. The proposed system also outperforms these systems.
Muhammad Imran Malik, Marcus Liwicki, Andreas Dengel 0001, Seiichi Uchida, Volkmar Frinken
ICFHR3
2014 A mixed reality head-mounted text translation system using eye gaze input
abstract
Efficient text recognition has recently been a challenge for augmented reality systems. In this paper, we propose a system with the ability to provide translations to the user in real-time. We use eye gaze for more intuitive and efficient input for ubiquitous text reading and translation in head mounted displays (HMDs). The eyes can be used to indicate regions of interest in text documents and activate optical-character-recognition (OCR) and translation functions. Visual feedback and navigation help in the interaction process, and text snippets with translations from Japanese to English text snippets, are presented in a see-through HMD. We focus on travelers who go to Japan and need to read signs and propose two different gaze gestures for activating the OCR text reading and translation function. We evaluate which type of gesture suits our OCR scenario best. We also show that our gaze-based OCR method on the extracted gaze regions provide faster access times to information than traditional OCR approaches. Other benefits include that visual feedback of the extracted text region can be given in real-time, the Japanese to English translation can be presented in real-time, and the augmentation of the synchronized and calibrated HMD in this mixed reality application are presented at exact locations in the augmented user view to allow for dynamic text translation management in head-up display systems.
Takumi Toyama, Daniel Sonntag, Andreas Dengel 0001, Takahiro Matsuda 0001, Masakazu Iwamura, Koichi Kise
IUI3
2014 Automatic Detection of CSA Media by Multi-modal Feature Fusion for Law Enforcement Support
abstract
The growing amounts of multimedia data being made available and shared via the Internet pose an increasing problem for law enforcement to investigate the distribution and possession of child sexual abuse (CSA) media. In this paper we address the automatic detection of CSA material in image and video data by multi-modal feature description. Instead of analyzing hash sums or file names, we propose the content-based analysis on visual and, in case of videos, also audio features. To this end, we apply multiple low level features as well as SentiBank, a novel mid-level representation of visual content. In collaboration with police partners and European cyber crime units, we conducted experiments on several datasets, including real world CSA media. Our quantitative evaluation reveals the challenging nature of child pornography detection, especially in the joint presence of non-illegal pornographic data, rendering skin detection, a popular feature for detecting pornography, less discriminative. Further, the utilization of SentiBank features shows high potential for detection and explainability of such content. Overall, multi-modal feature fusion can achieve an improved detection accuracy, reducing equal error rate from 17% to 10% for images and from 16% to 8% for videos as compared to best single feature performance for the challenging task of classifying CSA content from adult media.
Christian Schulze 0001, Dominik Henter, Damian Borth, Andreas Dengel 0001
ICMR4
2014 Intuitive justifications of medical semantic search results
Björn Forcher, Thomas Roth-Berghofer, Stefan Agne, Andreas Dengel 0001
Eng. Appl. Artif. Intell.4
2014 Automatic classifier selection for non-experts
Matthias Reif, Faisal Shafait, Markus Goldstein, Thomas M. Breuel, Andreas Dengel 0001
Pattern Anal. Appl.5
2014 Automatic analysis and sketch-based retrieval of architectural floor plans
Sheraz Ahmed, Marcus Liwicki, Christoph Langenhan, Andreas Dengel 0001, Frank Petzold
Pattern Recognit. Lett.5
2014 Bridging the gap between handwriting recognition and knowledge management
Marcus Liwicki, Sebastian Ebert, Andreas Dengel 0001
Pattern Recognit. Lett.3
2013 Collecting Links between Entities Ranked by Human Association Strengths
Jörn Hees, Mohamed Khamis, Ralf Biedert, Slim Abdennadher, Andreas Dengel 0001
ESWC5
2013 Enhancing Attentive Task Search with Information Gain Trees and Failure Detection Strategies
Kristin Stamm, Andreas Dengel 0001
ICAART (2)2
2013 Automatic Ground Truth Generation of Camera Captured Documents Using Document Image Retrieval
abstract
In this paper a novel method for automatic ground truth generation of camera captured document images is proposed. Currently, no dataset is available for camera captured documents. It is very difficult to build these datasets manually, as it is very laborious and costly. The proposed method is fully automatic, allowing building the very large scale (i.e., millions of images) labeled camera captured documents dataset, without any human intervention. Evaluation of samples generated by the proposed approach shows that 99.98% of the images are correctly labeled. Novelty of the proposed approach lies in the use of document image retrieval for automatic labeling, especially for camera captured documents, which contain different distortions specific to camera, e.g., blur, occlusion, perspective distortion, etc.
Sheraz Ahmed, Koichi Kise, Masakazu Iwamura, Marcus Liwicki, Andreas Dengel 0001
ICDAR5
2013 A Generic Method for Stamp Segmentation Using Part-Based Features
abstract
Traditionally, stamps are considered as a seal of authenticity for documents. For automatic processing and verification, segmentation of stamps from documents is pivotal. Existing methods for stamp extraction mostly employ color and/or shape based techniques, thereby limiting their applicability to only colored and specific shape stamps. In this paper, a novel, generic method based on part-based features is presented for segmentation of stamps from document images. The proposed method can segment black, colored, unseen, arbitrary shaped, textual, as well as graphical stamps. The proposed method is evaluated on a publicly available dataset for stamp detection and verification and achieved recall and precision of 73% and 83% respectively, for black stamps which were not addressed in the past.
Sheraz Ahmed, Faisal Shafait, Marcus Liwicki, Andreas Dengel 0001
ICDAR4
2013 Document Authentication Using Printing Technique Features and Unsupervised Anomaly Detection
abstract
Automatically identifying that a certain page in a set of documents is printed with a different printer than the rest of the documents can give an important clue for a possible forgery attempt. Different printers vary in their produced printing quality, which is especially noticeable at the edges of printed characters. In this paper, a system using the difference in edge roughness to distinguish laser printed ages from inkjet printed pages is presented. Several feature extraction methods have been developed and evaluated for that purpose. In contrast to previous work, this system uses unsupervised anomaly detection to detect documents printed by a different printing technique than the majority of the documents among a set. This approach has the advantage that no prior training using genuine documents has to be done. Furthermore, we created a dataset featuring 1200 document images from different domains (invoices, contracts, scientific papers) printed by 7 different inkjet and 13 laser printers. Results show that the presented feature extraction method achieves the best outlier rank score in comparison to state-of-the-art features.
Johann Gebhardt, Markus Goldstein, Faisal Shafait, Andreas Dengel 0001
ICDAR4
2013 FREAK for Real Time Forensic Signature Verification
abstract
This paper presents a novel signature verification system based on local features of signatures. The proposed system uses Fast Retina Key points (FREAK) which represent local features and are inspired by the human visual system, particularly the retina. To locate local points of interest in signatures, two local key point detectors, i.e., Features from Accelerated Segment Test (FAST) and Speeded-up Robust Features (SURF), have been used and their performance comparison in terms of Equal Error Rate (EER) and time is presented. The proposed system has been evaluated on publicly available dataset of forensic signature verification competition, 4NSigComp2010, which contains genuine, forged, and disguised signatures. The proposed system achieved an EER of 30%, which is considerably very low when compared against all the participants of the said competition. In addition to EER, the proposed system requires only 0.6 seconds on average to verify a 3000*1500 scanned signature. This shows that the proposed system has a potential and suitability for forensic signature verification as well as real time applications.
Muhammad Imran Malik, Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001
ICDAR4
2013 Part-Based Automatic System in Comparison to Human Experts for Forensic Signature Verification
abstract
The purpose of writing this paper is three-fold. First, it presents a novel local / part-based automatic system for forensic signature verification involving disguised signatures. Disguised signatures are written by authentic authors but with the intention of later denial. The proposed system reaches an equal error rate of 3.36% in classifying disguised and genuine signatures. Second, it compares the performance of the proposed system with various state-of-the-art signature verification systems on the same data, i.e., the publicly available dataset of 4NSigComp2010 signature verification competition. Third, it presents a performance comparison of the proposed system with human forensic handwriting examiners. It is important as it highlights the potential of the proposed system to assist humans in solving real world forensic signature verification cases.
Muhammad Imran Malik, Marcus Liwicki, Andreas Dengel 0001
ICDAR3
2013 Continuous Partial-Order Planning for Multichannel Document Analysis: A Process-Driven Approach
abstract
With the rise of email communication, enterprises strive to manage incoming documents from all input channels for achieving customer satisfaction. Their overall goal is to reduce request processing time and to increase processing quality. Previously, we proposed the approach of process-driven document analysis (DA) using the concepts of Attentive Tasks (ATs) and the Specialist Board (SB). The ATs formalize information expectations of the processes toward an incoming document, whereas the SB describes all available DA methods. In this paper, we propose to apply continuous partial order planning (CPOP) for guiding DA with the goal of optimal extraction accuracy and runtime. To our knowledge, this approach provides a novel method for integrating knowledge management with DA, in particular for processes. Since planning has not been applied to this field yet, we explore learning the suitability function (SF) and the adaptation of the DA plan. First evaluations indicate the applicability of the approach and preferences for calibration.
Kristin Stamm, Marcus Liwicki, Andreas Dengel 0001
ICDAR3
2013 Wearable Reading Assist System: Augmented Reality Document Combining Document Retrieval and Eye Tracking
abstract
We present a new system that assists people's reading activity by combining a wearable eye tracker, a see-through head mounted display, and an image based document retrieval engine. An image based document retrieval engine is used for identification of the reading document, whereas an eye tracker is used to detect which part of the document the reader is currently reading. The reader can refer to the glossary of the latest viewed key word by looking at the see-through head mounted display. This novel document reading assist application, which is the integration of a document retrieval system into an everyday reading scenario for the first time, enriches people's reading life. In this paper, we i) investigate the performance of the state-of-the-art image based document retrieval method using a wearable camera, ii) propose a method for identification of the word the reader is attendant, and iii) conduct pilot studies for evaluation of the system in this reading context. The results show the potential of a document retrieval system in combination with a gaze based user-oriented system.
Takumi Toyama, Andreas Dengel 0001, Wakana Suzuki, Koichi Kise
ICDAR2
2013 Predicting Classifier Combinations
Matthias Reif, Annika Leveringhaus, Faisal Shafait, Andreas Dengel 0001
ICPRAM4
2013 User attention oriented augmented reality on documents with document dependent dynamic overlay
abstract
When we read a document (any kind of, scientific papers, novels, etc.), we often encounter a situation that the information from the reading document is too less to comprehend what the author(s) would like to convey. In this paper, we demonstrate how the combination of a wearable eye tracker, a see-through head-mounted display (HMD) and an image based document retrieval engine enhances people's reading experiences. By using our proposed system, the reader can get supportive information in the see-through HMD when he wants. A wearable eye tracker and a document retrieval engine are used to detect which line in the document the reader is reading. We propose a method to detect the reader's attention on a word in a reading document, in order to present information at a preferable moment. Furthermore, we also propose a method to project a point of the document to a point of the HMD screen, by calculating the pose of the reading document in the camera image. This projection enables the system to overlay the information dynamically in an augmented view on the reading line. The results from the user study and the experiments show the potential of the proposed system in a practical use case.
Takumi Toyama, Wakana Suzuki, Andreas Dengel 0001, Koichi Kise
ISMAR3
2013 Analysis and forecasting of trending topics in online media streams
abstract
Among the vast information available on the web, social media streams capture what people currently pay attention to and how they feel about certain topics. Awareness of such trending topics plays a crucial role in multimedia systems such as trend aware recommendation and automatic vocabulary selection for video concept detection systems. Correctly utilizing trending topics requires a better understanding of their various characteristics in different social media streams. To this end, we present the first comprehensive study across three major online and social media streams, Twitter, Google, and Wikipedia, covering thousands of trending topics during an observation period of an entire year. Our results indicate that depending on one's requirements one does not necessarily have to turn to Twitter for information about current events and that some media streams strongly emphasize content of specific categories. As our second key contribution, we further present a novel approach for the challenging task of forecasting the life cycle of trending topics in the very moment they emerge. Our fully automated approach is based on a nearest neighbor forecasting technique exploiting our assumption that semantically similar topics exhibit similar behavior.
Tim Althoff, Damian Borth, Jörn Hees, Andreas Dengel 0001
ACM Multimedia4
2013 Graph-based retrieval of building information models for supporting the early design stages
Christoph Langenhan, Marcus Liwicki, Frank Petzold, Andreas Dengel 0001
Adv. Eng. Informatics5
2012 Extraction of Text Touching Graphics Using SURF
abstract
In this paper we propose a novel part-based method for the extraction of text touching graphic components. The Speeded Up Robust Features (SURF) are used to localize the text components and distinguish them from graphics. We introduce several post-processing steps to finally detect the text. We have tested our method on a publicly available data set of architectural floor plans and on real geographical maps. On floor plans we have located more than 95% of the text components which were not identified as text beforehand because they were touching graphic components.
Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001
Document Analysis Systems3
2012 Automatic Room Detection and Room Labeling from Architectural Floor Plans
abstract
This paper presents an automatic system for analyzing and labeling architectural floor plans. In order to detect the locations of the rooms, the proposed systems extracts both, structural and semantic information from given floor plans. Furthermore, OCR is applied on the text layer to retrieve the meaningful room labeling. Finally, a novel post-processing is proposed to split rooms into several sub-regions if several semantic rooms share the same physical room. Our fully automatic system is evaluated on a publicly available dataset of architectural floor plans. In our experiments, we could clearly outperform other state-of-the-art approaches for room detection.
Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001
Document Analysis Systems4
2012 Towards Understandable Explanations for Document Analysis Systems
abstract
smartFIX is a product portfolio for knowledge based extraction of data from any document format. The system automatically determines the document type and extracts all relevant data for the respective business process. Data that is unreliably recognized is forwarded to a verification workplace for manual checking. In general, users have no difficulties to interpret the document data and wonder why the system needs additional input. For that reason, we implemented an explanation component that is used to justify extraction results, thus, increasing confidence of users. The component is using a semantic log making it possible to provide understandable explanations. We illustrate the benefits of that kind of technology in contrast to the current smartFIX Log Viewer by means of a preliminary user experiment.
Björn Forcher, Stefan Agne, Andreas Dengel 0001, Michael Gillmann, Thomas Roth-Berghofer
Document Analysis Systems3
2012 Recognizing Words in Scenes with a Head-Mounted Eye-Tracker
abstract
Recognition of scene text using a hand-held camera is emerging as a hot topic of research. In this paper, we investigate the use of a head-mounted eye-tracker for scene text recognition. An eye-tracker detects the position of the user's gaze. Using gaze information of the user, we can provide the user with more information about his region/object of interest in a ubiquitous manner. Therefore, we can realize a service such as the user gazes at a certain word and soon obtain the related information of the word by combining a word recognition system with eye-tracking technology. Such a service is useful since the user has to do nothing but gazes at interested words. With a view to realize the service, we experimentally evaluate the effectiveness of using the eye-tracker for word recognition. The initial results show the recognition accuracy was around 70% in our word recognition experiment and the average computational time was less than one second per a query image.
Takuya Kobayashi, Takumi Toyama, Faisal Shafait, Masakazu Iwamura, Koichi Kise, Andreas Dengel 0001
Document Analysis Systems6
2012 Koios++: A Query-Answering System for Handwritten Input
abstract
In this paper we propose KOIOS++, which automatically processes natural language queries provided by handwritten input. The system integrates several recent achievements in the area of handwriting recognition, natural language processing, information retrieval, and human computer interaction. It uses a knowledge base described by the resource description framework (RDF). Our generic approach first generates a lexicon as background information for the handwritten text recognition. After recognizing a handwritten query, several output hypotheses are sent to a natural language processing system in order to generate a structured query (SPARQL query). Subsequently, the query is applied to the given knowledge base and a result graph visualizes the retrieved information. At all stages, the user can easily adjust the intermediate results if there is any undesired outcome. The system is implemented as a web-service and therefore works for handwritten input on digital paper as well as on input on Pen-enabled interactive surfaces. Furthermore, we build on the generic RDF-representation of semantic knowledge which is also used by the linked open data (LOD) initiative. As such, our system works well in various scenarios. We have implemented prototypes for querying company knowledge bases, the DBPedia1, the DBLP computer science bibliography2, and a knowledge base of the DAS 2012.
Marcus Liwicki, Björn Forcher, Philipp Jaeger, Andreas Dengel 0001
Document Analysis Systems4
2012 Seamless Integration of Handwriting Recognition into Pen-Enabled Displays for Fast User Interaction
abstract
This paper proposes a framework for the integration of handwriting recognition into natural user interfaces. As more and more pen-enabled touch displays are available, we make use of the distinction between touch actions and pen actions. Furthermore, we apply a recently introduced mode detection approach to distinguish between handwritten strokes and graphics drawn with the pen. These ideas are implemented in the Touch & Write SDK which can be used for various applications. In order to evaluate the effectiveness of our approach, we have conducted experiments for an annotation scenario. We asked several users to mark and label several objects in videos. We have measured the labeling time when using our novel user interaction system and compared it to the time needed when using common labeling tools. Furthermore, we compare our handwritten input paradigm to other existing systems. It turns out that the annotation is performed much faster when using our method and the user experience is also much better.
Marcus Liwicki, Tobias Zimmermann, Andreas Dengel 0001
Document Analysis Systems4
2012 A Signature Verification Framework for Digital Pen Applications
abstract
In this paper we present a framework for real-time online signature verification scenarios. The proposed framework is based on state-of-the-art feature extraction and Gaussian Mixture Model (GMM) classification. While our signature verification library is generally applicable to any input device using digital pens, we have implemented verification scenarios using the Anoto digital pen. As such our automated signature verification framework becomes an interesting commodity for industry, because the Anoto SDK is easy to apply and the GMM-based classification can be seamlessly integrated. The novelty of this work is the application of our framework that takes real-time online signature verification to every scenario where digital pens may potentially be used. In this paper we describe several scenarios where our framework has been applied, including signatures in financial contracts or ordering processes. We also propose a general approach to integrate the GMM-descriptions into electronic ID-cards in order to also store behavioral biometrics on these cards. In experiments we have measured the performance of the signature verification system when skilled forgeries were present. The interest shown by our partner financial institutions and the results of our initial evaluations indicate that our signature verification framework suits exactly the demands of our clients.
Muhammad Imran Malik, Sheraz Ahmed, Andreas Dengel 0001, Marcus Liwicki
Document Analysis Systems3
2012 How Salient is Scene Text?
abstract
Computational models of visual attention use image features to identify salient locations in an image that are likely to attract human attention. Attention models have been quite effectively used for various object detection tasks. However, their use for scene text detection is under-investigated. As a general observation, scene text often conveys important information and is usually prominent or salient in the scene itself. In this paper, we evaluate four state-of-the-art attention models for their response to scene text. Initial results indicate that saliency maps produced by these attention models can be used for aiding scene text detection algorithms by suppressing non-text regions.
Asif Shahab, Faisal Shafait, Andreas Dengel 0001, Seiichi Uchida
Document Analysis Systems3
2012 Attentive Tasks: Process-Driven Document Analysis for Multichannel Documents
abstract
The increasing amount of email data has led many companies to new challenges with their employees now having to deal with information overload while managing multiple communication channels at the same time, e.g., email, mail, and phone. Moreover, emails can contain attachments, i.e., files with additional information. Most existing approaches for reducing email processing time require significant domain specific customization efforts to achieve good performance and lack attachment handling. We aim at providing a more domain independent approach by integrating the process context and using the information expectations of a process to guide the document analysis (DA) schedule for emails and their attachments. We rely on the concepts of Attentive Tasks (ATs) and Specialist Board (SB). ATs are templates that describe all relevant and expected information about a process currently waiting for input. The SB provides a machine readable description of DA methods, so-called specialists, that extract all relevant information for further processes. We present our approach and demonstrate the benefits for a domain specific application, i.e., a financial institution.
Kristin Stamm, Andreas Dengel 0001
Document Analysis Systems2
2012 Reading and estimating gaze on smart phones
abstract
While lots of reading happens on mobile devices, little research has been performed on how the reading-interaction actually takes place. Therefore we describe our findings on a study conducted with 18 users which were asked to read a number of texts while their touch and gaze data was being recorded. We found three reader types and identified their preferred alignment of text on the screen. Based on our findings we are able to computationally estimate the reading area with an approximate .81 precision and .89 recall. Our computed reading speed estimate has an average 10.9% wpm error in contrast to the measured speed, and combining both techniques we can pinpoint the reading location at a given time with an overall word error of 9.26 words, or about three lines of text on our device.
Ralf Biedert, Andreas Dengel 0001, Georg Buscher, Arman Vartan
ETRA2
2012 Towards robust gaze-based objective quality measures for text
abstract
An increasing amount of text is being read digitally. In this paper we explore how eye tracking devices can be used to aggregate reading data of many readers in order to provide authors and editors with objective and implicitly gathered quality feedback. We present a robust way to jointly evaluate the gaze data of multiple readers, with respect to various reading-related features. We conducted an experiment in which a group of high school students composed essays subsequently read and rated by a group of seven other students. Analyzing the recorded data, we find that the amount of regression targets, the reading-to-skimming ratio, reading speed and reading count are the most discriminative features to distinguish very comprehensible from barely comprehensible text passages. By employing machine learning techniques, we are able to classify the comprehensibility of text automatically with an overall accuracy of 62%.
Ralf Biedert, Andreas Dengel 0001, Mostafa Elshamy, Georg Buscher
ETRA2
2012 Universal eye-tracking based text cursor warping
abstract
In this paper we present an approach to build an eye-tracking based text cursor placement system. When triggered, the system employs a computer vision based analysis of the screen's content around the current gaze position to find the most likely designated gaze target. Eventually it synthesizes a mouse event at that position, allowing for a rapid text cursor repositioning even in applications which do not support eye tracking explicitly. For our system we compared three different computer vision methods in a simulation run and evaluated the best candidate in two double blinded user studies. We used a total of 19 participants to assess the system's objective and perceived end user speed up. We can demonstrate that in terms of reposition time the OCR based method is superior to the other tested methods, it also beats common keyboard-mouse interaction for some users. We conclude that while the tool was almost universally preferred subjectively over keyboard-mouse interaction, the highest speed can be achieved by using the right amount of eye tracking.
Ralf Biedert, Andreas Dengel 0001, Christoph Käding
ETRA2
2012 A robust realtime reading-skimming classifier
abstract
Distinguishing whether eye tracking data reflects reading or skimming already proved to be of high analytical value. But with a potentially more widespread usage of eye tracking systems at home, in the office or on the road the amount of environmental and experimental control tends to decrease. This in turn leads to an increase in eye tracking noise and inaccuracies which are difficult to address with current reading detection algorithms. In this paper we propose a method for constructing and training a classifier that is able to robustly distinguish reading from skimming patterns. It operates in real time, considering a window of saccades and computing features such as the average forward speed and angularity. The algorithm inherently deals with distorted eye tracking data and provides a robust, linear classification into the two classes read and skimmed. It facilitates reaction times of 750ms on average, is adjustable in its horizontal sensitivity and provides confidence values for its classification results; it is also straightforward to implement. Trained on a set of six users and evaluated on an independent test set of six different users it achieved a 86% classification accuracy and it outperformed two other methods.
Ralf Biedert, Jörn Hees, Andreas Dengel 0001, Georg Buscher
ETRA3
2012 Gaze guided object recognition using a head-mounted eye tracker
abstract
Wearable eye trackers open up a large number of opportunities to cater for the information needs of users in today's dynamic society. Users no longer have to sit in front of a traditional desk-mounted eye tracker to benefit from the direct feedback given by the eye tracker about users' interest. Instead, eye tracking can be used as a ubiquitous interface in a real-world environment to provide users with supporting information that they need. This paper presents a novel application of intelligent interaction with the environment by combining eye tracking technology with real-time object recognition. In this context we present i) algorithms for guiding object recognition by using fixation points ii) algorithms for generating evidence of users' gaze on particular objects iii) building a next generation museum guide called Museum Guide 2.0 as a prototype application of gaze-based information provision in a real-world environment. We performed several experiments to evaluate our gaze-based object recognition methods. Furthermore, we conducted a user study in the context of Museum Guide 2.0 to evaluate the usability of the new gaze-based interface for information provision. These results show that an enormous amount of potential exists for using a wearable eye tracker as a human-environment interface.
Takumi Toyama, Thomas Kieninger, Faisal Shafait, Andreas Dengel 0001
ETRA4
2012 Signature Segmentation from Document Images
abstract
In this paper we propose a novel method for the extraction of signatures from document images. Instead of using a human defined set of features a part-based feature extraction method is used. In particular, we use the Speeded Up Robust Features (SURF) to distinguish the machine printed text from signatures. Using SURF features makes the approach generally more useful and reliable for different resolution documents. We have evaluated our system on the publicly available Tobacco-800 dataset in order to compare it to previous work. Finally, all signatures were found in the images and less than half of the found signatures are false positives. Therefore, our system can be applied for practical use.
Sheraz Ahmed, Muhammad Imran Malik, Marcus Liwicki, Andreas Dengel 0001
ICFHR4
2012 Local Feature Based Online Mode Detection with Recurrent Neural Networks
abstract
In this paper we propose a novel approach for online mode detection, where the task is to classify ink traces into several categories. In contrast to previous approaches working on global features, we introduce a system completely relying on local features. For classification, standard recurrent neural networks (RNNs) and the recently introduced long short-term memory (LSTM) networks are used. Experiments are performed on the publicly available IAMonDo-database which serves as a benchmark data set for several researches. In the experiments we investigate several RNN structures and classification sub-tasks of different complexities. The final recognition rate on the complete test set is 98.47% in average, which is significantly higher than the 97% achieved with an MCS in previous work. Further interesting results on different subsets are also reported in this paper.
Sebastian Otte, Dirk Krechel, Marcus Liwicki, Andreas Dengel 0001
ICFHR4
2012 Searching attentive tasks with document analysis evidences and Dempster-Shafer theory
Kristin Stamm, Andreas Dengel 0001
ICPR2
2012 Semantic E-Ink: Knowledge-Based Assistance for Making Mental Models Explicit
abstract
In this paper we describe a system which assists knowledge workers in making notes of their thoughts and transferring them to the computer. Our system processes handwritten notes written down with a digital pen. These notes are processed in order to recognize and understand their meaning. To realize such systems, two novel processing stages are proposed for the first time in literature. The first stage is the inclusion of knowledge bases into the Handwriting Recognition (HWR) process, where we make use of a person's mental model. The second stage is the transition from pure HWR to understanding of the handwritten notes, i.e. the system extracts knowledge in form of ontologies. For both novel approaches we performed a set of experiments on various data. With the proposed techniques, the recognition rate of the HWR system as well as the performance of the information extraction system are significantly increased.
Andreas Dengel 0001, Marcus Liwicki
Int. J. Pattern Recognit. Artif. Intell.1
2012 Meta-learning for evolutionary parameter optimization of classifiers
Matthias Reif, Faisal Shafait, Andreas Dengel 0001
Mach. Learn.3
2012 Faster subgraph isomorphism detection by well-founded total order indexing
Marcus Liwicki, Andreas Dengel 0001
Pattern Recognit. Lett.3
2012 Attentive documents: Eye tracking as implicit feedback for information retrieval and beyond
abstract
Reading is one of the most frequent activities of knowledge workers. Eye tracking can provide information on what document parts users read, and how they were read. This article aims at generating implicit relevance feedback from eye movements that can be used for information retrieval personalization and further applications. We report the findings from two studies which examine the relation between several eye movement measures and user-perceived relevance of read text passages. The results show that the measures are generally noisy, but after personalizing them we find clear relations between the measures and relevance. In addition, the second study demonstrates the effect of using reading behavior as implicit relevance feedback for personalizing search. The results indicate that gaze-based feedback is very useful and can greatly improve the quality of Web search. The article concludes with an outlook introducing attentive documents keeping track of how users consume them. Based on eye movement feedback, we describe a number of possible applications to make working with documents more effective.
Georg Buscher, Andreas Dengel 0001, Ralf Biedert, Ludger van Elst
ACM Trans. Interact. Intell. Syst.2
2011 Fast Subgraph Isomorphism Detection for Graph-Based Retrieval
Christoph Langenhan, Thomas Roth-Berghofer, Marcus Liwicki, Andreas Dengel 0001, Frank Petzold
ICCBR5
2011 Improved Automatic Analysis of Architectural Floor Plans
abstract
This paper proposes a novel complete system for automated floor plan analysis. Besides applying and improving state-of-the-art processing methods, we introduce novel preprocessing methods, e.g., the differentiation between thick, medium, and thin lines and the removal of components outside the convex hull of the outer walls. Especially the latter method increases the performance of the final system. In our experiments on a reference data set we compare our approach to other approaches available in the literature. We show that our system outperforms previous systems. The final room recognition accuracy is 79% that is 10% higher than the 69% achieved by a state-of-the-art approach from the literature.
Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001
ICDAR4
2011 Text/Graphics Segmentation in Architectural Floor Plans
abstract
In this paper, we propose an improved method for text/graphics segmentation. Text/graphics separation is a crucial preprocessing step in document analysis before further analysis and recognition can be applied. Our proposed system extends the method of Tombre et al. with a number of improvements to make it more suitable for architectural floor plans. A crucial novel preprocessing step is the detection and removal of walls before the actual segmentation. Furthermore, text components are then extracted by analyzing connected components and even considering text overlapping with graphics. Finally, a smearing approach is used to remove noise and extract the final text components. Evaluation results over the series of 90 floor plans which has also been used in reference work shows that our method has a recall of almost 99% and a precision greater then 97%.
Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001
ICDAR4
2011 Semantic Logging: Towards Explanation-Aware DAS
abstract
smartFIX is a product portfolio for knowledge-based extraction of data from any document format. smartFIX automatically determines the document type and extracts all relevant data for the respective business process. Data that is uncertainly recognized is forwarded to a verification workplace for manual checking. In general, users have no difficulties to interpret the document data and wonder why the system needs additional input. For that reason, we will integrate an explanation component that will be used to justify uncertain extraction results, thus, increasing confidence of users. The component will be based on semantic technologies in general and on a semantic log in particular. The log will contain all process relevant information enabling the explanation facility to generate customized and understandable explanations. In this paper, we will discuss the benefits of that kind of technology with reference to DAS.
Björn Forcher, Stefan Agne, Andreas Dengel 0001, Michael Gillmann, Thomas Roth-Berghofer
ICDAR3
2011 Bayesian Approach to Photo Time-Stamp Recognition
abstract
Time-stamps and URLs overlaid artificially on images add useful meta information which can be used for automatic indexing of images and videos. In this paper, we propose a method based on an attention-based model of visual saliency to extract overlaid text and time-stamps that are rendered on images. Our model of visual saliency is based on a Bayesian framework and works very well for the task of time-stamp detection and segmentation as is evident by overall object recall of 80% and precision of 70%. Our method produces a clean text segmented binarized image, which can be used for recognition directly by an OCR system. Furthermore, our technique is robust against variation of font styles and color of time-stamp and overlaid text.
Asif Shahab, Faisal Shafait, Andreas Dengel 0001
ICDAR3
2011 ICDAR 2011 Robust Reading Competition Challenge 2: Reading Text in Scene Images
abstract
Recognition of text in natural scene images is becoming a prominent research area due to the widespread availablity of imaging devices in low-cost consumer products like mobile phones. To evaluate the performance of recent algorithms in detecting and recognizing text from complex images, the ICDAR 2011 Robust Reading Competition was organized. Challenge 2 of the competition dealt specifically with detecting/recognizing text in natural scene images. This paper presents an overview of the approaches that the participants used, the evaluation measure, and the dataset used in the Challenge 2 of the contest. We also report the performance of all participating methods for text localization and word recognition tasks and compare their results using standard methods of area precision/recall and edit distance.
Asif Shahab, Faisal Shafait, Andreas Dengel 0001
ICDAR3
2011 MCS for Online Mode Detection: Evaluation on Pen-Enabled Multi-touch Interfaces
abstract
This paper proposes a new approach for drawing mode detection in online handwriting. The system classifies groups of ink traces into several categories. The main contributions of this work are as follows. First, we improve and optimize several state-of-the-art recognizers by adding new features and applying feature selections. Second, we use several classifiers for the recognition. Third, we perform multiple classifier combination strategies for combining the outputs. Finally, a large experimental evaluation on two data sets is performed: the publicly available Touch&Write database which has been acquired on a pen-enabled multi-touch surface, and the publicly available IAMonDo-database which serves as a benchmark. In our experiments on the IAM-OnDo-database we achieved a recognition rate of 97%, which is much higher than other results reported in the literature. On the more balanced multi-touch surface data set we achieved a recognition rate of close to 98%.
Marcus Liwicki, Yannik T. H. Schelske, Christopher Schölzel, Florian Strauß, Andreas Dengel 0001
ICDAR6
2011 Generic and Specific Object Recognition for Semantic Retrieval of Images
Martin Klinkigt, Koichi Kise, Andreas Dengel 0001
KES (1)3
2011 Semantic Retrieval of Images by Learning from Wikipedia
Martin Klinkigt, Koichi Kise, Heiko Maus, Andreas Dengel 0001
KES (4)4
2011 From Handwriting Recognition to Ontologie-Based Information Extraction of Handwritten Notes
Marcus Liwicki, Sebastian Ebert, Andreas Dengel 0001
KES (4)3
2011 An Intelligent Shopping List - Combining Digital Paper with Product Ontologies
Marcus Liwicki, Sandra Thieme, Gerrit Kahl, Andreas Dengel 0001
KES (4)4
2011 Topic-Based Recommendations for Enterprise 2.0 Resource Sharing Platforms
Rafael Schirru, Stephan Baumann 0001, Martin Memmel, Andreas Dengel 0001
KES (1)4
2010 A Social Network Analysis and Mining Methodology for the Monitoring of Specific Domains in the Blogosphere
abstract
Whenever the question arises how a product, a personality, a technology or some other specific entity is perceived by the public, the blogosphere is a very good source of information. This is what usually interests business users from marketing or PR. Modern search services offer a rich set of tools to monitor or track the blogosphere as a whole, but the analysis with respect to a certain domain is very limited. In this paper we lay some foundations to aggregate blog articles of a specific domain from multiple search services, to analyse the social authorities of articles and blogs, and to monitor the attention articles of the domain receive over time. These are the building blocks required for a monitoring application that presents users the currently most interesting articles. This methodology can be instantiated and combined with additional textual analysis methods to create highly automated business intelligence applications.
Darko Obradovic, Stephan Baumann 0001, Andreas Dengel 0001
ASONAM3
2010 Crumblr: Aggregation and Sharing of Spatial Content in Mobile Environments
abstract
In the growing mobile computing sector two trends are more and more wide-spread and gain further momentum: location-sensing technologies and mobile Internet access. In this paper we describe Crumblr, an application for semi-automatic capturing, aggregating and sharing spatial content in mobile environments. While existing Web 2.0 services focus mostly on points of interest or on routes, our application combines these two entities and additionally enriches them with contextual data. The result is the generation of an implicit network among people, linked via spatial information. This allows us to provide them with personalized recommendations of places, routes and related other people.
Dragan Sunjka, Darko Obradovic, Andreas Dengel 0001
ASONAM3
2010 Improving handwriting recognition by the use of semantic information
abstract
This paper proposes a first attempt to include real semantic information into the process of handwriting recognition. We take advantage of the fact that the main topic of handwritten notes is often known beforehand like in annotation or reviewing tasks. Using state-of-the-art technologies from the knowledge management research area it is possible to store a semantic representation of the user's knowledge in a Personal Information Model (PIMO). This PIMO stores the relations between semantic concepts and documents on the computer. In this paper we extract texts from related documents and concepts of the PIMO. The vocabulary of these texts is then used to aid the recognizer. In our multi-writer experiments, a significant improvement of the recognition accuracy by 8% on the text line level has been achieved.
Marcus Liwicki, Hassan Mohamed Abou Eisha, Andreas Dengel 0001
Document Analysis Systems3
2010 Touch & Write: a multi-touch table with pen-input
abstract
In this paper we present a novel rear-projection tabletop called Touch & Write. It combines the FTIR technology for touching with the Anoto-technology for handwriting. This allows an implicit switch between the modes object manipulation, and content editing. Our system incorporates real-time gesture and handwriting recognition. Drawn objects and written concepts can be converted to digital information immediately. We introduce a functional application, the LeCoOnt concept mapping software makes use of the full capability of the Touch & Write table. Touching actions are used for arranging the concepts like sheets on a normal table, and to recognizes guestures like zooming. Pen-actions are used for drawing, connecting concepts, and handwriting. The handwritten strokes are automatically recognized and converted into a machine-readable string. This system provides a reliable alternative to common approaches which try to reconstruct the information from photographs.
Marcus Liwicki, Oleg Rostanin, Saher Mohamed El-Neklawy, Andreas Dengel 0001
Document Analysis Systems4
2010 An open approach towards the benchmarking of table structure recognition systems
abstract
Table spotting and structural analysis are just a small fraction of tasks relevant when speaking of table analysis. Today, quite a large number of different approaches facing these tasks have been described in literature or are available as part of commercial OCR systems that claim to deal with tables on the scanned documents and to treat them accordingly.
Asif Shahab, Faisal Shafait, Thomas Kieninger, Andreas Dengel 0001
Document Analysis Systems4
2010 Epiphany: Adaptable RDFa Generation Linking the Web of Documents to the Web of Data
Benjamin Adrian, Jörn Hees, Ivan Herman, Michael Sintek, Andreas Dengel 0001
EKAW5
2010 Constructing Understandable Explanations for Semantic Search Results
Björn Forcher, Thomas Roth-Berghofer, Michael Sintek, Andreas Dengel 0001
EKAW4
2010 Representing the International Classification of Diseases Version 10 in OWL
Manuel Möller, Michael Sintek, Ralf Biedert, Patrick Ernst, Andreas Dengel 0001, Daniel Sonntag
KEOD5
2010 Case Acquisition from Text: Ontology-Based Information Extraction with SCOOBIE for myCBR
Thomas Roth-Berghofer, Benjamin Adrian, Andreas Dengel 0001
ICCBR3
2010 a.SCatch: Semantic Structure for Architectural Floor Plan Retrieval
Christoph Langenhan, Thomas Roth-Berghofer, Marcus Liwicki, Andreas Dengel 0001, Frank Petzold
ICCBR5
2010 Ontology-Based Information Extraction from Handwritten Documents
abstract
In this paper we introduce a new layer for the task of handwriting recognition. We add semantic information by means of ontologies. The task of our recognizer therefore is not only to recognize the ASCII transcription of the handwritten document, but also to identify the semantic concepts which appear in the text. This task is called ontology-based information extraction (OBIE), which has been applied to electronic documents recently. OBIE methods first segment the text into tokens, then identify their values and their corresponding instances of the ontology, and finally try to generate new facts based on the text. To the authors' knowledge, in this paper OBIE is proposed for the first time in handwriting literature. In our experiments we have evaluated the process up to the instantiation. We have found that using not only the top alternative, but also the k-best alternatives increases the performance of information extraction. Furthermore, the use of an ontology-based lexicon results in another performance increase.
Sebastian Ebert, Marcus Liwicki, Andreas Dengel 0001
ICFHR3
2010 Handwriting Reconstruction for a Camera Pen Using Random Dot Patterns
abstract
This paper proposes a new method of handwriting reconstruction using a camera pen. We print random dot patterns on the document background to enable retrieval of both the current document and the pen position on this document. Dot arrangements are stored in a hash table using Locally Likely Arrangement Hashing. For retrieval, they are extracted from the camera image and matched to the corresponding points in the hash table. We were able to achieve high retrieval accuracy (81.1~100.0%), given a sufficient amount of visible dots. Using a two-step homography approximation, an accurate image of handwriting can be reconstructed. By using knowledge about document context and a client-server architecture, our method allows real-time processing on ordinary hardware.
Matthias Sperber, Martin Klinkigt, Koichi Kise, Masakazu Iwamura, Benjamin Adrian, Andreas Dengel 0001
ICFHR6
2010 a.SCAtch - A Sketch-Based Retrieval for Architectural Floor Plans
abstract
Architects' daily routine means working with drawings. They use either a pen or a computer sketching their ideas or drawing to scale. When beginning a new project they often have to search for similar projects in the past. In this paper a sketch-based approach is proposed to query the floor plan repository. The user searches for semantically similar floor plans just by drawing the new plan. An algorithm extracts the semantic structure sketched by the architect on DFKI's Touch & Write table and compares the structure of the sketch with the ones from the floor plan repository. The a SCatch system enables the user to easily access knowledge from past projects. While in the current prototype only sketches with a predefined structure are recognized, we will extend the system to work with normal floor plans.
Marcus Liwicki, Andreas Dengel 0001
ICFHR3
2010 Combining Patient Metadata Extraction and Automatic Image Parsing for the Generation of an Anatomic Atlas
Manuel Möller, Patrick Ernst, Michael Sintek, Sascha Seifert, Gunnar Aastrand Grimnes, Alexander Cavallaro, Andreas Dengel 0001
KES (1)7
2009 Helping People Remember: Coactive Assistance for Personal Information Management on a Semantic Desktop
Andreas Dengel 0001, Benjamin Adrian
IC3K1
2009 Seizing the Treasure: Transferring Knowledge in Invoice Analysis
abstract
This paper deals with the transfer of knowledge on invoice document layout and extraction strategies, collected by users of the invoice recognition software smartFIX over several years of productive use, to other user's systems. The results of a project analyzing this 'treasure' of knowledge and putting it to use in the smartFIX system are presented. The evaluation shows that this transfer of knowledge using state-of-the-art techniques in transfer learning achieves significantly higher initial recognition rates than the unaugmented system, delivering instant economic advantages by reducing accountant personnel workload.
Frederick Schulz, Markus Ebbecke, Michael Gillmann, Benjamin Adrian, Stefan Agne, Andreas Dengel 0001
ICDAR6
2009 Segment-level display time as implicit feedback: a comparison to eye tracking
abstract
We examine two basic sources for implicit relevance feedback on the segment level for search personalization: eye tracking and display time. A controlled study has been conducted where 32 participants had to view documents in front of an eye tracker, query a search engine, and give explicit relevance ratings for the results. We examined the performance of the basic implicit feedback methods with respect to improved ranking and compared their performance to a pseudo relevance feedback baseline on the segment level and the original ranking of a Web search engine.
Georg Buscher, Ludger van Elst, Andreas Dengel 0001
SIGIR3
2008 Attention-Based Document Classifier Learning
abstract
We describe an approach for creating precise personalized document classifiers based on the user's attention. The general idea is to observe which parts of a document the user was interested in just before he or she comes to a classification decision. Having information about this manual classification decision and the document parts the decision was based on, we can learn precise classifiers. For observing the user's focus point of attention we use an unobtrusive eye tracking device and apply an algorithm for reading behavior detection. On this basis, we can extract terms characterizing the text parts interesting to the user and employ them for describing the class the document was assigned to by the user. Having learned classifiers in that way, new documents can be classified automatically using techniques of passage-based retrieval. We prove the very strong improvement of incorporating the user's visual attention by a case study that evaluates an attention-based term extraction method.
Georg Buscher, Andreas Dengel 0001
Document Analysis Systems2
2008 The HCI Paradigm of HyperPrinting
abstract
Today, printing and reverse printing (scanning, OCR, logical labeling etc.) technologies have become quite mature and thus allow for an easy transition of documents between physical and electronic world. However, there is no technology today which supports the lossless interpretation of paper-based user interaction with direct effects upon the electronic representation of that document. The HyperPrinting environment tries to fill in this gap and thus accounts for the personal favors of a majority of office workers: Not only managers and knowledge workers prefer to read longer documents, articles or news from paper in contrast to a computer monitor or handheld computer. With the help of HyperPrinting, users can annotate, send notes or initiate tasks and it thus offers a completely new paradigm in the usage and treatment of paper documents. As a side-effect, the use of HyperPrinting builds up a document repository which is not only searchable by full text but also by meta-information, which in turn is depending on the selected user scenario.
Thomas Kieninger, Andreas Dengel 0001
Document Analysis Systems2
2008 Contextualized Knowledge Acquisition in a Personal Semantic Wiki
Ludger van Elst, Malte Kiesel, Sven Schwarz, Georg Buscher, Andreas Lauer, Andreas Dengel 0001
EKAW6
2008 Managing a document-based information space
abstract
We present a novel user interface in the form of a complementary virtual environment for managing personal document archives, i.e., for document filing and retrieval. Our implementation of a spatial medium for document interaction, exploratory search and active navigation plays to the strengths of human visual information processing and further stimulates it.
Matthias Deller, Stefan Agne, Achim Ebert, Andreas Dengel 0001, Hans Hagen, Bertin Klein, Tony Bernardin, Bernd Hamann
IUI4
2008 Query expansion using gaze-based feedback on the subdocument level
abstract
We examine the effect of incorporating gaze-based attention feedback from the user on personalizing the search process. Employing eye tracking data, we keep track of document parts the user read in some way. We use this information on the subdocument level as implicit feedback for query expansion and reranking.
Georg Buscher, Andreas Dengel 0001, Ludger van Elst
SIGIR2
2008 Visualizing personal trend on gnowsis Semantic Desktop
abstract
To let a user know the user's personal trend on the user's computer, this paper proposes a method to explore and visualize one's personal trend on the premise that the Gnowsis semantic desktop is used. The proposed method classifies information items into four attention areas with a popularity change measure by using three indicators (direct attention score (DAS), indirect attention score (IAS) (using spreading activation) and popularity change (PC)). Comparing the proposed method with two existing tools (PIMO timeline and PIMO cloud) in the gnowsis, this paper describes the following advantages about the proposed method: (1) To deal with both touched concepts and neighbor concepts, (2) To consider frequency, (3) To consider trend, (4) To catch seven operations in the gnowsis. Moreover, from an experimental usage, which shows more than three subjects answered the proposed method provided good results in every question, this paper suggests the proposed method would be useful to show one's personal trend.
Shingo Kubo, Hiroshi Tsuji, Heiko Maus, Andreas Dengel 0001
SMC4
2008 IVIP - A Scientific Workflow System to Support Experts in Spatial Planning of Crop Production
Christopher J. Tuot, Michael Sintek, Andreas Dengel 0001
SSDBM3
2007 Learning of Pattern-Based Rules for Document Classification
abstract
Automatic processing of office documents, such as orders, invoices, or offers entails a significant potential for saving costs. Because such domains have a high percentage of special vocabulary, purely statistical approaches fail in automatic classification. The inherent structure and short text messages require specific approaches. We propose a rule-based method to classify mixed stacks of documents into a set of hierarchically organized classes. Rules are learned by extracting patterns of different types from a document sample. The paper focuses on the architecture and on the learning process, presents comparing results to other techniques, and gives an outlook on how to further improve the system.
Andreas Dengel 0001
ICDAR1
2007 Knowledge Technologies for the Social Semantic Desktop
Andreas Dengel 0001
KSEM1
2006 Task-based process know-how reuse and proactive information delivery in TaskNavigator
abstract
Knowledge management approaches for weakly-structured, adhoc knowledge work processes need to be lightweight, i.e., they cannot rely on high upfront modeling efforts. This paper presents TaskNavigator, a novel prototype to support weakly-structured processes by integrating a standard task list application with a state-of-the-art document classification system. The resulting system allows for a task-oriented view on office workers' personal knowledge spaces in order to realize a proactive and contextsensitive information support during daily, knowledge-intensive tasks. Moreover, TaskNavigator supports process know-how reuse by proactively suggesting similar tasks or relevant process models, based on textual similarities. Finally, we report on a feasibility test and a case study that have been conducted in order to evaluate the system in the context of daily research task management and software requirements analysis.
Harald Holz, Oleg Rostanin, Andreas Dengel 0001, Takeshi Suzuki, Kaoru Maeda, Katsumi Kanasaki
CIKM3
2006 On Benchmarking of Invoice Analysis Systems
Bertin Klein, Stefan Agne, Andreas Dengel 0001
Document Analysis Systems3
2006 Semantic Desktop 2.0: The Gnowsis Experience
Leo Sauermann, Gunnar Aastrand Grimnes, Malte Kiesel, Christiaan Fluit, Heiko Maus, Dominik Heim, Danish Nadeem, Benjamin Horak, Andreas Dengel 0001
ISWC9
2005 An Approach towards Benchmarking of Table Structure Recognition Results
abstract
After we developed a model free table recognition system we had the desire to automatically register the effect of minor changes to parameters upon the overall performance quality of our system in order to tune parameters. Therefore we developed a complete benchmarking environment, containing a user front-end to acquire ground truth data as well as mechanisms to evaluate the quality of the recognition results. The tasks involved in the analysis systems were the locating of table regions, identification of cells and mapping of cells to rows and columns. This paper presents our approach towards the comparison of recognition results with the ground truth. The established definitions of recall and precision did not meet our requirements, as we wanted to register even smallest improvements (or changes in general) in the results, even when both results were imperfect. We therefore extended the measures recall and precision in order to deal with recognition probabilities of objects rather than just with Boolean values.
Thomas Kieninger, Andreas Dengel 0001
ICDAR2
2004 Results of a Study on Invoice-Reading Systems in Germany
Bertin Klein, Stefan Agne, Andreas Dengel 0001
Document Analysis Systems3
2003 Evaluating SEE - A Benchmarking System for Document Page Segmentation
abstract
The decomposition of a document into segments such as text regions and graphics is a significant part of the document analysis process. The basic requirement for rating and improvement of page segmentation algorithms is systematic evaluation. The approaches known from the literature have the disadvantage that manually generated reference data (zoning ground truth) are needed for the evaluation task. The effort and cost of the creation of these data are very high. This paper describes the evaluation system SEE and presents an assessment of its quality. The system requires the OCR generated text and the original text of the document in correct reading order (text ground truth) as input. No manually generated zoning ground truth is needed. The implicit structure information that is contained in the text ground truth is used for the evaluation of the automatic zoning. Therefore, an assignment of the corresponding text regions in the text ground truth and those in the OCR generated text (matches) is sought. A fault tolerant string matching algorithm underlies a method, able to tolerate OCR errors in the text. The segmentation errors are determined as a result of the evaluation of the matching. Subsequently, the edit operations which are necessary for the correction of the recognized segmentation errors are computed to estimate the correction costs. Furthermore, SEE provides a version of the OCR generated text, which is corrected from the detected page segmentation errors.
Stefan Agne, Andreas Dengel 0001, Bertin Klein
ICDAR2
2003 Making Documents Work: Challenges for Document Understanding
abstract
In this paper I will try to explain the nature of documentunderstanding in all of its dimensions. Therefore I willfirst describe the characteristics of data, knowledge, andinformation in order to describe their synergetic inter-weaving.After that I will try to structure the inherentcomplexity of sub-problems of document understandingwhich may not be solved serially, but rather are attributesof individual documents. Thus, this paper focuses onsystem engineering challenges. However, I will showsome recent work done on the different topics and givesome insights in the individual techniques we chose atDFKI.
Andreas Dengel 0001
ICDAR1
2003 Problem-adaptable document analysis and understanding for high-volume applications
Bertin Klein, Andreas Dengel 0001
Int. J. Document Anal. Recognit.2
2002 smartFIX: A Requirements-Driven System for Document Analysis and Understanding
Andreas Dengel 0001, Bertin Klein
Document Analysis Systems1
2002 Improving Document Retrieval by Automatic Query Expansion Using Collaborative Learning of Term-Based Concepts
Stefan Klink, Armin Hust, Markus Junker 0002, Andreas Dengel 0001
Document Analysis Systems4
2002 Collaborative Learning of Term-Based Concepts for Automatic Query Expansion
Stefan Klink, Armin Hust, Markus Junker 0002, Andreas Dengel 0001
ECML4
2001 Passage-Based Document Retrieval as a Tool for Text Mining with User's Information Needs
Koichi Kise, Markus Junker 0002, Andreas Dengel 0001, Keinosuke Matsumoto
Discovery Science3
2001 Applying the T-Recs Table Recognition System to the Business Letter Domain
abstract
This paper summarizes the core idea of the T-Recs table recognition system, an integrated system covering block-segmentation, table location and a model-free structural analysis of tables. T-Recs works on the output of commercial OCR systems that provide the word bounding box geometry together with the text itself (e.g. Xerox ScanWorX). While T-Recs performs well on a number of document categories, business letters still remained a challenging domain because the T-Recs location heuristics are mislead by their header or footer resulting in a low recognition precision. Business letters such as invoices are a very interesting domain for industrial applications due to the large amount of documents to be analyzed and the importance of the data carried within their tables. Hence, we developed a more restrictive approach which is implemented in the T-Recs++ prototype. This paper describes the ideas of the T-Recs++ location and also proposes a quality evaluation measure that reflects the bottom-up strategy of either T-Recs or T-Recs++. Finally, some results comparing both systems on a collection of business letters are given.
Thomas Kieninger, Andreas Dengel 0001
ICDAR2
2001 Experimental Evaluation of Passage-Based Document Retrieval
abstract
Retrieval of electronic documents is a fundamental component for intelligent access to the contents of documents. For the retrieval of long documents, a method called passage-based document retrieval has proven to be effective. In this paper we experimentally show that the passage-based retrieval is also advantageous for dealing with short queries on condition that documents are long. We employ a passage-based method based on density distributions of query terms in documents, and compare it with three conventional methods: the vector space model, pseudo-feedback and latent semantic indexing.
Koichi Kise, Markus Junker 0002, Andreas Dengel 0001, Keinosuke Matsumoto
ICDAR3
2001 Three Approaches to "Industrial" Table Spotting
abstract
This paper introduces three approaches for an industrial, comprehensive document analysis system to enable it to spot tables in documents. Searching for a set of known table headers (approach 1) works rather well in a significant number of documents. But this approach (though it is implemented tolerant to OCR errors) is not tolerant enough towards some kinds of even minor aberrations. This not only decreases the recognition results, but also, even worse, makes users feel uncomfortable. Pragmatically trying to mimic for what the human eyes might key, leads to our two further, complementary approaches: searching for layout structures which resemble parts of columns (approach 2), and searching for groupings of similar lines (approach 3). The suitability of the approaches for our system requires them to be very simple to implement and simple to explain to users, computationally cheap, and combinable. In the domain of health insurances who receive huge amounts of so called medical liquidations on a daily basis we obtain very good results. On document samples representative for the every day practice of five customers-health insurance companies-tables were spotted as good and as fast as the customers expected the system to be. We thus consider our current approaches as a step towards cognitive adequacy.
Bertin Klein, Serdar Gökkus, Thomas Kieninger, Andreas Dengel 0001
ICDAR4
2001 Guest Editorial
Andreas Dengel 0001, Markus Junker 0002
Int. J. Document Anal. Recognit.1
2001 Editorial
Andreas Dengel 0001, Markus Junker 0002
Int. J. Document Anal. Recognit.1
1999 On the Evaluation of Document Analysis Components by Recall, Precision, and Accuracy
abstract
In document analysis, it is common to prove the usefulness of a component by an experimental evaluation. By applying the respective algorithms to a test sample, effectiveness measures such as recall, precision, and accuracy are computed. The goal of such an evaluation is two-fold: on the one hand it shows that the absolute effectiveness of the algorithm is acceptable for practical use. On the other hand the evaluation can prove that the algorithm has a better or worse effectiveness than another algorithm. We argue that the experimental evaluation on relative small test sets-as is very common in document analysis has to be taken with extreme care from a statistical point of view. In fact, it is surprising how weak statements derived from such evaluations are.
Markus Junker 0002, Andreas Dengel 0001, Rainer Hoch
ICDAR2
1999 Quality Evaluation of Document Segmentation Results
abstract
Summary form only given, as follows. Increasing the performance of document analysis systems requires a detailed quality evaluation of the achieved results. By focussing on segmentation algorithms, we point out that the results produced from the module under consideration should be evaluated directly; we show that the text based evaluation method which is often used in the document analysis domain is not sufficient for the purpose of a detailed quality evaluation of the segmentation module. Therefore, we propose a general evaluation approach for comparing segmentation results which is based on the segments directly. This approach is able to handle both algorithms which produce complete segmentations (partition) and algorithms which only extract objects of interest (extraction). Classes of errors are defined in a systematic way and frequencies for each class can be computed. The evaluation approach is applicable to segmentation or extraction algorithms in a wide range. We have chosen the character segmentation task as an example to demonstrate the applicability of our evaluation approach and we suggest applying our approach to other segmentation tasks.
Michael Thulke, Volker Märgner, Andreas Dengel 0001
ICDAR3
1998 The T-Recs Table Recognition and Analysis System
Thomas Kieninger, Andreas Dengel 0001
Document Analysis Systems2
1998 Text-Line Extraction as Selection of Paths in the Neighbor Graph
Koichi Kise, Motoi Iwata, Andreas Dengel 0001, Keinosuke Matsumoto
Document Analysis Systems3
1998 A General Approach to Quality Evaluation of Document Segmentation Results
Michael Thulke, Volker Märgner, Andreas Dengel 0001
Document Analysis Systems3
1997 Message Extraction from Printed Documents - A Complete Solution
abstract
The task to be solved within our core research was the design and development of a document analysis toolbox covering typical document analysis tasks such as document understanding, information extraction and text recognition. In order to prove the feasibility of our concepts, we have developed the prototypical analysis system OfficeMAID (Office Mail Analysis, Interpretation and Delivery). The system analyses documents, as used in the daily work of a purchasing department, by a priori knowledge about workflows and document features. In this way, the system provides goal-directed information extraction, shallow understanding and process identification for given documents (paper, fax, e-mail).
Stephan Baumann 0001, Majdi Ben Hadj Ali, Andreas Dengel 0001, Thorsten Jäger, Michael Malburg, Achim Weigel, Claudia Wenzel
ICDAR3
1997 Real Time Object Detection, Tracking and Classification in Monocular Image Sequences of Road Traffic Scenes
abstract
Automatic traffic scene analysis is very interesting in the context of traffic planning and monitoring. In cooperation with the Traffic and Transport Research Group at the University of Kaiserslautern we conceived and implemented an automatic real rime object detection, tracking and classification system working on color image sequences, taken by a static camera. The system is employed by researcher to detect conflict situations in traffic scenes. The main requirements for the system were real time ability and implementation on cheap standard hardware.
Majdi Ben Hadj Ali, Markus Ebbecke, Andreas Dengel 0001
ICIP (2)3
1997 Selecting distinctive attributes for concept learning
abstract
This paper presents an innovative approach for learning the distinctive attributes of uncertain objects. The proposed system takes instances, clusters them into different concepts and consequently induces a hierarchy which is used for later classification. We introduce the major steps of the approach using a set of city attributes and further illustrate the applicability for a real world problem, namely the learning of structural concepts of business letters.
Andreas Dengel 0001, Frank Dubiel
KES (1)1
1996 Formclas - a System for OCR Free identification of Forms
Frank Dubiel, Andreas Dengel 0001
DAS2
1996 Document Analysis and Learning: das'96 Working Group Report
Dar-Shyang Lee, Andreas Dengel 0001
DAS2
1995 Clustering and classification of document structure-a machine learning approach
abstract
We describe a system which is capable of learning the presentation of document logical structures, exemplarily shown for business letters. Presenting a set of instances to the system, it clusters them into structural concepts and induces a concept hierarchy. This concept hierarchy is taken as a source for classifying future input. The paper introduces the different learning steps, describes how the resulting concept hierarchy is applied for logical labeling and reports on the results.
Andreas Dengel 0001, Frank Dubiel
ICDAR1
1995 Post-processing of OCR results for automatic indexing
abstract
The indexing of inaccurately recognized OCR text yields unsatisfactory results, where the quality of the index terms decreases rapidly when the quality of the documents get worse. Index terms of OCR processed documents can be used for archiving or classification tasks. We present an indexing component whose input are character hypothesis lattices which are post-processed by a generate-and-test component feeding a morphology, a rule based substitution system, and a trigram correction component with word candidates. Stop words are filtered by a Levenshtein-based elimination routine. The recognized words are subsequently processed by our indexing component. Our system minimizes the number of generated index terms which are correct German words. The experiments have shown an increase in accuracy of next to 10%.
Lars Wiedenhifer, Hans-Günther Hein, Andreas Dengel 0001
ICDAR3
1995 Syntactic Analysis and Representation of Spatial Structures by Puzzletrees
abstract
The objective of this paper is to propose a syntactic formalism for space representation, which besides the well known advantages of hierarchical data structure, has the additional strength of self-adapting to a spatial structure at hand. The formalism is called puzzletree because its generation results in a number of blocks which in a certain order — like a puzzle — reconstruct the original space. The approach may be applied to any higher-dimensioned space (e.g. images, volumes). The paper concentrates on the principles of puzzletrees by explaining the underlying heuristic for their generation with respect to 2D spaces, i.e. images.
Andreas Dengel 0001
Int. J. Pattern Recognit. Artif. Intell.1
1993 Initial learning of document structure
abstract
Proposes an approach for automatically generating a decision tree which is applied as a model for the logical labeling of business letters. Instead of top-down determination of the discriminating attributes, the system inspects a finite set of document instances that are presented to a learner in a bottom-up position. The learner itself figures out local similarities, rates them with respect to the overall structure, and determines the best structural match of two instances (neighborhood). The entire decision tree is grown step by step deducing subtrees by forming generalizations from a neighborhood. Consequently, heuristics are learned for structurally discriminating documents during subsequent classification.>
Andreas Dengel 0001
ICDAR1
1993 The role of document analysis and understanding in multi-media information systems
abstract
One obstacle for introducing a general and homogeneous information system is the lack of possibilities for combining the information processing capabilities of the electronic medium and traditional ways in which information is generated, interchanged, categorized, processed or stored. In this context, document analysis and understanding technology seems to be the key to overcome this problem. An overview of efforts focusing on the integration of document analysis and understanding techniques into active interfaces for multimedia information systems is given. Moreover, the role and the potential of this technology are outlined using sample application scenarios.>
Andreas Dengel 0001
ICDAR1
1992 Fragmentary string matching by selective access to hybrid tries
abstract
The authors propose a dictionary look-up method as a contextual postprocessing for character hypotheses forming word candidates. In particular, a hybrid trie organization is combined with a selective-access-matrix (SAM) that allows an efficient matching of fragmentary input strings against legal words. Experiments prove that the method achieves some respectable results concerning speed. Furthermore, the additional memory needed for the SAM is smaller than the memory saved by the hybrid organization of the trie.>
Andreas Dengel 0001, Adolf Pleyer, Rainer Hoch
ICPR (2)1
1989 ANASTASIL: A Hybrid Knowledge-Based System for Document Layout Analysis
Andreas Dengel 0001, Gerhard Barth
IJCAI1
1988 High Level Document Analysis Guided by Geometric Aspects
abstract
The realization of the paper-free office seems to be difficult that expected. Therefore, good paper-computer interfaces are necessary to transform paper documents into an electronic form, which allows the use of a filing and retrieval system. An electronic document page is an optically scanned and digitized representation of a printed page. Document analysis is the problem of interpreting and labeling the constitutents of the document. Although there are very reliable optical character recognition (OCR) methods, the process could be very inefficient. To prune the search space and to become more efficient, some search supporting methods have to be developed. This article proposes an approach to identify the layout of a document page by dividing it recursively into nested rectangular areas. The procedure is used as a basis for a document layout model, which is able to control an automatic interpretation mechanism for deriving a high level representation of the contents of a document. We have implemented our method in Common Lisp on a Symbolies 3640 Workstation and have run it for a large population of office documents. The results obtained have been very encouraging and have convincingly confirmed the soundness of our approach.
Andreas Dengel 0001, Gerhard Barth
Int. J. Pattern Recognit. Artif. Intell.1