VLDB 2026 Research / reviewers in the wild / expert
Sheraz Ahmed
dblp:81/10400
· DBLP profile ↗
94ranked-venue papers
8as first author
52since 2021 · last 2026
0000-0002-4239-6520ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 81 · 5 first-author · 46 since 2021Databases, data management, data science and information retrieval · 27 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can Generative Adversarial Networks Compete against Diffusion Models at Generating Time Series?
Philipp Engler, Ludger van Elst, Sheraz Ahmed, Andreas Dengel 0001 |
ICAART (4) | 3 |
| 2026 | FREQuency ATTribution: benchmarking frequency-based occlusion for time series data
Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed |
Appl. Intell. | 3 |
| 2025 | SAT: Segment and Track Anything for MicroscopyabstractIntegrating cell segmentation with tracking is critical for achieving a detailed and dynamic understanding of cellular behavior. This integration facilitates the study and quantification of cell morphology, movement, and interactions, offering valuable insights into a wide range of biological processes and diseases. However, traditional methods rely on labor-intensive and costly annotations, such as full segmentation masks or bounding boxes for each cell. To address this limitation, we present SAT: Segment and Track Anything for Microscopy, a novel pipeline that leverages point annotations in the first frame to automate cell segmentation and tracking across all subsequent frames. By significantly reducing annotation time and effort, SAT enables efficient and scalable analysis, making it well-suited for large-scale studies. The pipeline was evaluated on two diverse datasets, achieving over 80% Multiple Object Tracking Accuracy (MOTA), demonstrating its robustness and effectiveness across various imaging modalities and cell types. These results highlight SAT’s potential to streamline biomedical research and enable deeper exploration of cellular behavior. Nabeel Khalid, Mohammadmahdi Koochali, Khola Naseem, Maria Caroprese, Gillian Lovell, Daniel A. Porto, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (2) | 9 |
| 2025 | Synthesizing Annotated Cell Microscopy Images with Generative Adversarial Networks
Duway Nicolas Lesmes-Leon, Miro Miranda, Maria Caroprese, Gillian Lovell, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (3) | 6 |
| 2025 | YeastFormer: An End-to-End Instance Segmentation Approach for Yeast Cells in Microstructure Environment
Khola Naseem, Nabeel Khalid, Lea Bertgen, Johannes M. Herrmann, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (2) | 6 |
| 2025 | How to Box Your Cells: An Introduction to Box Supervision for 2.5D Cell Instance Segmentation and a Study of Applications
Fabian Schmeisser, Maria Caroprese, Gillian Lovell, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (3) | 5 |
| 2025 | Time Series Generation for Augmenting Multi-channel Automotive Audio Data
Philipp Engler, Ludger van Elst, Peter Schichtel, Andreas Dengel 0001, Sheraz Ahmed |
ICANN (3) | 5 |
| 2025 | DocForgeNet: Dual Cross-Stream Fusion Network for Robust Forgery Detection in Scanned Documents
Nauman Riaz, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (4) | 4 |
| 2025 | DP-DocLDM: Differentially Private Document Image Generation Using Latent Diffusion Models
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (4) | 4 |
| 2025 | VAEneu: a new avenue for VAE application on probabilistic forecastingabstractThis paper introduces VAEneu, a novel autoregressive method for multistep ahead univariate probabilistic time series forecasting, designed to address the challenges of generating sharp and well-calibrated probabilistic forecasts without assuming a specific parametric form for the predictive distribution. VAEneu leverages the Conditional VAE framework and optimizes the likelihood of the predictive distribution using the Continuous Ranked Probability Score (CRPS), a strictly proper scoring rule, as the loss function. This approach enables the model to learn flexible, sharp, and well-calibrated predictive distributions without the need for a tractable likelihood function. In a comprehensive empirical study, VAEneu is rigorously benchmarked against 12 baseline models across 12 datasets, demonstrating superior performance in both forecasting accuracy and uncertainty quantification. VAEneu provides a valuable tool for quantifying future uncertainties, and our extensive empirical study lays the foundation for future comparative studies for univariate multistep ahead probabilistic forecasting. Alireza Koochali, Ensiye Tahaei, Andreas Dengel 0001, Sheraz Ahmed |
Appl. Intell. | 4 |
| 2025 | KRNN: A hybrid data and knowledge oriented time series forecasting approach for health care applications
Muhammad Ali Chattha, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
Expert Syst. Appl. | 4 |
| 2025 | Addressing data dependency in neural networks: introducing the Knowledge Enhanced Neural Network (KENN) for time series forecasting +
Muhammad Ali Chattha, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
Mach. Learn. | 4 |
| 2025 | Cloud segmentation in satellite imagery using attention-based deep learning
Khola Naseem, Muhammad Imran Malik, Sheraz Ahmed |
Multim. Tools Appl. | 3 |
| 2024 | Medi-CAT: Contrastive Adversarial Training for Medical Image Classification
Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (3) | 3 |
| 2024 | A Unique Training Strategy to Enhance Language Models Capabilities for Health Mention Detection from Social Media Content
Pervaiz Iqbal Khan, Muhammad Nabeel Asim, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (3) | 4 |
| 2024 | CellSpot: Deep Learning-Based Efficient Cell Center Detection in Microscopic Images
Nabeel Khalid, Maria Caroprese, Gillian Lovell, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed |
ICANN (8) | 6 |
| 2024 | Point-Based Weakly Supervised 2.5D Cell Segmentation
Fabian Schmeisser, Andreas Dengel 0001, Sheraz Ahmed |
ICANN (8) | 3 |
| 2024 | Latent Diffusion for Guided Document Table Generation
Syed Jawwad Haider Hamdani, Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (5) | 5 |
| 2024 | StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation
Nauman Riaz, Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (2) | 5 |
| 2024 | DocXplain: A Novel Model-Agnostic Explainability Method for Document Image Classification
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (4) | 4 |
| 2024 | CCATS: Moving Forward with Class-Conditional Time Series Generation
Philipp Engler, Alireza Koochali, Ludger van Elst, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (3) | 5 |
| 2024 | Improving Text Representation for Disease Detection from Social Media via Self-augmentation and Contrastive Learning
Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (5) | 3 |
| 2024 | Generating Counterfactual Trajectories with Latent Diffusion Models for Concept Discovery
Payal Varshney, Adriano Lucieri, Christoph Peter Balada, Andreas Dengel 0001, Sheraz Ahmed |
ICPR (12) | 5 |
| 2024 | From private to public: benchmarking GANs in the context of private time series classification
Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed |
Appl. Intell. | 3 |
| 2024 | DocXclassifier: towards a robust and interpretable deep neural network for document image classification
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
Int. J. Document Anal. Recognit. | 4 |
| 2024 | Towards privacy preserved document image classification: a comprehensive benchmark
Saifullah Saifullah, Dominique Mercier, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
Int. J. Document Anal. Recognit. | 5 |
| 2023 | Randout-KD: Finetuning Foundation Models for Text Classification via Random Noise and Knowledge Distillation
Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (3) | 3 |
| 2023 | Knowledge Forcing: Fusing Knowledge-Driven Approaches with LSTM for Time Series Forecasting
Muhammad Ali Chattha, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
ICANN (6) | 4 |
| 2023 | PACE: Point Annotation-Based Cell Segmentation for Efficient Microscopic Image Analysis
Nabeel Khalid, Tiago Comassetto Fróes, Maria Caroprese, Gillian Lovell, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed |
ICANN (2) | 7 |
| 2023 | ColDBin: Cold Diffusion for Document Image Binarization
Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (5) | 4 |
| 2023 | Deep Learning Architectures for the Prediction of YY1-Mediated Chromatin Loops
Ahtisham Fazeel Abbasi, Muhammad Nabeel Asim, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed |
ISBRA | 5 |
| 2023 | Quantifying quality of class-conditional generative models in time series domainabstractAbstract Despite recent breakthroughs in the domain of implicit generative models, the task of evaluating these models remains a challenging task. With no single metric to assess overall performance, various existing metrics only offer partial information. This issue is further compounded for unintuitive data types such as time series, where manual inspection is infeasible. This deficiency hinders the confident application of modern implicit generative models on time series data. To alleviate this problem, we propose two new metrics, the InceptionTime Score (ITS) and the Fréchet InceptionTime Distance (FITD), to assess the quality of class-conditional generative models on time series data. We conduct extensive experiments on 80 different datasets to study the discriminative capabilities of proposed metrics alongside two existing evaluation metrics: Train on Synthetic Test on Real (TSTR) and Train on Real Test on Synthetic (TRTS). Our evaluations reveal that the proposed assessment evaluation metrics, i.e., ITS and FITD in combination with TSTR, can accurately assess class-conditional generative model performance and detect common issues in implicit generative models. Our findings suggest that the proposed evaluation framework can be a valuable tool for confidently applying modern implicit generative models in time series analysis. Alireza Koochali, Maria Walch, Sankrutyayan Thota, Peter Schichtel, Andreas Dengel 0001, Sheraz Ahmed |
Appl. Intell. | 6 |
| 2023 | DNA-MP: a generalized DNA modifications predictor for multiple species based on powerful sequence encoding methodabstractAccurate prediction of deoxyribonucleic acid (DNA) modifications is essential to explore and discern the process of cell differentiation, gene expression and epigenetic regulation. Several computational approaches have been proposed for particular type-specific DNA modification prediction. Two recent generalized computational predictors are capable of detecting three different types of DNA modifications; however, type-specific and generalized modifications predictors produce limited performance across multiple species mainly due to the use of ineffective sequence encoding methods. The paper in hand presents a generalized computational approach "DNA-MP" that is competent to more precisely predict three different DNA modifications across multiple species. Proposed DNA-MP approach makes use of a powerful encoding method "position specific nucleotides occurrence based 117 on modification and non-modification class densities normalized difference" (POCD-ND) to generate the statistical representations of DNA sequences and a deep forest classifier for modifications prediction. POCD-ND encoder generates statistical representations by extracting position specific distributional information of nucleotides in the DNA sequences. We perform a comprehensive intrinsic and extrinsic evaluation of the proposed encoder and compare its performance with 32 most widely used encoding methods on $17$ benchmark DNA modifications prediction datasets of $12$ different species using $10$ different machine learning classifiers. Overall, with all classifiers, the proposed POCD-ND encoder outperforms existing $32$ different encoders. Furthermore, combinedly over 5-fold cross validation benchmark datasets and independent test sets, proposed DNA-MP predictor outperforms state-of-the-art type-specific and generalized modifications predictors by an average accuracy of 7% across 4mc datasets, 1.35% across 5hmc datasets and 10% for 6ma datasets. To facilitate the scientific community, the DNA-MP web application is available at https://sds_genetic_analysis.opendfki.de/DNA_Modifications/. Muhammad Nabeel Asim, Muhammad Ali Ibrahim, Ahtisham Fazeel, Andreas Dengel 0001, Sheraz Ahmed |
Briefings Bioinform. | 5 |
| 2023 | IAMonSense: multi-level handwriting classification using spatiotemporal information
Ahmad Mustafid, Junaid Younas, Paul Lukowicz, Sheraz Ahmed |
Int. J. Document Anal. Recognit. | 4 |
| 2023 | Analyzing the potential of active learning for document image classificationabstractAbstract Deep learning has been extensively researched in the field of document analysis and has shown excellent performance across a wide range of document-related tasks. As a result, a great deal of emphasis is now being placed on its practical deployment and integration into modern industrial document processing pipelines. It is well known, however, that deep learning models are data-hungry and often require huge volumes of annotated data in order to achieve competitive performances. And since data annotation is a costly and labor-intensive process, it remains one of the major hurdles to their practical deployment. This study investigates the possibility of using active learning to reduce the costs of data annotation in the context of document image classification, which is one of the core components of modern document processing pipelines. The results of this study demonstrate that by utilizing active learning (AL), deep document classification models can achieve competitive performances to the models trained on fully annotated datasets and, in some cases, even surpass them by annotating only 15–40% of the total training dataset. Furthermore, this study demonstrates that modern AL strategies significantly outperform random querying, and in many cases achieve comparable performance to the models trained on fully annotated datasets even in the presence of practical deployment issues such as data imbalance, and annotation noise, and thus, offer tremendous benefits in real-world deployment of deep document classification models. The code to reproduce our experiments is publicly available at https://github.com/saifullah3396/doc_al . Saifullah Saifullah, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
Int. J. Document Anal. Recognit. | 4 |
| 2023 | Performance Comparison of Transformer-Based Models on Twitter Health Mention ClassificationabstractHealth mention classification classifies a given piece of text as a health mention or not. However, figurative usage of disease words makes the classification task challenging. To address this challenge, consideration of emojis and surrounding words of the disease names in the text can be helpful. Transformer-based methods are better at capturing the meaning of a word based on its surrounding words compared to traditional methods. However, there are numerous transformer-based methods available and pretrained on natural language processing (NLP) data that are inherently different from Twitter data. Moreover, the size of these models varies in terms of the number of parameters. Hence, it is challenging to decide and choose one of these methods for fine-tuning it on the downstream tasks such as tweet classification. In this work, we experiment with nine widely used transformer methods and compare their performance on the personal health mention classification of tweet data. Furthermore, we analyze the impact of model size on the classification task and provide a brief interpretation of the classification decision made by the best performing classifier. Experimental results show that RoBERTa outperforms all other models by achieving an F1 score of 93%, while two other models perform similarly by achieving an F1 score of 92.5%. Pervaiz Iqbal Khan, Muhammad Imran Razzak, Andreas Dengel 0001, Sheraz Ahmed |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | Time to Focus: A Comprehensive Benchmark using Time Series Attribution Methods
Dominique Mercier, Jwalin Bhatt, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (2) | 4 |
| 2022 | A Novel Approach to Train Diverse Types of Language Models for Health Mention Classification of Tweets
Pervaiz Iqbal Khan, Muhammad Imran Razzak, Andreas Dengel 0001, Sheraz Ahmed |
ICANN (2) | 4 |
| 2022 | Are Deep Models Robust against Real Distortions? A Case Study on Document Image ClassificationabstractAs deep learning models in the context of document image classification are reaching diminishing returns with nearperfect recognition scores, their robustness characteristics are poorly understood. In order to evaluate the robustness of existing state-of-the-art document image classifiers against different types of distortions that are commonly encountered in the real world, we present two separate benchmark datasets, namely RVL-CDIPD and Tobacco3482-D. The proposed benchmarks are generated by augmenting the well-known pre-existing document image classification datasets (RVL-CDIP and Tobacco3482) with 21 different types of distortions including varying severity levels. We leverage the proposed benchmark datasets to analyze the robustness characteristics of existing document image classification systems. Our analysis reveals that despite higher accuracy models exhibiting relatively higher robustness, they still severely underperform on some specific distortions, with classification accuracies dropping from ~90% to as low as ~40% in some cases. Interestingly, some of these high accuracy models perform even worse than the baseline AlexNet model in the presence of distortions, with the relative decline in their accuracy sometimes reaching as high as 300-450%. We envision these benchmarks to serve as a strong signal of progress in document image classification tasks, beyond the saturated accuracy metrics. The datasets and code to reproduce them is publicly available: https://github.com/saifullah3396/docrobustness. Saifullah Saifullah, Shoaib Ahmed Siddiqui, Stefan Agne, Andreas Dengel 0001, Sheraz Ahmed |
ICPR | 5 |
| 2022 | CircNet: an encoder-decoder-based convolution neural network (CNN) for circular RNA identification
Marco Stricker, Muhammad Nabeel Asim, Andreas Dengel 0001, Sheraz Ahmed |
Neural Comput. Appl. | 4 |
| 2022 | Evaluating Privacy-Preserving Machine Learning in Critical Infrastructures: A Case Study on Time-Series ClassificationabstractWith the advent of machine learning in applications of critical infrastructure such as healthcare and energy, privacy is a growing concern in the minds of stakeholders. It is pivotal to ensure that neither the model nor the data can be used to extract sensitive information used by attackers against individuals or to harm whole societies through the exploitation of critical infrastructure. The applicability of machine learning in these domains is mostly limited due to a lack of trust regarding the transparency and the privacy constraints. Various safety-critical use cases (mostly relying on time-series data) are currently underrepresented in privacy-related considerations. By evaluating several privacy-preserving methods regarding their applicability on time-series data, we validated the inefficacy of encryption for deep learning, the strong dataset dependence of differential privacy, and the broad applicability of federated methods. Dominique Mercier, Adriano Lucieri, Mohsin Munir, Andreas Dengel 0001, Sheraz Ahmed |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | ImpactCite: An XLNet-based Solution Enabling Qualitative Citation Impact Analysis Utilizing Sentiment and IntentabstractCitations play a vital role in understanding the impact of scientific literature. Generally, citations are analyzed quantitatively whereas qualitative analysis of citations can reveal deeper insights into the impact of a scientific artifact in the community. Therefore, citation impact analysis (which includes sentiment and intent classification) enables us to quantify the quality of the citations which can eventually assist us in the estimation of ranking and impact. The contribution of this paper is two-fold. First, we benchmark the well-known language models like BERT and ALBERT along with several popular networks for both tasks of sentiment and intent classification. Second, we provide ImpactCite, which is XLNet-based method for citation impact analysis. All evaluations are performed on a set of publicly available citation analysis datasets. Evaluation results reveal that ImpactCite achieves a new state-of-the-art performance for both citation intent and sentiment classification by outperforming the existing approaches by 3.44% and 1.33% in F1-score. Therefore, we emphasize ImpactCite (XLNet-based solution) for both tasks to better understand the impact of a citation. Additional efforts have been performed to come up with CSC-Clean corpus, which is a clean and reliable dataset for citation sentiment classification. Dominique Mercier, Syed Tahseen Raza Rizvi, Vikas Rajashekar, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (2) | 5 |
| 2021 | Understanding and Mitigating the Impact of Model Compression for Document Image Classification
Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (1) | 3 |
| 2021 | Analyzing the Potential of Zero-Shot Recognition for Document Image Classification
Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR (4) | 3 |
| 2021 | L2S-MirLoc: A Lightweight Two Stage MiRNA Sub-Cellular Localization Prediction FrameworkabstractA comprehensive understanding of miRNA sub-cellular localization may leads towards better understanding of physiological processes and support the fixation of diverse irregularities present in a variety of organisms. To date, diverse computational methodologies have been proposed to automatically infer sub-cellular localization of miR-NAs solely using sequence information, however, existing approaches lack in performance. Considering the success of data transformation approaches in Natural Language Processing which primarily transform multi-label classification problem into multi-class classification problem, here, we introduce three different data transformation approaches namely binary relevance, label power set, and classifier chains. Using data transformation approaches, at 1ststage, multi-label miRNA sub-cellular localization problem is transformed into multi-class problem. Then, at 2ndstage, 3 different machine learning classifiers are used to estimate which classifier performs better with what data transformation approach for hand on task. Empirical evaluation on independent test set indicates that L2S-MirLoc selected combination based on binary relevance and deep random forest outperforms state-of-the-art performance values by significant margin. Muhammad Nabeel Asim, Muhammad Ali Ibrahim, Christoph Zehe, Olivier Cloarec, Rickard Sjögren, Johan Trygg, Andreas Dengel 0001, Sheraz Ahmed |
IJCNN | 8 |
| 2021 | DeepCeNS: An end-to-end Pipeline for Cell and Nucleus Segmentation in Microscopic ImagesabstractWith the evolution of deep learning in the past decade, more biomedical related problems that seemed strenuous, are now feasible. The introduction of U-net and Mask R-CNN architectures has paved a way for many object detection and segmentation tasks in numerous applications ranging from security to biomedical applications. In the cell biology domain, light microscopy imaging provides a cheap and accessible source of raw data to study biological phenomena. By leveraging such data and deep learning techniques, human diseases can be easily diagnosed and the process of treatment development can be greatly expedited. In microscopic imaging, accurate segmentation of individual cells is a crucial step to allow better insight into cellular heterogeneity. To address the aforementioned challenges, DeepCeNS is proposed in this paper to detect and segment cells and nucleus in microscopic images. We have used EVICAN2 dataset which contains microscopic images from a variety of microscopes having numerous cell cultures, to evaluate the proposed pipeline. DeepCeNS outperforms EVICAN-MRCNN by a significant margin on the EVICAN2 dataset. Nabeel Khalid, Mohsin Munir, Christoffer Edlund, Timothy R. Jackson, Johan Trygg, Rickard Sjögren, Andreas Dengel 0001, Sheraz Ahmed |
IJCNN | 8 |
| 2021 | PatchX: Explaining Deep Models by Intelligible Pattern Patches for Time-series ClassificationabstractClassification of time-series data is pivotal for a wide range of applications and comes with many challenges. Although the amount of publicly available datasets increases rapidly, deep neural models are only partially exploited in contrast to the traditional methods. These methods get preferred in safety-critical, financial, or medical fields because of their interpretable results. However, their performance and scalability are limited, and finding suitable explanations for time-series classification tasks is challenging due to the intrinsic nature of concealed concepts in the time-series data. Visual analysis of complete time-series comes with an extensive cognitive overload, as it is difficult to perceive and leads to confusion. Therefore, we believe that patch-wise processing of the data results in a more interpretable representation. To bridge this gap, and to reduce the cognitive overload for interpretation of time series data, we propose a novel hybrid approach that utilizes deep neural networks and traditional machine learning algorithms for an interpretable and scale-able time-series classification approach. Both quantitively and qualitatively PatchX shows superiority to its counterparts with an edge of interoperability. Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed |
IJCNN | 3 |
| 2021 | Sense the pen: Classification of online handwritten sequences (text, mathematical expression, plot/graph)
Junaid Younas, Muhammad Imran Malik, Sheraz Ahmed, Faisal Shafait, Paul Lukowicz |
Expert Syst. Appl. | 3 |
| 2021 | Correction to: Benchmarking performance of machine and deep learning-based methodologies for Urdu text document classification
Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Ali Ibrahim, Waqar Mahmood, Andreas Dengel 0001, Sheraz Ahmed |
Neural Comput. Appl. | 6 |
| 2021 | Benchmarking performance of machine and deep learning-based methodologies for Urdu text document classification
Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Ali Ibrahim, Waqar Mahmood, Andreas Dengel 0001, Sheraz Ahmed |
Neural Comput. Appl. | 6 |
| 2021 | Learning the micro deformations by max-pooling for offline signature verification
Yuchen Zheng 0001, Brian Kenji Iwana, Muhammad Imran Malik, Sheraz Ahmed, Wataru Ohyama, Seiichi Uchida |
Pattern Recognit. | 4 |
| 2021 | Guest Editorial Computational Social Systems for COVID-19 Emergency Management and BeyondabstractSince early 2020, the COVID-19 global pandemic has significantly impacted almost every aspect of the human society throughout the world. Until now, middle of 2021, although with all the efforts on pandemic intervention and vaccination, COVID-19 is still hovering around the world, resulting in more than 177 million confirmed cases and 3.8 million deaths. Jun Jason Zhang, Fei-Yue Wang 0001, Yong Yuan 0003, Guandong Xu, Huan Liu 0001, Wei Gao 0001, Shoaib Jameel, Muhammad Imran Razzak, Peter W. Eklund, Sheraz Ahmed, Rui Qin 0002, Juanjuan Li, Xiao Wang 0002, De-Nian Yang, Damla Turgut, Abderrahim Benslimane, Neeli Prasad, Kwang-Cheng Chen |
IEEE Trans. Comput. Soc. Syst. | 10 |
| 2020 | From Automatic Keyword Detection to Ontology-Based Topic Modeling
Marc Beck, Syed Tahseen Raza Rizvi, Andreas Dengel 0001, Sheraz Ahmed |
DAS | 4 |
| 2020 | Enhancer-DSNet: A Supervisedly Prepared Enriched Sequence Representation for the Identification of Enhancers and Their Strength
Muhammad Nabeel Asim, Muhammad Ali Ibrahim, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (3) | 5 |
| 2020 | DeepEquaL: Deep Learning Based Mathematical Equation to Latex Generation
Ghaith Bilbeisi, Sheraz Ahmed, Rupak Majumdar |
ICONIP (5) | 2 |
| 2020 | Improving Personal Health Mention Detection on Twitter Using Permutation Based Word Representation Learning
Pervaiz Iqbal Khan, Muhammad Imran Razzak, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (1) | 4 |
| 2020 | Explaining AI-Based Decision Support Systems Using Concept Localization MapsabstractHuman-centric explainability of AI-based Decision Support Systems (DSS) using visual input modalities is directly related to reliability and practicality of such algorithms. An otherwise accurate and robust DSS might not enjoy trust of experts in critical application areas if it is not able to provide reasonable justification of its predictions. This paper introduces Concept Localization Maps (CLMs), which is a novel approach towards explainable image classifiers employed as DSS. CLMs extend Concept Activation Vectors (CAVs) by locating significant regions corresponding to a learned concept in the latent space of a trained image classifier. They provide qualitative and quantitative assurance of a classifier's ability to learn and focus on similar concepts important for humans during image recognition. To better understand the effectiveness of the proposed method, we generated a new synthetic dataset called Simple Concept DataBase (SCDB) that includes annotations for 10 distinguishable concepts, and made it publicly available. We evaluated our proposed method on SCDB as well as a real-world dataset called CelebA. We achieved localization recall of above 80% for most relevant concepts and average recall above 60% for all concepts using SE-ResNeXt-50 on SCDB. Our results on both datasets show great promise of CLMs for easing acceptance of DSS in practice. Adriano Lucieri, Muhammad Naseer Bajwa, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (4) | 4 |
| 2020 | P2ExNet: Patch-Based Prototype Explanation Network
Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (3) | 3 |
| 2020 | Benchmarking Adversarial Attacks and Defenses for Time-Series Data
Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (3) | 3 |
| 2020 | K-mer Neural Embedding Performance Analysis Using Amino Acid CodonsabstractExponential growth of genome-wide assays of gene expressions and their public access open new horizons for machine learning methodologies to effectively perform genetic analysis. In this work, domain specific pre-train k-mer embeddings of DNA sequences are generated by utilising FastText approach. Sequence co-expression pattern information is embedded into 200 dimensional vectors by training Fasttext model on 317,151 samples of DNA sequences (with k-mers representation). We propose a novel idea to utilize the information of various codons present in amino acids for the evaluation of learned sequence vectors. We employ two diverse techniques to compare the performance of generated task-specific k-mer embeddings with state-of-the-art publicly available generic k-mer embeddings of genome. Firstly, we utilize a dimensionality reduction approach namely PCA to alleviate the dimensions of DNA sequences upto 50 features by preserving almost 85% of sequence features information. Afterwards, TSNE algorithms is used to visualize k-mer embeddings and to make sure whether different codons representing the same amino acid are more closer to each other than the ones representing different amino acids. Secondly, to assess the analogy of k-mer embeddings, generated domain specific k-mer embeddings are compared with state-of-the-art k-mer embeddings by estimating the cosine similarity among those codons vectors which represent same amino acid. Overall, we believe that task-specific distributed representation of k-mers would be useful for DNA methylation and Histone occupancy prediction tasks. Muhammad Nabeel Asim, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
IJCNN | 4 |
| 2020 | G1020: A Benchmark Retinal Fundus Image Dataset for Computer-Aided Glaucoma DetectionabstractScarcity of large publicly available retinal fundus image datasets for automated glaucoma detection has been the bottleneck for successful application of artificial intelligence towards practical Computer-Aided Diagnosis (CAD). A few small datasets that are available for research community usually suffer from impractical image capturing conditions and stringent inclusion criteria. These shortcomings in already limited choice of existing datasets make it challenging to mature a CAD system so that it can perform in real-world environment. In this paper we present a large publicly available retinal fundus image dataset for glaucoma classification called G1020. The dataset is curated by conforming to standard practices in routine ophthalmology and it is expected to serve as standard benchmark dataset for glaucoma detection. This database consists of 1020 high resolution colour fundus images and provides ground truth annotations for glaucoma diagnosis, optic disc and optic cup segmentation, vertical cup-to-disc ratio, size of neuroretinal rim in inferior, superior, nasal and temporal quadrants, and bounding box location for optic disc. We also report baseline results by conducting extensive experiments for automated glaucoma diagnosis and segmentation of optic disc and optic cup. Muhammad Naseer Bajwa, Gur Amrit Pal Singh, Wolfgang Neumeier, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
IJCNN | 6 |
| 2020 | Conceptual Explanations of Neural Network Prediction for Time SeriesabstractDeep neural networks are black boxes by construction. Explanation and interpretation methods therefore are pivotal for a trustworthy application. Existing methods are mostly based on heatmapping and focus on locally determining the relevant input parts triggering the network prediction. However, these methods struggle to uncover global causes. While this is a rare case in the image or NLP modality, it is of high relevance in the time series domain. This paper presents a novel framework, i.e. Conceptual Explanation, designed to evaluate the effect of abstract (local or global) input features on the model behavior. The method is model-agnostic and allows utilizing expert knowledge. On three time series datasets Conceptual Explanation demonstrates its ability to pinpoint the causes inherent to the data to trigger the correct model prediction. Ferdinand Küsters, Peter Schichtel, Sheraz Ahmed, Andreas Dengel 0001 |
IJCNN | 3 |
| 2020 | On Interpretability of Deep Learning based Skin Lesion Classifiers using Concept Activation VectorsabstractDeep learning based medical image classifiers have shown remarkable prowess in various application areas like ophthalmology, dermatology, pathology, and radiology. However, the acceptance of these Computer-Aided Diagnosis (CAD) systems in real clinical setups is severely limited primarily because their decision-making process remains largely obscure. This work aims at elucidating a deep learning based medical image classifier by verifying that the model learns and utilizes similar disease-related concepts as described and employed by dermatologists. We used a well-trained and high performing neural network developed by REasoning for COmplex Data (RECOD) Lab for classification of three skin tumours, i.e. Melanocytic Naevi, Melanoma and Seborrheic Keratosis and performed a detailed analysis on its latent space. Two well established and publicly available skin disease datasets, PH2and derm7pt, are used for experimentation. Human understandable concepts are mapped to RECOD image classification model with the help of Concept Activation Vectors (CAVs), introducing a novel training and significance testing paradigm for CAVs. Our results on an independent evaluation set clearly shows that the classifier learns and encodes human understandable concepts in its latent representation. Additionally, TCAV scores (Testing with CAVs) suggest that the neural network indeed makes use of disease-related concepts in the correct way when making predictions. We anticipate that this work can not only increase confidence of medical practitioners on CAD but also serve as a stepping stone for further development of CAV-based neural network interpretation methods. Adriano Lucieri, Muhammad Naseer Bajwa, Stephan Alexander Braun, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
IJCNN | 6 |
| 2020 | Interpreting Deep Models through the Lens of DataabstractIdentification of input data points relevant for the classifier (i.e. serve as the support vector) has recently spurred the interest of researchers for both interpretability as well as dataset debugging. This paper presents an in-depth analysis of the methods which attempt to identify the influence of these data points on the resulting classifier. To quantify the quality of the influence, we curated a set of experiments where we debugged and pruned the dataset based on the influence information obtained from different methods. To do so, we provided the classifier with mislabeled examples that hampered the overall performance. Since the classifier is a combination of both the data and the model, therefore, it is essential to also analyze these influences for the interpretability of deep learning models. Analysis of the results shows that some interpretability methods can detect mislabels better than using a random approach, however, contrary to the claim of these methods, the sample selection based on the training loss showed a superior performance. Dominique Mercier, Shoaib Ahmed Siddiqui, Andreas Dengel 0001, Sheraz Ahmed |
IJCNN | 4 |
| 2019 | DeepEX: Bridging the Gap Between Knowledge and Data Driven Techniques for Time Series Forecasting
Muhammad Ali Chattha, Shoaib Ahmed Siddiqui, Mohsin Munir, Muhammad Imran Malik, Ludger van Elst, Andreas Dengel 0001, Sheraz Ahmed |
ICANN (2) | 7 |
| 2019 | A Robust Hybrid Approach for Textual Document ClassificationabstractText document classification is an important task for diverse natural language processing based applications. Traditional machine learning approaches mainly focused on reducing dimensionality of textual data to perform classification. This although improved the overall classification accuracy, the classifiers still faced sparsity problem due to lack of better data representation techniques. Deep learning based text document classification, on the other hand, benefitted greatly from the invention of word embeddings that have solved the sparsity problem and researchers focus mainly remained on the development of deep architectures. Deeper architectures, however, learn some redundant features that limit the performance of deep learning based solutions. In this paper, we propose a two stage text document classification methodology which combines traditional feature engineering with automatic feature engineering (using deep learning). The proposed methodology comprises a filter based feature selection (FSE) algorithm followed by a deep convolutional neural network. This methodology is evaluated on the two most commonly used public datasets, i.e., 20 Newsgroups data and BBC news data. Evaluation results reveal that the proposed methodology outperforms the state-of-the-art of both the (traditional) machine learning and deep learning based text document classification methodologies with a significant margin of 7.7% on 20 Newsgroups and 6.6% on BBC news datasets. Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR | 5 |
| 2019 | Two Stream Deep Network for Document Image ClassificationabstractThis paper presents a novel two-stream approach for document image classification. The proposed approach leverages textual and visual modalities to classify document images into ten categories, including letter, memo, news article, etc. In order to alleviate dependency of textual stream on performance of underlying OCR (which is the case with general content based document image classifiers), we utilize a filter based feature-ranking algorithm. This algorithm ranks the features of each class based on their ability to discriminate document images and selects a set of top 'K' features that are retained for further processing. In parallel, the visual stream uses deep CNN models to extract structural features of document images.Finally, textual and visual streams are concatenated together using an average ensembling method. Experimental results reveal that the proposed approach outperforms the state-of-the-art system with a significant margin of 4.5% on publicly available Tobacco-3482 dataset. Muhammad Nabeel Asim, Muhammad Usman Ghani Khan, Muhammad Imran Malik, Khizar Razzaque, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR | 6 |
| 2019 | DeepTabStR: Deep Learning based Table Structure RecognitionabstractThis paper presents a novel method for the analysis of tabular structures in document images using the potential of deformable convolutional networks. In order to assess the suitability of the model to the task of table structure recognition, most of the prior methods have been tested on the smaller ICDAR-13 table structure recognition dataset comprising of just 156 tables. We curated a new image-based table structure recognition dataset, TabStructDB2, comprising of 1081 tables densely labeled with row and column information. Instead of collecting new images for this purpose, we leveraged the famous Page-Object Detection dataset from ICDAR-17, and added structural information for all the tabular regions present in the dataset. This new publicly available dataset will enable the development of more sophisticated table structure recognition techniques in the future. We performed extensive evaluation on the two datasets (ICDAR-13 and TabStructDB) including cross-dataset testing in order to evaluate the efficacy of the proposed approach. We achieved state-of-the-art results with deformable models on ICDAR-13 with an average F-Measure of 92.98% (89.42% for rows and 96.55% for columns) and report baseline results on TabStructDB for guiding future research efforts with an F-Measure of 93.72% (91.26% for rows and 95.59% for columns). Despite promising results, structural analysis of tables with arbitrary layouts is still far from achievable at this point. Shoaib Ahmed Siddiqui, Imran Ali Fateh, Syed Tahseen Raza Rizvi, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR | 5 |
| 2019 | Rethinking Semantic Segmentation for Table Structure Recognition in DocumentsabstractBased on the recent advancements in the domain of semantic segmentation, Fully-Convolutional Networks (FCN) have been successfully applied for the task of table structure recognition in the past. We analyze the efficacy of semantic segmentation networks for this purpose and simplify the problem by proposing prediction tiling based on the consistency assumption which holds for tabular structures. For an image of dimensions H × W, we predict a single column for the rows (ŷrowϵ H) and a predict a single row for the columns (ŷrowϵ W). We use a dual-headed architecture where initial feature maps (from the encoder-decoder model) are shared while the last two layers generate class specific (row/column) predictions. This allows us to generate predictions using a single model for both rows and columns simultaneously, where previous methods relied on two separate models for inference. With the proposed method, we were able to achieve state-of-the-art results on ICDAR-13 image-based table structure recognition dataset with an average F-Measure of 92.39% (91.90% and 92.88% F-Measure for rows and columns respectively). With the proposed method, we were able to achieve state-of-the-art results on ICDAR-13. The obtained results advocate that constraining the problem space in the case of FCN by imposing valid constraints can lead to significant performance gains. Shoaib Ahmed Siddiqui, Pervaiz Iqbal Khan, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR | 4 |
| 2019 | A Comparative Analysis of Traditional and Deep Learning-Based Anomaly Detection Methods for Streaming DataabstractWith the Internet of Things (IoT) devices becoming an integral part of human life, the need for robust anomaly detection in streaming data has also been elevated. Dozens of distance-based, density-based, kernel-based, and cluster-based algorithms have been proposed in the area of anomaly detection. Recently, because of the robustness of the deep neural networks (DNN), different deep learning-based anomaly detection methods have also been proposed. With all these rapid developments, there exists a small number of comparative studies for anomaly detection methods. Even in those studies, the comparison is done only in typical anomaly detection settings without taking the streaming data into consideration. The presence of intrinsic time-series characteristics like trend, seasonality, and change-point makes it important to study the behavior of commonly used anomaly detection methods on streaming data. Moreover, the comparison of traditional methods with deep learning-based methods also brings exciting insights about the data which are generally overlooked by traditional methods. In this study, we compare 13 anomaly detection methods on two commonly used streaming data sets. We used four different evaluation metrics to evaluate the methods from different perspectives. Our analysis reveals that the deep learning-based anomaly detection methods are superior to traditional anomaly detection methods. Mohsin Munir, Muhammad Ali Chattha, Andreas Dengel 0001, Sheraz Ahmed |
ICMLA | 4 |
| 2018 | Evaluating similarity measures for gaze patterns in the context of representational competence in physics educationabstractThe competent handling of representations is required for understanding physics' concepts, developing problem-solving skills, and achieving scientific expertise. Using eye-tracking methodology, we present the contributions of this paper as follows: We first investigated the preferences of students with the different levels of knowledge; experts, intermediates, and novices, in representational competence in the domain of physics problem-solving. It reveals that experts more likely prefer to use vector than other representations. Besides, a similar tendency of table representation usage was observed in all groups. Also, diagram representation has been used less than others. Secondly, we evaluated three similarity measures; Levenshtein distance, transition entropy, and Jensen-Shannon divergence. Conducting Recursive Feature Elimination technique suggests Jensen-Shannon divergence is the best discriminating feature among the three. However, investigation on mutual dependency of the features implies transition entropy mutually links between two other features where it has mutual information with Levenshtein distance (Maximal Information Coefficient = 0.44) and has a correlation with Jensen-Shannon divergence (r(18313) = 0.70, p < .001). Seyyed Saleh Mozafari Chanijani, Pascal Klein 0002, Jouni Viiri, Sheraz Ahmed, Jochen Kuhn, Andreas Dengel 0001 |
ETRA | 4 |
| 2018 | Hierarchical Model for Zero-shot Activity Recognition using Wearable Sensors
Mohammad Al-Naser, Hiroki Ohashi, Sheraz Ahmed, Katsuyuki Nakamura, Takayuki Akiyama, Takuto Sato, Phong Xuan Nguyen, Andreas Dengel 0001 |
ICAART (2) | 3 |
| 2018 | Towards a Digital Personal Trainer for Health Clubs - Sport Exercise Recognition Using Personalized Models and Deep Learning
Sebastian Baumbach, Arun Bhatt, Sheraz Ahmed, Andreas Dengel 0001 |
ICAART (2) | 3 |
| 2018 | SentiCite - An Approach for Publication Sentiment AnalysisabstractWith the rapid growth in the number of scientific publications, year after year, it is becoming increasingly difficult to identify quality authoritative work on a single topic. Though there is an availability of scientometric measures which promise to offer a solution to this problem, these measures are mostly quantitative and rely, for instance, only on the number of times an article is cited. With this approach, it becomes irrelevant if an article is cited 10 times in a positive, negative or neutral way. In this context, it is quite important to study the qualitative aspect of a citation to understand its significance. This paper presents a novel system for sentiment analysis of citations in scientific documents (SentiCite) and is also capable of detecting nature of citations by targeting the motivation behind a citation, e.g., reference to a dataset, reading reference. Furthermore, the paper also presents two datasets (SentiCiteDB and IntentCiteDB) containing about 2,600 citations with their ground truth for sentiment and nature of citation. SentiCite along with other state-of-the-art methods for sentiment analysis are evaluated on the presented datasets. Evaluation results reveal that SentiCite outperforms state-of-the-art methods for sentiment analysis in scientific publications by achieving a F1-measure of 0.71. Dominique Mercier, Akansha Bhardwaj, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (2) | 4 |
| 2018 | Ontology-based Information Extraction from Technical Documents
Syed Tahseen Raza Rizvi, Dominique Mercier, Stefan Agne, Steffen Erkel, Andreas Dengel 0001, Sheraz Ahmed |
ICAART (2) | 6 |
| 2017 | Cutting the Error by Half: Investigation of Very Deep CNN and Advanced Training Strategies for Document Image ClassificationabstractWe present an exhaustive investigation of recent Deep Learning architectures, algorithms, and strategies for the task of document image classification to finally reduce the error by more than half. Existing approaches, such as the DeepDoc-Classifier, apply standard Convolutional Network architectures with transfer learning from the object recognition domain. The contribution of the paper is threefold: First, it investigates recently introduced very deep neural network architectures (GoogLeNet, VGG, ResNet) using transfer learning (from real images). Second, it proposes transfer learning from a huge set of document images, i.e. 400; 000 documents. Third, it analyzes the impact of the amount of training data (document images) and other parameters to the classification abilities. We use two datasets, the Tobacco-3482 and the large-scale RVL-CDIP dataset. We achieve an accuracy of 91:13% for the Tobacco-3482 dataset while earlier approaches reach only 77:6%. Thus, a relative error reduction of more than 60% is achieved. For the large dataset RVL-CDIP, an accuracy of 90:97% is achieved, corresponding to a relative error reduction of 11:5%. Muhammad Zeshan Afzal, Andreas Kölsch, Sheraz Ahmed, Marcus Liwicki |
ICDAR | 3 |
| 2017 | Academic Community Explorer (ACE) for Syntactic, Semantic and Pragmatic Document AnalysisabstractThis paper presents a novel Academic Community Explorer (ACE) which performs syntactic, semantic and pragmatic document analysis of scientific publications. Firstly, ACE uses syntactic structure to extract relevant information from a scientific document. Secondly, semantic analysis is performed to derive an article based co-authorship and citation network. Finally, ACE uses these document based networks to build a complete community network for pragmatic analysis. Furthermore, scientometric analysis is performed to extract the pragmatics by analyzing authors and publication community networks through micro and macro indicators. Two novel micro indicators Senti-Index, reflecting the sentiment present in citations and, Overlap index, reflecting community behavior have been introduced. This is a step in the direction of automatic qualitative assessment of scientific documents. In addition, ACE provides a rich visualization interface which helps in exploratory analysis of the community to identify hidden patterns, e.g, isolated small groups in the community which collaborate and cite each other frequently. A feasibility study is performed on the corpus of ICDAR publications from 1993-2015 to show the insights and benefits of the ACE framework. The results reveals that ICDAR is a highly collaborative community which has most likely arrived at its 'phase transition' stage with 70% of the community closely connected to each other. Akansha Bhardwaj, Dominique Mercier, Hisham Hashmi, Sheraz Ahmed, Andreas Dengel 0001 |
ICDAR | 4 |
| 2017 | AirScript - Creating Documents in AirabstractThis paper presents a novel approach, called AirScript, for creating, recognizing and visualizing documents in air. We present a novel algorithm, called 2-DifViz, that converts the hand movements in air (captured by a Myo-armband worn by a user) into a sequence of x, y coordinates on a 2D Cartesian plane, and visualizes them on a canvas. Existing sensor-based approaches either do not provide visual feedback or represent the recognized characters using prefixed templates. In contrast, AirScript stands out by giving freedom of movement to the user, as well as by providing a real-time visual feedback of the written characters, making the interaction natural. AirScript provides a recognition module to predict the content of the document created in air. To do so, we present a novel approach based on deep learning, which uses the sensor data and the visualizations created by 2-DifViz. The recognition module consists of a Convolutional Neural Network (CNN). and two Gated Recurrent Unit (GRU) Networks. The output from these three networks is fused to get the final prediction about the characters written in air. AirScript can be used in highly sophisticated environments like a smart classroom, a smart factory or a smart laboratory, where it would enable people to annotate pieces of texts wherever they want without any reference surface. We have evaluated AirScript against various well-known learning models (HMM, KNN, SVM, etc.) on the data of 12 participants. Evaluation results show that the recognition module of AirScript largely outperforms all of these models by achieving an accuracy of 91.7% in a person independent evaluation and a 96.7% accuracy in a person dependent evaluation. Ayushman Dash, Amit Sahu, Rajveer Shringi, John Cristian Borges Gamboa, Muhammad Zeshan Afzal, Muhammad Imran Malik, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR | 8 |
| 2017 | DeepDeSRT: Deep Learning for Detection and Structure Recognition of Tables in Document ImagesabstractThis paper presents a novel end-to-end system for table understanding in document images called DeepDeSRT. In particular, the contribution of DeepDeSRT is two-fold. First, it presents a deep learning-based solution for table detection in document images. Secondly, it proposes a novel deep learning-based approach for table structure recognition, i.e. identifying rows, columns, and cell positions in the detected tables. In contrast to existing rule-based methods, which rely on heuristics or additional PDF metadata (like, for example, print instructions, character bounding boxes, or line segments), the presented system is data-driven and does not need any heuristics or metadata to detect as well as to recognize tabular structures in document images. Furthermore, in contrast to most existing table detection and structure recognition methods, which are applicable only to PDFs, DeepDeSRT processes document images, which makes it equally suitable for born-digital PDFs (as they can automatically be converted into images) as well as even harder problems, e.g. scanned documents. To gauge the performance of DeepDeSRT, the system is evaluated on the publicly available ICDAR 2013 table competition dataset containing 67 documents with 238 pages overall. Evaluation results reveal that DeepDeSRT outperforms state-of-the-art methods for table detection and structure recognition and achieves F1-measures of 96.77% and 91.44% for table detection and structure recognition, respectively. Additionally, DeepDeSRT is evaluated on a closed dataset from a real use case of a major European aviation company comprising documents which are highly unlike those in ICDAR 2013. Tested on a randomly selected sample from this dataset, DeepDeSRT achieves high detection accuracy for tables which demonstrates the sound generalization capabilities of our system. Sebastian Schreiber 0001, Stefan Agne, Ivo Wolf, Andreas Dengel 0001, Sheraz Ahmed |
ICDAR | 5 |
| 2017 | D-StaR: A Generic Method for Stamp Segmentation from Document ImagesabstractB This paper presents a novel approach, named D-StaR, for stamp segmentation from scanned document images. The presented approach is generic (applicable to stamps of any color, shape, size, and orientation) and based on deep learning. In particular, it uses Fully Convolutional networks for semantic analysis of documents to extract stamps. The presented approach is evaluated on a publicly available stamp dataset. Evaluation results show that the presented approach outperforms the state-of-the-art methods for stamp segmentation and achieves pixel based precision and recall of 87% and 84%, respectively. Deeper analysis of the evaluation reveals that the presented approach can segment both overlapping and non-overlapping stamps, which was always a problem for existing systems in the literature. Junaid Younas, Muhammad Zeshan Afzal, Muhammad Imran Malik, Faisal Shafait, Paul Lukowicz, Sheraz Ahmed |
ICDAR | 6 |
| 2017 | DeepBIBX: Deep Learning for Image Based Bibliographic Data Extraction
Akansha Bhardwaj, Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed |
ICONIP (2) | 4 |
| 2016 | Automatic Signature Segmentation Using Hyper-Spectral ImagingabstractIn this paper, we propose a method for automatic signature segmentation using hyper-spectral imaging. The proposed method first uses the connected component analysis and local features to segment the printed text and signatures. Secondly, it uses spectral response of text, signature, and background to extract signature pixels. The proposed method is robust, and remains unaffected by color and intensity of the ink, and by any structural information of the text, as the classification relies exclusively on the spectral response of the document. The proposed method can extract signature pixels either overlapping or non-overlapping from different backgrounds like, logos, tables, stamps, and printed text. We used high-resolution hyper-spectral imaging to study and classify 300 documents with varying backgrounds. We evaluated the proposed classification method and compared results with the state-of-the art system. The proposed method outperformed the state-of-the-art system and achieved 100% precision and 84% recall. Umair Muneer Butt, Sheraz Ahmed, Faisal Shafait, Christian Nansen, Ajmal Mian, Muhammad Imran Malik |
ICFHR | 2 |
| 2015 | ICDAR2015 competition on signature verification and writer identification for on- and off-line skilled forgeries (SigWIcomp2015)abstractThis paper presents the results of the ICDAR 2015 competition on signature verification and writer identification for on- and off-line skilled forgeries jointly organized by PR-researchers and Forensic Handwriting Examiners (FHEs). The aim is to bridge the gap between recent technological developments and forensic casework. Two modalities (signatures and handwritten text) are considered and training and evaluation data are collected and provided by FHEs and PR-researchers. Four tasks are defined for four different languages; Bengali off-line signature verification, Italian off-line signature verification, German on-line signature verification, and English handwritten text based writer identification. In total, 40 systems have participated in this competition. The participants of the signatures modality were motivated to report their results in Likelihood Ratios (LRs). This has made the systems even more interesting for application in forensic casework. For evaluating the performance of the systems, we have used the forensically substantial Cost of Log Likelihood Ratios (Ĉllr) in the case of signatures, and the F-measure in the case of handwritten text. Muhammad Imran Malik, Sheraz Ahmed, Angelo Marcelli, Umapada Pal 0001, Michael Blumenstein, Linda Alewijnse, Marcus Liwicki |
ICDAR | 2 |
| 2014 | Statistical segmentation and structural recognition for floor plan interpretation - Notation invariant structural element recognition
Lluís-Pere de las Heras, Sheraz Ahmed, Marcus Liwicki, Ernest Valveny, Gemma Sánchez |
Int. J. Document Anal. Recognit. | 2 |
| 2014 | Automatic analysis and sketch-based retrieval of architectural floor plans
Sheraz Ahmed, Marcus Liwicki, Christoph Langenhan, Andreas Dengel 0001, Frank Petzold |
Pattern Recognit. Lett. | 1 |
| 2013 | Automatic Ground Truth Generation of Camera Captured Documents Using Document Image RetrievalabstractIn this paper a novel method for automatic ground truth generation of camera captured document images is proposed. Currently, no dataset is available for camera captured documents. It is very difficult to build these datasets manually, as it is very laborious and costly. The proposed method is fully automatic, allowing building the very large scale (i.e., millions of images) labeled camera captured documents dataset, without any human intervention. Evaluation of samples generated by the proposed approach shows that 99.98% of the images are correctly labeled. Novelty of the proposed approach lies in the use of document image retrieval for automatic labeling, especially for camera captured documents, which contain different distortions specific to camera, e.g., blur, occlusion, perspective distortion, etc. Sheraz Ahmed, Koichi Kise, Masakazu Iwamura, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 1 |
| 2013 | A Generic Method for Stamp Segmentation Using Part-Based FeaturesabstractTraditionally, stamps are considered as a seal of authenticity for documents. For automatic processing and verification, segmentation of stamps from documents is pivotal. Existing methods for stamp extraction mostly employ color and/or shape based techniques, thereby limiting their applicability to only colored and specific shape stamps. In this paper, a novel, generic method based on part-based features is presented for segmentation of stamps from document images. The proposed method can segment black, colored, unseen, arbitrary shaped, textual, as well as graphical stamps. The proposed method is evaluated on a publicly available dataset for stamp detection and verification and achieved recall and precision of 73% and 83% respectively, for black stamps which were not addressed in the past. Sheraz Ahmed, Faisal Shafait, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 1 |
| 2013 | FREAK for Real Time Forensic Signature VerificationabstractThis paper presents a novel signature verification system based on local features of signatures. The proposed system uses Fast Retina Key points (FREAK) which represent local features and are inspired by the human visual system, particularly the retina. To locate local points of interest in signatures, two local key point detectors, i.e., Features from Accelerated Segment Test (FAST) and Speeded-up Robust Features (SURF), have been used and their performance comparison in terms of Equal Error Rate (EER) and time is presented. The proposed system has been evaluated on publicly available dataset of forensic signature verification competition, 4NSigComp2010, which contains genuine, forged, and disguised signatures. The proposed system achieved an EER of 30%, which is considerably very low when compared against all the participants of the said competition. In addition to EER, the proposed system requires only 0.6 seconds on average to verify a 3000*1500 scanned signature. This shows that the proposed system has a potential and suitability for forensic signature verification as well as real time applications. Muhammad Imran Malik, Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 2 |
| 2012 | Extraction of Text Touching Graphics Using SURFabstractIn this paper we propose a novel part-based method for the extraction of text touching graphic components. The Speeded Up Robust Features (SURF) are used to localize the text components and distinguish them from graphics. We introduce several post-processing steps to finally detect the text. We have tested our method on a publicly available data set of architectural floor plans and on real geographical maps. On floor plans we have located more than 95% of the text components which were not identified as text beforehand because they were touching graphic components. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
Document Analysis Systems | 1 |
| 2012 | Automatic Room Detection and Room Labeling from Architectural Floor PlansabstractThis paper presents an automatic system for analyzing and labeling architectural floor plans. In order to detect the locations of the rooms, the proposed systems extracts both, structural and semantic information from given floor plans. Furthermore, OCR is applied on the text layer to retrieve the meaningful room labeling. Finally, a novel post-processing is proposed to split rooms into several sub-regions if several semantic rooms share the same physical room. Our fully automatic system is evaluated on a publicly available dataset of architectural floor plans. In our experiments, we could clearly outperform other state-of-the-art approaches for room detection. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
Document Analysis Systems | 1 |
| 2012 | A Signature Verification Framework for Digital Pen ApplicationsabstractIn this paper we present a framework for real-time online signature verification scenarios. The proposed framework is based on state-of-the-art feature extraction and Gaussian Mixture Model (GMM) classification. While our signature verification library is generally applicable to any input device using digital pens, we have implemented verification scenarios using the Anoto digital pen. As such our automated signature verification framework becomes an interesting commodity for industry, because the Anoto SDK is easy to apply and the GMM-based classification can be seamlessly integrated. The novelty of this work is the application of our framework that takes real-time online signature verification to every scenario where digital pens may potentially be used. In this paper we describe several scenarios where our framework has been applied, including signatures in financial contracts or ordering processes. We also propose a general approach to integrate the GMM-descriptions into electronic ID-cards in order to also store behavioral biometrics on these cards. In experiments we have measured the performance of the signature verification system when skilled forgeries were present. The interest shown by our partner financial institutions and the results of our initial evaluations indicate that our signature verification framework suits exactly the demands of our clients. Muhammad Imran Malik, Sheraz Ahmed, Andreas Dengel 0001, Marcus Liwicki |
Document Analysis Systems | 2 |
| 2012 | Signature Segmentation from Document ImagesabstractIn this paper we propose a novel method for the extraction of signatures from document images. Instead of using a human defined set of features a part-based feature extraction method is used. In particular, we use the Speeded Up Robust Features (SURF) to distinguish the machine printed text from signatures. Using SURF features makes the approach generally more useful and reliable for different resolution documents. We have evaluated our system on the publicly available Tobacco-800 dataset in order to compare it to previous work. Finally, all signatures were found in the images and less than half of the found signatures are false positives. Therefore, our system can be applied for practical use. Sheraz Ahmed, Muhammad Imran Malik, Marcus Liwicki, Andreas Dengel 0001 |
ICFHR | 1 |
| 2011 | Improved Automatic Analysis of Architectural Floor PlansabstractThis paper proposes a novel complete system for automated floor plan analysis. Besides applying and improving state-of-the-art processing methods, we introduce novel preprocessing methods, e.g., the differentiation between thick, medium, and thin lines and the removal of components outside the convex hull of the outer walls. Especially the latter method increases the performance of the final system. In our experiments on a reference data set we compare our approach to other approaches available in the literature. We show that our system outperforms previous systems. The final room recognition accuracy is 79% that is 10% higher than the 69% achieved by a state-of-the-art approach from the literature. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 1 |
| 2011 | Text/Graphics Segmentation in Architectural Floor PlansabstractIn this paper, we propose an improved method for text/graphics segmentation. Text/graphics separation is a crucial preprocessing step in document analysis before further analysis and recognition can be applied. Our proposed system extends the method of Tombre et al. with a number of improvements to make it more suitable for architectural floor plans. A crucial novel preprocessing step is the detection and removal of walls before the actual segmentation. Furthermore, text components are then extracted by analyzing connected components and even considering text overlapping with graphics. Finally, a smearing approach is used to remove noise and extract the final text components. Evaluation results over the series of 90 floor plans which has also been used in reference work shows that our method has a recall of almost 99% and a precision greater then 97%. Sheraz Ahmed, Marcus Liwicki, Andreas Dengel 0001 |
ICDAR | 1 |