VLDB 2026 Research / reviewers in the wild / expert
Andreas K. Maier
dblp:131/7133 · also Andreas Maier 0001
· DBLP profile ↗
235ranked-venue papers
6as first author
120since 2021 · last 2026
0000-0002-9550-5284ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 138 · 6 first-author · 65 since 2021Applied, interdisciplinary, general and emerging computing · 115 · 54 since 2021Artificial intelligence and machine learning · 76 · 5 first-author · 49 since 2021Databases, data management, data science and information retrieval · 13 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
Uddipan Basu Bir, Vincent Christlein, Andreas K. Maier, Mathias Zinnen |
ICDAR (3) | 3 |
| 2026 | A speech-to-video synthesis approach using spatio-temporal diffusion for vocal tract MRI
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Fangxu Xing, Xiaofeng Liu 0001, Maureen Stone 0001, Jiachen Zhuo, Juan Rafael Orozco-Arroyave, Elmar Nöth, Jana Hutter, Jerry L. Prince, Andreas K. Maier, Jonghye Woo |
Medical Image Anal. | 11 |
| 2026 | Comparison Study: Glacier Calving Front Delineation in Synthetic Aperture Radar Images With Deep LearningabstractContinuous monitoring of glacier calving fronts is essential for sea level rise projections. This study benchmarks Deep Learning systems for front delineation in Synthetic Aperture Radar imagery. While Deep Learning systems exhibit errors up to 221 m, human annotators deviate by only 38 m, underscoring the need for further research. Nora Gourmelon, Konrad Heidler, Erik Loebel, Daniel Cheng, Julian Klink, Anda Dong, Fei Wu 0025, Noah Maul, Moritz Koch, Marcel Dreier, Dakota Pyles, Thorsten Seehaus, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Pattern Anal. Mach. Intell. | 14 |
| 2026 | Vision transformer Hook for dense predictionsabstractPre-trained vision transformers (ViTs) have demonstrated remarkable capability in learning semantically rich image representations. However, their underlying plain architectures yield low-resolution feature maps, lacking essential fine-grained spatial details required for dense prediction tasks. To better transfer the learned visual features, we present ViT-Hook, a novel hybrid backbone compatible with plain ViTs that effectively bridges the gap between global semantic understanding and local spatial encodings. Specifically, our method aims to broaden the scope and impact of ViT from the following perspectives: (1) We propose a simple transformer-decoder-inspired hook module that receives hierarchical CNN features as spatial queries and interacts with expressive ViT features from large-scale pre-training, therefore instantiating general-purpose representations into task-suited ones. (2) ViT-Hook is a plug-and-play solution for powerful vision foundation models, such as DINOv2 and RADIO. In this case, we find that only partially fine-tuning several intermediate ViT layers can outperform previous full fine-tuning methods, while substantially reducing compute and memory burdens with most parameters frozen. (3) We evaluate ViT-Hook with various pre-trained sources on multiple dense prediction tasks, including semantic segmentation, instance segmentation, and object detection. Notably, tested on the unified UperNet and Mask R-CNN frameworks, our ViT-Hook surpasses state-of-the-art by a large margin, achieving 59.7 (+4.7) mIoU on ADE20K val, 55.0 (+3.6) box AP and 48.5 (+3.3) mask AP on COCO val2017. • We propose ViT-Hook, a hybrid backbone that effectively enhances Vision Transformer performance on various dense prediction tasks. • The proposed spatial query and hook modules are lightweight yet powerful, achieving competitive results compared to SoTA on widely used benchmarks. • We introduce a novel partial fine-tuning strategy, which outperforms full fine-tuning while using only a fraction of compute and memory. • We validate the generalizability of ViT-Hook on multiple types of upstream pre-training methods, including the most recent vision foundation models. Siyuan Mei, Mareike Thies, Yan Xia 0002, Yipeng Sun, Fei Wu 0025, Fuxin Fan, Mingxuan Gu, Chengze Ye, Yixing Huang, Vincent Christlein, Andreas K. Maier |
Pattern Recognit. | 11 |
| 2025 | Same Semantics of the Signal - What Do We Cluster with what RepresentationabstractSemantic clustering of bioacoustic signals is crucial for a deeper understanding of intra-class differences. This is particularly important for understanding killer whale signals, as their vocalizations are learned behaviors and determination of matrilineal-specific dialects is reliant upon subtle differences within instances which may be characterized into a single larger category. Aspects of data collection may have an effect on how these calls are grouped, and it is therefore necessary to understand what the focus of the feature generation algorithm is. This study addresses the impact of factors such as recording conditions and environment, together with its respective relative noise levels, by first analyzing two different deep learning and data-driven feature representations, either derived by an undercomplete autoencoder or a supervised call type classifier. These are then compared with representations generated by two state-of-the-art transformer-based tools, namely HuBERT and Wav2Vec2. Alexander Barnhill, Oliver Traub, Andreas K. Maier, Elmar Nöth, Christian Bergler |
ICASSP | 3 |
| 2025 | A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power GridsabstractThe integration of renewable energy sources into the electrical grid introduces complex challenges in fault detection and coordination of grid recovery mechanisms. Traditional relay protection systems, which operate based on static rules and predefined thresholds, are inadequate for addressing these challenges, particularly in detecting and isolating faults such as short circuits. Consequently, the conventional methodologies applied to electrical network protection frequently fail to achieve optimal performance in fault detection, especially in terms of adherence to safety standards and the selective limitation of damage. Recent research indicates that machine learning (ML)-based approaches can effectively tackle these issues; however, variations in grid configurations and analysis windows have impeded consistent comparative assessments. In this study, we assess the efficacy of various ML models in detecting electrical faults and pinpointing defective transmission lines within a 10 ms measurement interval—a critical time-frame for real-time operational viability, for the first time. The most effective model attained an F1 score of 0.991±0.018 and demonstrated a processing time of 0.342ms±0.509ms. Julian Oelhaf, Georg Kordowich, Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Andreas K. Maier, Johann Jaeger, Siming Bayer |
ICASSP | 5 |
| 2025 | Refusal Behavior in Large Language Models: A Nonlinear PerspectiveabstractRefusal behavior in large language models (LLMs) enables them to decline responding to harmful, unethical, or inappropriate prompts, ensuring alignment with ethical standards. This paper investigates refusal behavior across six LLMs from three architectural families. We challenge the assumption of refusal as a linear phenomenon by employing dimensionality reduction techniques, including PCA, t-SNE, and UMAP. Our results reveal that refusal mechanisms exhibit nonlinear, multidimensional characteristics that vary by model architecture and layer. These findings highlight the need for nonlinear interpretability to improve alignment research and inform safer AI deployment strategies. Fabian Hildebrandt, Andreas K. Maier, Patrick Krauss, Achim Schilling |
IJCNN | 2 |
| 2025 | Exemplar Med-DETR: Toward Generalized and Robust Lesion Detection in Mammogram Images and Beyond
Sheethal Bhat, Bogdan Georgescu, Adarsh Bhandary Panambur, Mathias Zinnen, Tri-Thien Nguyen, Awais Mansoor, Karim Khalifa Elbarbary, Siming Bayer, Florin C. Ghesu, Sasa Grbic, Andreas K. Maier |
MICCAI (6) | 11 |
| 2025 | CXR-CML: Improved Zero-Shot Classification of Long-Tailed Multi-label Diseases in Chest X-Rays
Rajesh Madhipati, Sheethal Bhat, Lukas Buess, Andreas K. Maier |
MICCAI (7) | 4 |
| 2025 | DINO Adapted to X-Ray (DAX): Foundation Models for Intraoperative X-Ray Imaging
Joshua Scheuplein, Maximilian Rohleder, Andreas K. Maier, Björn W. Kreher |
MICCAI (10) | 3 |
| 2025 | FIND-Net - Fourier-Integrated Network with Dictionary Kernels for Metal Artifact Reduction
Farid Tasharofi, Fuxin Fan, Melika Qahqaie, Mareike Thies, Andreas K. Maier |
MICCAI (13) | 5 |
| 2025 | Pancreatic duct centerline extraction for image unfolding in photon-counting CTabstractPancreatic diseases are often only diagnosed at a late stage, and pancreatic cancer is the most feared due to a very high mortality. Abnormalities of the main pancreatic duct, such as blockages and dilatation, are often (early) signs of such pancreatic diseases, but are difficult to detect in standard Computed Tomography image series. Photon-Counting Computed Tomography with its higher resolution improves the detectability of this duct, allowing diagnostic assessment. A comprehensive visualization in a single view requires a centerline-based unfolding of the duct and pancreas. However, manual centerline annotation is tedious. To automate this process, we introduce a fully automated pipeline for pancreatic duct unfolding by robustly extracting the centerline using Dijkstra’s algorithm on a cost map derived from a segmentation probability map. The core contribution of this work lies in the processing of the data-driven cost map leading to a consistent centerline for generating CPR visualizations of the pancreas. To improve individual steps within the pipeline, we investigate further enhancements such as segmentation filtering and the topology-preserving skeleton recall loss. In the evaluation, we assess performance of our method on both ultra-high-resolution and regular PCCT images. We find that the centerline can be consistently extracted from both scan types, where the centerlines from the ultra-high resolution images exhibit a slightly lower median error of 0.58 mm compared to the 0.73 mm using the regular resolution. Jie Yi Tan, Leonhard Rist, Abraham Ayala Hernandez, Michael Sühling, Erik Gudman Steuble Brandt, Andreas K. Maier, Oliver Taubmann |
Comput. Graph. | 6 |
| 2025 | Unsupervised motion artifacts reduction for cone-beam CT via enhanced landmark detection
Thanaporn Viriyasaranon, Serie Ma, Mareike Thies, Andreas K. Maier, Jang Hwan Choi 0001 |
Expert Syst. Appl. | 4 |
| 2025 | Zero-Shot Paragraph-level Handwriting Imitation with Latent Diffusion ModelsabstractAbstract The imitation of cursive handwriting is mainly limited to generating handwritten words or lines. Multiple synthetic outputs must be stitched together to create paragraphs or whole pages, whereby consistency and layout information are lost. To close this gap, we propose a method for imitating handwriting at the paragraph level that also works for unseen writing styles. Therefore, we introduce a modified latent diffusion model that enriches the encoder-decoder mechanism with specialized loss functions that explicitly preserve the style and content. We enhance the attention mechanism of the diffusion model with adaptive 2D positional encoding and the conditioning mechanism to work with two modalities simultaneously: a style image and the target text. This significantly improves the realism of the generated handwriting. We set a new benchmark in our comprehensive evaluation, achieving 61 % mAP and 56 % top-1 accuracy in style preservation, significantly outperforming the previous best method (37 % mAP, 30 % top-1). We are making our code publicly available for reproducibility, supporting research in this area and research into potential countermeasures: https://github.com/M4rt1nM4yr/paragraph_handwriting_imitation_ldm Martin Mayr, Marcel Dreier, Florian Kordon, Mathias Seuret, Jochen Zöllner, Fei Wu 0025, Andreas K. Maier, Vincent Christlein |
Int. J. Comput. Vis. | 7 |
| 2025 | Lightweight cross-attention-based HookNet for historical handwritten document layout analysis
Fei Wu 0025, Mathias Seuret, Martin Mayr, Florian Kordon, Jochen Zöllner, Sebastian Wind, Andreas K. Maier, Vincent Christlein |
Int. J. Document Anal. Recognit. | 7 |
| 2025 | Multi-modal cognitive maps for language and vision based on neural successor representationsabstractCognitive maps are a proposed concept on how the brain efficiently organizes memories and retrieves context out of them. The entorhinal-hippocampal complex is heavily involved in episodic and relational memory processing, as well as spatial navigation and is thought to built cognitive maps via place and grid cells. To make use of the promising properties of cognitive maps, we set up a multi-modal neural network using successor representations which are able to model place cell dynamics and cognitive map representations. Here, we use multi-modal inputs consisting of images and word embeddings. The network learns the similarities between novel inputs and the training database and therefore the representation of the cognitive map successfully. Subsequently, the prediction of the network can be used to infer from one modality to another with over 90% accuracy. The proposed method could therefore be a building block to improve current AI systems for better understanding of the environment and the different modalities in which objects appear. The association of specific modalities with certain encounters can therefore lead to context awareness in novel situations when similar encounters with less information occur and additional information can be inferred from the learned cognitive map. Cognitive maps, as represented by the entorhinal-hippocampal complex in the brain, organize and retrieve context from memories, suggesting that large language models (LLMs) like ChatGPT could harness similar architectures to function as a high-level processing center, akin to how the hippocampus operates within the cortex hierarchy. Finally, by utilizing multi-modal inputs, LLMs can potentially bridge the gap between different forms of data (like images and words), paving the way for context-awareness and grounding of abstract concepts through learned associations, addressing the symbol grounding problem in AI. Paul Stoewer, Achim Schilling, Pegah Ramezani, Hassane Kissane, Andreas K. Maier, Patrick Krauss |
Neurocomputing | 5 |
| 2025 | Low-dose computed tomography perceptual image quality assessmentabstractIn computed tomography (CT) imaging, optimizing the balance between radiation dose and image quality is crucial due to the potentially harmful effects of radiation on patients. Although subjective assessments by radiologists are considered the gold standard in medical imaging, these evaluations can be time-consuming and costly. Thus, objective methods, such as the peak signal-to-noise ratio and structural similarity index measure, are often employed as alternatives. However, these metrics, initially developed for natural images, may not fully encapsulate the radiologists' assessment process. Consequently, interest in developing deep learning-based image quality assessment (IQA) methods that more closely align with radiologists' perceptions is growing. A significant barrier to this development has been the absence of open-source datasets and benchmark models specific to CT IQA. Addressing these challenges, we organized the Low-dose Computed Tomography Perceptual Image Quality Assessment Challenge in conjunction with the Medical Image Computing and Computer Assisted Intervention 2023. This event introduced the first open-source CT IQA dataset, consisting of 1,000 CT images of various quality, annotated with radiologists' assessment scores. As a benchmark, this challenge offers a comprehensive analysis of six submitted methods, providing valuable insight into their performance. This paper presents a summary of these methods and insights. This challenge underscores the potential for developing no-reference IQA methods that could exceed the capabilities of full-reference IQA methods, making a significant contribution to the research community with this novel dataset. The dataset is accessible at https://zenodo.org/records/7833096. Wonkyeong Lee, Fabian Wagner, Adrian Galdran, Yongyi Shi, Wenjun Xia, Ge Wang 0001, Xuanqin Mou, Md. Atik Ahamed, Abdullah-Al-Zubaer Imran, Jieun Oh, Kyung Sang Kim, Jong Tak Baek, Dongheon Lee 0002, Boohwi Hong, Philip Tempelman, Donghang Lyu, Adrian Kuiper, Lars van Blokland, Maria Baldeon Calisto, Scott S. Hsieh, Minah Han, Jongduk Baek, Andreas K. Maier, Adam S. Wang, Garry Gold, Jang Hwan Choi 0001 |
Medical Image Anal. | 23 |
| 2025 | Data-efficient handwritten text recognition of diplomatic historical textabstractAbstract Traditional methods in handwritten text recognition primarily focus on generating basic transcriptions, which often fall short for in-depth humanities research. Our study enhances this by providing diplomatic transcriptions for German studies, meticulously reproducing the original manuscripts, including layout and expanded abbreviations. State-of-the-art sequence-to-sequence approaches for handwritten text recognition predominantly use Connectionist Temporal Classification (CTC) as an auxiliary loss of the encoder output to improve robustness and accuracy. This is not possible in this task due to the great differences in the length of diplomatic transcriptions. We propose using the basic transcription instead of the diplomatic one as an additional target for the CTC feedback. Additionally, we introduce positional encoding at the intersection between the encoder and decoder to resolve the conflict of competing encoder objectives, balancing CTC loss reduction with the maintenance of implicit positional encoding for the decoder. Our empirical tests on the newly created dataset “Nuremberg Letterbooks” demonstrate significant data efficiency improvements. With only 4000 training lines (about 130 transcribed pages), we achieve a Character Error Rate (CER) of 9.39% without expanded abbreviations and 12.07% with expanded abbreviations, outperforming the baseline errors of 14.26% and 68.21%, respectively. Martin Mayr, Katharina Neumeier, Julian Krenz, Simon Bürcky, Florian Kordon, Mathias Seuret, Jochen Zöllner, Fei Wu 0025, Andreas K. Maier, Vincent Christlein |
Multim. Tools Appl. | 9 |
| 2025 | Recognizing sensory gestures in historical artworksabstractAbstract The automatic recognition of sensory gestures in artworks provides the opportunity to open up methods of computational humanities to modern paradigms like sensory studies or everyday history. We introduce SensoryArt, a dataset of multisensory gestures in historical artworks, annotated with person boxes, pose estimation key points and gesture labels. We analyze algorithms for each label type and explore their combination for gesture recognition without intermediate supervision. These combined algorithms are evaluated for their ability to recognize and localize depicted persons performing sensory gestures. Our experiments show that direct detection of smell gestures is the most effective method for both detecting and localizing gestures. After applying post-processing, this method outperforms even image-level classification algorithms in image-level classification metrics, despite not being the primary training objective. This work aims to open up the field of sensory history to the computational humanities and provide humanities-based scholars with a solid foundation to complement their methodological toolbox with quantitative methods. Mathias Zinnen, Azhar Hussian, Andreas K. Maier, Vincent Christlein |
Multim. Tools Appl. | 3 |
| 2025 | Nonlinear Neural Dynamics and Classification Accuracy in Reservoir ComputingabstractReservoir computing information processing based on untrained recurrent neural networks with random connections is expected to depend on the nonlinear properties of the neurons and the resulting oscillatory, chaotic, or fixed-point dynamics of the network. However, the degree of nonlinearity required and the range of suitable dynamical regimes for a given task remain poorly understood. To clarify these issues, we study the classification accuracy of a reservoir computer in artificial tasks of varying complexity while tuning both the neuron's degree of nonlinearity and the reservoir's dynamical regime. We find that even with activation functions of extremely reduced nonlinearity, weak recurrent interactions, and small input signals, the reservoir can compute useful representations. These representations, detectable only in higher-order principal components, make complex classification tasks linearly separable for the readout layer. Increasing the recurrent coupling leads to spontaneous dynamical behavior. Nevertheless, some input-related computations can "ride on top" of oscillatory or fixed-point attractors with little loss of accuracy, whereas chaotic dynamics often reduces task performance. By tuning the system through the full range of dynamical phases, we observe in several classification tasks that accuracy peaks at both the oscillatory/chaotic and chaotic/fixed-point phase boundaries, supporting the edge of chaos hypothesis. We also present a regression task with the opposite behavior. Our findings, particularly the robust weakly nonlinear operating regime, may offer new perspectives for both technical and biological neural networks with random connectivity. Claus Metzner, Achim Schilling, Andreas K. Maier, Patrick Krauss |
Neural Comput. | 3 |
| 2025 | SSL4SAR: Self-Supervised Learning for Glacier Calving Front Extraction From SAR ImageryabstractGlaciers are losing ice mass at unprecedented rates, increasing the need for accurate, year-round monitoring to understand frontal ablation, particularly the factors driving the calving process. Deep learning models can extract calving front positions from Synthetic Aperture Radar imagery to track seasonal ice losses at the calving fronts of marine- and lake-terminating glaciers. The current state-of-the-art model relies on ImageNet-pretrained weights. However, they are suboptimal due to the domain shift between the natural images in ImageNet and the specialized characteristics of remote sensing imagery, in particular for Synthetic Aperture Radar imagery. To address this challenge, we propose two novel self-supervised multimodal pretraining techniques that leverage SSL4SAR, a new unlabeled dataset comprising 9,563 Sentinel-1 and 14 Sentinel-2 images of Arctic glaciers, with one optical image per glacier in the dataset. Additionally, we introduce a novel hybrid model architecture that combines a Swin Transformer encoder with a residual Convolutional Neural Network (CNN) decoder. When pretrained on SSL4SAR, this model achieves a mean distance error of 293m on the “CAlving Fronts and where to Find thEm” (CaFFe) benchmark dataset, outperforming the prior best model by 67 m. Evaluating an ensemble of the proposed model on a multi-annotator study of the benchmark dataset reveals a mean distance error of 75 m, approaching the human performance of 38 m. This advancement enables precise monitoring of seasonal changes in glacier calving fronts. Nora Gourmelon, Marcel Dreier, Martin Mayr, Thorsten Seehaus, Dakota Pyles, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | A Gradient-Based Approach to Fast and Accurate Head Motion Compensation in Cone-Beam CTabstractCone-beam computed tomography (CBCT) systems, with their flexibility, present a promising avenue for direct point-of-care medical imaging, particularly in critical scenarios such as acute stroke assessment. However, the integration of CBCT into clinical workflows faces challenges, primarily linked to long scan duration resulting in patient motion during scanning and leading to image quality degradation in the reconstructed volumes. This paper introduces a novel approach to CBCT motion estimation using a gradient-based optimization algorithm, which leverages generalized derivatives of the backprojection operator for cone-beam CT geometries. Building on that, a fully differentiable target function is formulated which grades the quality of the current motion estimate in reconstruction space. We drastically accelerate motion estimation yielding a 19-fold speed-up compared to existing methods. Additionally, we investigate the architecture of networks used for quality metric regression and propose predicting voxel-wise quality maps, favoring autoencoder-like architectures over contracting ones. This modification improves gradient flow, leading to more accurate motion estimation. The presented method is evaluated through realistic experiments on head anatomy. It achieves a reduction in reprojection error from an initial average of 3mm to 0.61mm after motion compensation and consistently demonstrates superior performance compared to existing approaches. The analytic Jacobian for the backprojection operation, which is at the core of the proposed method, is made publicly available. In summary, this paper contributes to the advancement of CBCT integration into clinical workflows by proposing a robust motion estimation approach that enhances efficiency and accuracy, addressing critical challenges in time-sensitive scenarios. Mareike Thies, Fabian Wagner, Noah Maul, Manuela Goldmann, Linda-Sophie Schneider, Mingxuan Gu, Siyuan Mei, Lukas Folle, Alexander Preuhs, Michael Manhart 0001, Andreas K. Maier |
IEEE Trans. Medical Imaging | 12 |
| 2024 | Style-Extracting Diffusion Models for Semi-supervised Histopathology Segmentation
Mathias Öttl, Frauke Wilm, Jana Steenpass, Jingna Qiu, Matthias Rübner, Arndt Hartmann, Matthias W. Beckmann, Peter A. Fasching, Andreas K. Maier, Ramona Erber, Bernhard Kainz, Katharina Breininger |
ECCV (75) | 9 |
| 2024 | Longitudinal Modeling of Depression Shifts Using Speech and LanguageabstractSpeech analysis can provide a potential non-invasive and objective means of assessing and monitoring an individual’s mental health. Most studies to date have focused on cross-sectional analysis and have not explored the benefits of speech analysis as a longitudinal monitoring tool that can assist in the management of chronic conditions such as major depressive disorder (MDD). Objectively monitoring for shifts in depression symptom severity levels over time presents a notable challenge, which we address through an automated approach using longitudinal English and Spanish speech samples collected from a clinical population. We employ time–frequency representations and linguistic embeddings to enhance the early recognition of alterations in depression levels in individuals with MDD. We investigate the suitability of using siamese-based training for modeling these changes, intending to enable personalized and adaptive interventions. Paula Andrea Pérez-Toro, Judith Dineley, Agnieszka Kaczkowska, Pauline Conde, Yuezhou Zhang 0001, Faith Matcham, Sara Siddi, Josep Maria Haro, Stuart Bruce, Til Wykes, Raquel Bailón, Srinivasan Vairavan, Richard J. B. Dobson, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave, Vaibhav A. Narayan, Nicholas Cummins |
ICASSP | 14 |
| 2024 | Transforming Cardiovascular Health: a Transformer-Based Approach to Continuous, Non-Invasive Blood Pressure Estimation via Radar SensingabstractHypertension is considered to be one of the most critical risk factors for cardiovascular diseases. As such, continuous, accurate and non-invasive monitoring of blood pressure (BP) is of utmost importance and research on such approaches is gaining momentum. In this study, we propose a novel transformer-based model architecture that leverages historic pressure wave information for accurate blood pressure regression. We achieve remarkable results that satisfy both the British Hypertension Society (BHS) and the Association for the Advancement of Medical Instrumentation (AAMI) blood pressure monitoring standards, with 97% of errors less than 5mmHg and 1.02 ± 1.77mmHg (mean absolute error ± standard deviation) accuracy for systolic (SBP), and 93% and 1.57 ± 2.36mmHg for diastolic BP (DBP). To the best of our knowledge, this is the first approach that utilizes transformers for BP regression and the first radar approach to satisfy BP standards, demonstrating the predictive power of the proposed model and the suitability of radar for the task. Nastassia Vysotskaya, Noah Maul, Alessandra Fusco, Souvik Hazra, Jens Harnisch, Tomás Arias-Vergara, Andreas K. Maier |
ICASSP | 7 |
| 2024 | Utilizing Deep Incomplete Classifiers to Implement Semantic Clustering for Killer Whale Photo Identification Data
Alexander Barnhill, Jared R. Towers, Elmar Nöth, Andreas K. Maier, Christian Bergler |
ICPR (1) | 4 |
| 2024 | SNOBERT: A Benchmark for Clinical Notes Entity Linking in the SNOMED CT Clinical Terminology
Mikhail Kulyabin, Gleb Sokolov, Aleksandr Galaida, Andreas K. Maier, Tomás Arias-Vergara |
ICPR (31) | 4 |
| 2024 | Generalist Segmentation Algorithm for Photoreceptors Analysis in Adaptive Optics Imaging
Mikhail Kulyabin, Aline Sindel, Hilde R. Pedersen, Stuart J. Gilson, Rigmor C. Baraas, Andreas K. Maier |
ICPR (28) | 6 |
| 2024 | Contrastive Learning Approach for Assessment of Phonological Precision in Patients with Tongue Cancer Using MRI DataabstractMagnetic Resonance Imaging (MRI) allows analyzing speech production by capturing high-resolution images of the dynamic processes in the vocal tract. In clinical applications, combining MRI with synchronized speech recordings leads to improved patient outcomes, especially if a phonological-based approach is used for assessment. However, when audio signals are unavailable, the recognition accuracy of sounds is decreased when using only MRI data. We propose a contrastive learning approach to improve the detection of phonological classes from MRI data when acoustic signals are not available at inference time. We demonstrate that frame-wise recognition of phonological classes improves from an f1 of 0.74 to 0.85 when the contrastive loss approach is implemented. Furthermore, we show the utility of our approach in the clinical application of using such phonological classes to assess speech disorders in patients with tongue cancer, yielding promising results in the recognition task. Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Maria Schuster, Elmar Nöth, Jonghye Woo, Andreas K. Maier |
INTERSPEECH | 11 |
| 2024 | ANIMAL-CLEAN - A Deep Denoising Toolkit for Animal-Independent Signal Enhancement
Alexander Barnhill, Elmar Nöth, Andreas K. Maier, Christian Bergler |
INTERSPEECH | 3 |
| 2024 | Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis
Kubilay Can Demir, Belén Lojo Rodríguez, Tobias Weise, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 4 |
| 2024 | Multilingual Speech and Language Analysis for the Assessment of Mild Cognitive Impairment: Outcomes from the Taukadial Challenge
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Philipp Klumpp, Tobias Weise, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier |
INTERSPEECH | 8 |
| 2024 | Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
Tobias Weise, Philipp Klumpp, Kubilay Can Demir, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Björn Heismann, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 8 |
| 2024 | DIREGA - Building Decision Support for German Register LawabstractThe interdisciplinary project DIREGA aims to analyse the joint application of linguistic, symbolic, and sub-symbolic AI techniques to German register law. This analysis is based on a dataset consisting of all past applications, related documents, and decisions of German register courts in the Free State of Bavaria. Corpus queries and sub-symbolic AI methods will be used for information extraction which then instantiate facts for the symbolic reasoning pipeline based on a manual formalization of the relevant laws. The goal is to build a prototypical implementation checking register applications and providing a detailed explanation for acceptance (or rejection) to aid legal professionals such as notaries in drafting and reviewing such documents. Axel Adrian, Osman Anil Basaran, Nathan Dykes, Stephanie Evert, Michael Gritz, Merlin Humml, Michael Kohlhase, Johannes Lindner, Andreas K. Maier, Stephan Prettner, Max Rapp, Lutz Schröder, Verena Stürmer |
JURIX | 9 |
| 2024 | Unsupervised Domain Adaptation Using Soft-Labeled Contrastive Learning with Reversed Monte Carlo Method for Cardiac Image Segmentation
Mingxuan Gu, Mareike Thies, Siyuan Mei, Fabian Wagner, Mingcheng Fan, Yipeng Sun, Zhaoya Pan, Sulaiman Vesal, Ronak Kosti, Dennis Possart, Jonas Utz, Andreas K. Maier |
MICCAI (9) | 12 |
| 2024 | Deep Learning for Cancer Prognosis Prediction Using Portrait Photos by StyleGAN Embedding
Amr Hagag, Ahmed Gomaa, Dominik Kornek, Andreas K. Maier, Rainer Fietkau, Christoph Bert, Yixing Huang, Florian Putz |
MICCAI (5) | 4 |
| 2024 | A Novel Tracking Framework for Devices in X-ray Leveraging Supplementary Cue-Driven Self-supervised Features
Saahil Islam, Venkatesh N. Murthy, Dominik Neumann, Serkan Çimen, Andreas K. Maier, Dorin Comaniciu, Florin C. Ghesu |
MICCAI (6) | 6 |
| 2024 | Cut to the Mix: Simple Data Augmentation Outperforms Elaborate Ones in Limited Organ Segmentation Datasets
Fuxin Fan, Annette Schwarz, Andreas K. Maier |
MICCAI (8) | 4 |
| 2024 | Tagged-to-Cine MRI Sequence Synthesis via Light Spatial-Temporal Transformer
Xiaofeng Liu 0001, Fangxu Xing, Zhangxing Bian, Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Andreas K. Maier, Maureen Stone 0001, Jiachen Zhuo, Jerry L. Prince, Jonghye Woo |
MICCAI (7) | 6 |
| 2024 | No-New-Denoiser: A Critical Analysis of Diffusion Models for Medical Image Denoising
Laura Pfaff, Fabian Wagner, Nastassia Vysotskaya, Mareike Thies, Noah Maul, Siyuan Mei, Tobias Würfl, Andreas K. Maier |
MICCAI (10) | 8 |
| 2024 | Death by Retrospective Undersampling - Caveats and Solutions for Learning-Based MRI Reconstructions
Junaid R. Rajput, Simon Weinmueller, Jonathan Endres, Peter Dawood, Florian Knoll, Andreas K. Maier, Moritz Zaiss |
MICCAI (7) | 6 |
| 2024 | Differentiable Score-Based Likelihoods: Learning CT Motion Compensation from Clean Images
Mareike Thies, Noah Maul, Siyuan Mei, Laura Pfaff, Nastassia Vysotskaya, Mingxuan Gu, Jonas Utz, Dennis Possart, Lukas Folle, Fabian Wagner, Andreas K. Maier |
MICCAI (7) | 11 |
| 2024 | CheXReport: A transformer-based architecture to generate chest X-ray reports suggestions
Felipe André Zeiser, Cristiano André da Costa, Gabriel de Oliveira Ramos, Andreas K. Maier, Rodrigo da Rosa Righi |
Expert Syst. Appl. | 4 |
| 2024 | Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
Mathias Zinnen, Prathmesh Madhu, Inger Leemans, Peter Bell 0007, Azhar Hussian, Hang Tran, Ali Hurriyetoglu, Andreas K. Maier, Vincent Christlein |
Expert Syst. Appl. | 8 |
| 2024 | The Artificial Neural Twin - Process optimization and continual learning in distributed process chainsabstractIndustrial process optimization and control is crucial to increase economic and ecologic efficiency. However, data sovereignty, differing goals, or the required expert knowledge for implementation impede holistic implementation. Further, the increasing use of data-driven AI-methods in process models and industrial sensory often requires regular fine-tuning to accommodate distribution drifts. We propose the Artificial Neural Twin, which combines concepts from model predictive control, deep learning, and sensor networks to address these issues. Our approach introduces decentral, differentiable data fusion to estimate the state of distributed process steps and their dependence on input data. By treating the interconnected process steps as a quasi neural-network, we can backpropagate loss gradients for process optimization or model fine-tuning to process parameters or AI models respectively. The concept is demonstrated on a virtual machine park simulated in Unity, consisting of bulk material processes in plastic recycling. Johannes Emmert, Ronald Mendez, Houman Mirzaalian Dastjerdi, Christopher Syben, Andreas K. Maier |
Neural Networks | 5 |
| 2024 | Contextual HookFormer for Glacier Calving Front SegmentationabstractPosition changes of glacier calving fronts are important indicators for evaluating the health of ice sheet outlet glaciers and changes in ice dynamics. However, manual delineation of calving fronts in remote sensing imagery is a time-consuming task, resulting in potential large costs. Deep learning-based methods have made remarkable progress in automatically segmenting and delineating glacier calving fronts from remote sensing imagery. The relatively few remote sensing images and the limited geometric changes for glacier observations both reduce the diversity of the data and exacerbate the difficulty of accurate segmentation. Here we describe a novel automatic method for detecting glacier calving fronts in synthetic aperture radar (SAR) images, termed HookFormer. Our approach processes high-resolution (target) and low-resolution (context) inputs with a unified Transformer architecture. The global-local tokens from the context and the target branches are integrated purely by the proposed cross-attention mechanism and cross-interaction module to complement and enhance each other. Moreover, we redesign the HookFormer architecture based on the CNN model AMD-HookNet aiming to improve computational efficiency while achieving significant performance gains with only half of the model parameters/FLOPs. We conduct an in-depth analysis and make extensive comparisons based on the challenging glacier segmentation benchmark dataset CaFFe. As the first pure Transformer approach, HookFormer sets a new state of the art with a mean distance error of 353 m to the ground truth, outperforming the baseline, Swin-Unet, and AMD-HookNet by 53 %, 39 %, and 19 %, respectively. Fei Wu 0025, Nora Gourmelon, Thorsten Seehaus, Jianlin Zhang 0001, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | PLIKS: A Pseudo-Linear Inverse Kinematic Solver for 3D Human Body EstimationabstractWe introduce PLIKS (Pseudo-Linear Inverse Kinematic Solver) for reconstruction of a 3D mesh of the human body from a single 2D image. Current techniques directly regress the shape, pose, and translation of a parametric model from an input image through a non-linear mapping with minimal flexibility to any external influences. We approach the task as a model-in-the-loop optimization problem. PLIKS is built on a linearized formulation of the parametric SMPL model. Using PLIKS, we can analytically reconstruct the human model via 2D pixel-aligned vertices. This enables us with the flexibility to use accurate camera calibration information when available. PLIKS offers an easy way to introduce additional constraints such as shape and translation. We present quantitative evaluations which confirm that PLIKS achieves more accurate reconstruction with greater than 10% improvement compared to other state-of-the-art methods with respect to the standard 3D human pose and shape benchmarks while also obtaining a reconstruction error improvement of 12.9 mm on the newer AGORA dataset. Karthik Shetty, Annette Birkhold, Srikrishna Jaganathan, Norbert Strobel, Markus Kowarschik, Andreas K. Maier, Bernhard Egger 0001 |
CVPR | 6 |
| 2023 | Transferring Quantified Emotion Knowledge for the Detection of Depression in Alzheimer's Disease Using ForestnetsabstractProgressive loss of memory is the most known symptom of Alzheimer’s Disease (AD); however, it also affects other cognitive skills and leads to depression symptoms. This paper presents a transfer learning strategy for automatically detecting AD and depression in AD patients using acoustic information and ForestNet, an artificial neural network that allows computing the contribution of a set of features to a model’s decision. The methodology consists of training ForestNet with a dataset commonly used for emotion recognition; then, we fine-tune the pre-trained model to detect AD and depression in AD. We trained the models with several acoustic features commonly used for emotion and AD applications. Unweighted average recalls of up to 0.87 were achieved to classify the disease and up to 0.82 to detect depression in AD. Our results indicate that the information obtained from the Arousal Valence plane may be suitable for discriminating and analyzing depression in AD. Paula Andrea Pérez-Toro, Dalia Rodríguez-Salas, Tomás Arias-Vergara, Sebastian P. Bayerl, Philipp Klumpp, Korbinian Riedhammer, Maria Schuster, Elmar Nöth, Andreas K. Maier, Juan Rafael Orozco-Arroyave |
ICASSP | 9 |
| 2023 | Word class representations spontaneously emerge in a deep neural network trained on next word predictionabstractHow do humans learn language, and can the first language be learned at all? These fundamental questions are still hotly debated. In contemporary linguistics, there are two major schools of thought that give completely opposite answers. According to Chomsky's theory of universal grammar, language cannot be learned because children are not exposed to sufficient data in their linguistic environment. In contrast, usage-based models of language assume a profound relationship between language structure and language use. In particular, contextual mental processing and mental representations are assumed to have the cognitive capacity to capture the complexity of actual language use at all levels. The prime example is syntax, i.e., the rules by which words are assembled into larger units such as sentences. Typically, syntactic rules are expressed as sequences of word classes. However, it remains unclear whether word classes are innate, as implied by universal grammar, or whether they emerge during language acquisition, as suggested by usage-based approaches. Here, we address this issue from a machine learning and natural language processing perspective. In particular, we trained an artificial deep neural network on predicting the next word, provided sequences of consecutive words as input. Subsequently, we analyzed the emerging activation patterns in the hidden layers of the neural network. Strikingly, we find that the internal representations of nine-word input sequences cluster according to the word class of the tenth word to be predicted as output, even though the neural network did not receive any explicit information about syntactic rules or word classes during training. This surprising result suggests, that also in the human brain, abstract representational categories such as word classes may naturally emerge as a consequence of predictive coding and processing during language acquisition. Kishore Surendra, Achim Schilling, Paul Stoewer, Andreas K. Maier, Patrick Krauss |
ICMLA | 4 |
| 2023 | Conditional Random Fields for Improving Deep Learning-Based Glacier Calving Front DelineationsabstractBig advancements in the field of Deep Learning allow the automated extraction of glacier calving fronts from satellite imagery. However, current efforts on Synthetic Aperture Radar (SAR) imagery still produce coarse and partly spurious predictions. Therefore, in this study, a two-dimensional fully-connected Conditional Random Field (CRF) is incorporated into the post-processing of a deep learning-based calving front delineation pipeline. The CRF takes as input the artificial neural network prediction and the original SAR image. Experiments are undertaken using the newly introduced benchmark dataset CaFFe and results are compared to the associated baseline. By introducing the CRF into the post-processing of the baseline’s pipeline, the mean distance error of the front prediction is improved by an average of 27 meters. The code is available at https://github.com/EntChanelt/GlacierCRF. Nora Gourmelon, Julian Klink, Thorsten Seehaus, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IGARSS | 5 |
| 2023 | Caffe - A Benchmark Dataset for Glacier Calving Front Extraction from Synthetic Aperture Radar ImageryabstractFrontal dynamics of marine-terminating glaciers play a crucial role in glacier projections. To reduce manual effort, deep learning methods can be employed to extract calving front positions from satellite imagery automatically. The newly introduced benchmark dataset "CaFFe" (Calving fronts and where to find them: a benchmark dataset and methodology for automatic glacier calving front extraction from SAR imagery) [1] provides multi-mission Synthetic Aperture Radar (SAR) imagery along with manually annotated calving fronts. CaFFe establishes a standardized framework for comparing deep learning techniques in glacier calving front extraction. By utilizing CaFFe to benchmark forthcoming deep learning models, researchers can identify the most promising directions for future research. A leaderboard of models can be accessed at https://paperswithcode.com/sota/calving-front-delineation-in-synthetic. Nora Gourmelon, Thorsten Seehaus, Julian Klink, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IGARSS | 5 |
| 2023 | Federated Learning for Secure Development of AI Models for Parkinson's Disease Detection Using Speech from Different LanguagesabstractParkinson's disease (PD) is a neurological disorder impacting a person's speech. Among automatic PD assessment methods, deep learning models have gained particular interest. Recently, the community has explored cross-pathology and cross-language models which can improve diagnostic accuracy even further. However, strict patient data privacy regulations largely prevent institutions from sharing patient speech data with each other. In this paper, we employ federated learning (FL) for PD detection using speech signals from 3 real-world language corpora of German, Spanish, and Czech, each from a separate institution. Our results indicate that the FL model outperforms all the local models in terms of diagnostic accuracy, while not performing very differently from the model based on centrally combined training sets, with the advantage of not requiring any data sharing among collaborators. This will simplify inter-institutional collaborations, resulting in enhancement of patient outcomes. Soroosh Tayebi Arasteh, Cristian D. Ríos-Urrego, Elmar Nöth, Andreas K. Maier, Seung-Hee Yang, Jan Rusz, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 4 |
| 2023 | Measuring Phonological Precision in Children with Cleft Lip and Palate
Tomás Arias-Vergara, Elizabeth Londoño-Mora, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier |
INTERSPEECH | 7 |
| 2023 | PoCaPNet: A Novel Approach for Surgical Phase Recognition Using Speech and X-Ray ImagesabstractSurgical phase recognition is a challenging and necessary task for the development of context-aware intelligent systems that can support medical personnel for better patient care and effective operating room management.In this paper, we present a surgical phase recognition framework that employs a Multi-Stage Temporal Convolution Network using speech and X-Ray images for the first time.We evaluate our proposed approach using our dataset that comprises 31 port-catheter placement operations and report 82.56 % frame-wise accuracy with eight surgical phases.Additionally, we investigate the design choices in the temporal model and solutions for the class-imbalance problem.Our experiments demonstrate that speech and X-Ray data can be effectively utilized for surgical phase recognition, providing a foundation for the development of speech assistants in operating rooms of the future. Kubilay Can Demir, Tobias Weise, Matthias S. May, Axel Schmid, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 5 |
| 2023 | Speaking Clearly, Understanding Better: Predicting the L2 Narrative Comprehension of Chinese Bilingual Kindergarten Children Based on Speech Intelligibility Using a Machine Learning Approach
Hiuching Hung, Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Andreas K. Maier, Elmar Nöth |
INTERSPEECH | 4 |
| 2023 | Automatic Assessment of Alzheimer's across Three Languages Using Speech and Language Features
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Franziska Braun, Florian Hönig, Carlos Tobon 0001, David Aguillón, Francisco Lopera, Liliana Hincapié-Henao, Maria Schuster, Korbinian Riedhammer, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 11 |
| 2023 | DeepFilterNet: Perceptually Motivated Real-Time Speech Enhancement
Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier |
INTERSPEECH | 4 |
| 2023 | Deep Multi-Frame Filtering for Hearing Aids
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
INTERSPEECH | 4 |
| 2023 | Robust Hough and Spatial-To-Angular Transform Based Rotation Estimation for Orthopedic X-Ray Images
Magdalena Bachmaier, Maximilian Rohleder, Benedict Swartman, Maxim Privalov, Andreas K. Maier, Holger Kunze |
MICCAI (7) | 5 |
| 2023 | Virtual Heart Models Help Elucidate the Role of Border Zone in Sustained Monomorphic Ventricular Tachycardia
Eduardo Castañeda, Masahito Suzuki, Hiroshi Ashikaga, Èric Lluch, Felix Meister, Viorel Mihalef, Chloé Audigier, Andreas K. Maier, Henry R. Halperin, Tiziano Passerini |
MICCAI (7) | 8 |
| 2023 | Deep Learning-Based Anonymization of Chest Radiographs: A Utility-Preserving Measure for Patient Privacy
Kai Packhäuser, Sebastian Gündel, Florian Thamm, Felix Denzinger, Andreas K. Maier |
MICCAI (3) | 5 |
| 2023 | Physics-Informed Conditional Autoencoder Approach for Robust Metabolic CEST MRI at 7T
Junaid R. Rajput, Tim A. Möhle, Moritz S. Fabian, Angelika Mennecke, Jochen A. Sembill, Joji B. Kuramatsu, Manuel Schmidt, Arnd Dörfler, Andreas K. Maier, Moritz Zaiss |
MICCAI (8) | 9 |
| 2023 | Flexible Unfolding of Circular Structures for Rendering Textbook-Style Cerebrovascular Maps
Leonhard Rist, Oliver Taubmann, Hendrik Ditt, Michael Sühling, Andreas K. Maier |
MICCAI (5) | 5 |
| 2023 | Enabling Geometry Aware Learning Through Differentiable Epipolar View Translation
Maximilian Rohleder, Charlotte Pradel, Fabian Wagner, Mareike Thies, Noah Maul, Felix Denzinger, Andreas K. Maier, Björn W. Kreher |
MICCAI (3) | 7 |
| 2023 | Self-Supervised 2D/3D Registration for X-Ray to CT Image FusionabstractDeep Learning-based 2D/3D registration enables fast, robust, and accurate X-ray to CT image fusion when large annotated paired datasets are available for training. However, the need for paired CT volume and X-ray images with ground truth registration limits the applicability in interventional scenarios. An alternative is to use simulated X-ray projections from CT volumes, thus removing the need for paired annotated datasets. Deep Neural Networks trained exclusively on simulated X-ray projections can perform significantly worse on real X-ray images due to the domain gap. We propose a self-supervised 2D/3D registration framework combining simulated training with unsupervised feature and pixel space domain adaptation to overcome the domain gap and eliminate the need for paired annotated datasets. Our framework achieves a registration accuracy of 1.83 ± 1.16 mm with a high success ratio of 90.1% on real X-ray images showing a 23.9% increase in success ratio compared to reference annotation-free algorithms. Srikrishna Jaganathan, Maximilian Kukla, Jian Wang 0009, Karthik Shetty, Andreas K. Maier |
WACV | 5 |
| 2023 | Pleural Effusion Classification on Chest X-Ray Images with Contrastive Learning
Felipe André Zeiser, Ismael Santos 0002, Henrique Bohn, Cristiano André da Costa, Gabriel de Oliveira Ramos, Rodrigo da Rosa Righi, Andreas K. Maier, José Rodrigo M. Andrade, Alexandre Bacelar |
WEBIST | 7 |
| 2023 | Updating Siamese trackers using peculiar mixup
Fei Wu 0025, Jianlin Zhang 0001, Zhiyong Xu 0007, Andreas K. Maier, Vincent Christlein |
Appl. Intell. | 4 |
| 2023 | Bayesian Convolutional Neural Networks for Limited Data Hyperspectral Remote Sensing Image ClassificationabstractHyperspectral remote sensing (HSRS) images have high dimensionality, and labeling HSRS data is expensive and therefore limited to small amounts of pixels. This makes it challenging to use deep neural networks for HSRS image classification. In extreme cases, deep neural networks are even outperformed by traditional models. In this work, we propose to use Bayesian convolutional neural networks (BCNNs) as a potential alternative to convolutional neural networks (CNNs). BCNNs benefit from Bayesian learning, which is more robust against overfitting and inherently provides a measure for uncertainty. We show in experiments on the Pavia Centre, Salinas, and Botswana datasets that a BCNN outperforms a similarly constructed non-Bayesian CNN, an off-the-shelf random forest (RF), and a state-of-the-art Bayesian neural network (BNN). We also show that BCNN is more robust against overfitting compared with the CNN. Furthermore, the BCNN exhibits a remarkably larger capacity for model compression, which makes BCNN a better candidate in hardware-constrained settings. Finally, we show that the BCNN’s uncertainty measure can effectively identify misclassified samples. This useful property can be used to detect mislabeled data or to reject predictions with low confidence. Mohammad Joshaghani, AmirAbbas Davari, Faezeh Nejati Hatamian, Andreas K. Maier, Christian Riess |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Mitosis domain generalization in histopathology images - The MIDOG challenge
Marc Aubreville, Nikolas Stathonikos, Christof Bertram, Robert Klopfleisch, Natalie D. ter Hoeve, Francesco Ciompi, Frauke Wilm, Christian Marzahl, Taryn A. Donovan, Andreas K. Maier, Jack Breen, Nishant Ravikumar, Youjin Chung, Jinah Park, Ramin Nateghi, Fattaneh Pourakpour, Rutger H. J. Fick, Saima Ben Hadj, Mostafa Jahanifar, Adam J. Shephard, Jakob Dexl, Thomas Wittenberg, Satoshi Kondo, Maxime W. Lafarge, Viktor H. Koelzer, Jingtang Liang, Yubo Wang 0001, Jingxin Liu 0005, Salar Razavi, April Khademi, Sen Yang 0006, Ramona Erber, Andrea Klang, Karoline Lipnik, Pompei Bolfa, Michael J. Dark, Gabriel Wasinger, Mitko Veta, Katharina Breininger |
Medical Image Anal. | 10 |
| 2023 | ICC++: Explainable feature learning for art history using image compositions
Prathmesh Madhu, Tilman Marquart, Ronak Kosti, Dirk Suckow, Peter Bell 0007, Andreas K. Maier, Vincent Christlein |
Pattern Recognit. | 6 |
| 2023 | AMD-HookNet for Glacier Front SegmentationabstractKnowledge on changes in glacier calving front positions is important for assessing the status of glaciers. Remote sensing imagery provides the ideal database for monitoring calving front positions; however, it is not feasible to perform this task manually for all calving glaciers globally due to time constraints. Deep-learning-based methods have shown great potential for glacier calving front delineation from optical and radar satellite imagery. The calving front is represented as a single thin line between the ocean and the glacier, which makes the task vulnerable to inaccurate predictions. The limited availability of annotated glacier imagery leads to a lack of data diversity (not all possible combinations of different weather conditions, terminus shapes, sensors, etc. are present in the data), which exacerbates the difficulty of accurate segmentation. In this article, we propose attention-multihooking-deep-supervision HookNet (AMD-HookNet), a novel glacier calving front segmentation framework for synthetic aperture radar (SAR) images. The proposed method aims to enhance the feature representation capability through multiple information interactions between low-resolution and high-resolution inputs based on a two-branch U-Net. The attention mechanism, integrated into the two branch U-Net, aims to interact between the corresponding coarse and fine-grained feature maps. This allows the network to automatically adjust feature relationships, resulting in accurate pixel classification predictions. Extensive experiments and comparisons on the challenging glacier segmentation benchmark dataset CaFFe show that our AMD-HookNet achieves a mean distance error (MDE) of 438 m to the ground truth outperforming the current state of the art by 42%, which validates its effectiveness. Fei Wu 0025, Nora Gourmelon, Thorsten Seehaus, Jianlin Zhang 0001, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Deep Learning in Surgical Workflow Analysis: A Review of Phase and Step RecognitionabstractOBJECTIVE: In the last two decades, there has been a growing interest in exploring surgical procedures with statistical models to analyze operations at different semantic levels. This information is necessary for developing context-aware intelligent systems, which can assist the physicians during operations, evaluate procedures afterward or help the management team to effectively utilize the operating room. The objective is to extract reliable patterns from surgical data for the robust estimation of surgical activities performed during operations. The purpose of this article is to review the state-of-the-art deep learning methods that have been published after 2018 for analyzing surgical workflows, with a focus on phase and step recognition. METHODS: Three databases, IEEE Xplore, Scopus, and PubMed were searched, and additional studies are added through a manual search. After the database search, 343 studies were screened and a total of 44 studies are selected for this review. CONCLUSION: The use of temporal information is essential for identifying the next surgical action. Contemporary methods used mainly RNNs, hierarchical CNNs, and Transformers to preserve long-distance temporal relations. The lack of large publicly available datasets for various procedures is a great challenge for the development of new and robust models. As supervised learning strategies are used to show proof-of-concept, self-supervised, semi-supervised, or active learning methods are used to mitigate dependency on annotated data. SIGNIFICANCE: The present study provides a comprehensive review of recent methods in surgical workflow analysis, summarizes commonly used architectures, datasets, and discusses challenges. Kubilay Can Demir, Hannah Schieber, Tobias Weise, Daniel Roth 0001, Matthias S. May, Andreas K. Maier, Seung-Hee Yang |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Is Multitask Learning Always Better?
Alexander Mattick, Martin Mayr, Andreas K. Maier, Vincent Christlein |
DAS | 3 |
| 2022 | Combining Visual and Linguistic Models for a Robust Recipient Line Recognition in Historical Documents
Martin Mayr, Alex Felker, Andreas K. Maier, Vincent Christlein |
DAS | 3 |
| 2022 | ORCA-PARTY: An Automatic Killer Whale Sound Type Separation Toolkit Using Deep LearningabstractData-driven and machine-based analysis of massive bioacoustic data collections, in particular acoustic regions containing a substantial number of vocalizations events, is essential and extremely valuable to identify recurring vocal paradigms. However, these acoustic sections are usually characterized by a strong incidence of overlapping vocalization events, a major problem severely affecting subsequent human-/machine-based analysis and interpretation. Robust machine-driven signal separation of species-specific call types is extremely challenging due to missing ground truth data, speaker/source-relevant information, limited knowledge about inter- and intra-call type variations, next to diverse recording conditions. The current study is the first introducing a fully-automated deep signal separation approach for overlapping orca vocalizations, addressing all of the previously mentioned challenges, together with one of the largest bioacoustic data archives recorded on killer whales (Orcinus Orca). Incorporating ORCA-PARTY as additional data enhancement step for downstream call type classification demonstrated to be extremely valuable. Besides the proof of cross-domain applicability and consistently promising results on non-overlapping signals, significant improvements were achieved when processing acoustic orca segments comprising a multitude of vocal activities. Apart from auspicious visual inspections, a final numerical evaluation on an unseen dataset proved that about 30 % more known sound patterns could be identified. Christian Bergler, Manuel Schmitt, Andreas K. Maier, Rachael Cheng, Volker Barth, Elmar Nöth |
ICASSP | 3 |
| 2022 | Deepfilternet: A Low Complexity Speech Enhancement Framework for Full-Band Audio Based On Deep FilteringabstractComplex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks (CM) are usually preferred over real-valued masks due to their ability to modify the phase. Recent work proposed to use a complex filter instead of a point-wise multiplication with a mask. This allows to incorporate information from previous and future time steps exploiting local correlations within each frequency band.In this work, we propose DeepFilterNet, a two stage speech enhancement framework utilizing deep filtering. First, we enhance the spectral envelope using ERB-scaled gains modeling the human frequency perception. The second stage employs deep filtering to enhance the periodic components of speech. Additionally to taking advantage of perceptual properties of speech, we enforce network sparsity via separable convolutions and extensive grouping in linear and recurrent layers to design a low complexity architecture.We further show that our two stage deep filtering approach outperforms complex masks over a variety of frequency resolutions and latencies and demonstrate convincing performance compared to other state-of-the-art models. Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier |
ICASSP | 4 |
| 2022 | Supervised Contrastive Learning for Robust and Efficient Multi-modal Emotion and Sentiment AnalysisabstractExpression of human emotion and sentiment are often multi-modal consisting use of spoken speech, vision, and text. Combining multiple modalities allows learning-based models to benefit with the complementary information present across modalities to produce more accurate predictions. One of the bigger challenges in multi-modal affective computing is performance consistency in non-ideal scenarios. Most benchmarks fail to generalize in non-ideal scenarios where one of the modalities is missing or highly corrupted due to occlusion, sensor errors, or change of orientation. Consequently, various modality fusion approaches were proposed. However, most of these fusion approaches assume that each modality is equally useful. To address the challenge of performance consistency, in this work we propose to use supervised contrastive learning (SCL). We demonstrate through various experiments and comparison with state-of-the-art (SOTA) methods that the model robustness against corrupted and missing modalities improves when trained with SCL. Next, we use the Perceiver architecture [1] in order to efficiently combine the representations of different modalities. Its iterative attention mechanism allows to create a reduced latent representation in an efficient manner. We observe that it can accommodate a wide range of modality combinations, allowing for robust information fusion. Our approach allows reduction of model complexity and efficient fusion of different modalities, while maintaining the performance consistency and model robustness. We conduct ablation experiments to study the effect of each contribution in different scenarios, and we show that the proposed methods outperform the state-of-art, while simultaneously being robust to corrupted modalities. Our method also outperforms its counterparts and SOTA while using less numerical complexity (inference times and compute operations). Ahmed Gomaa, Andreas K. Maier, Ronak Kosti |
ICPR | 2 |
| 2022 | ODOR: The ICPR2022 ODeuropa Challenge on Olfactory Object RecognitionabstractThe Odeuropa Challenge on Olfactory Object Recognition aims to foster the development of object detection in the visual arts and to promote an olfactory perspective on digital heritage. Object detection in historical artworks is particularly challenging due to varying styles and artistic periods. Moreover, the task is complicated due to the particularity and historical variance of predefined target objects, which exhibit a large intra-class variance, and the long tail distribution of the dataset labels, with some objects having only very few training examples. These challenges should encourage participants to create innovative approaches using domain adaptation or few-shot learning. We provide a dataset of 2647 artworks annotated with 20 120 tightly fit bounding boxes that are split into a training and validation set (public). A test set containing 1140 artworks and 15 480 annotations is kept private for the challenge evaluation. Mathias Zinnen, Prathmesh Madhu, Ronak Kosti, Peter Bell 0007, Andreas K. Maier, Vincent Christlein |
ICPR | 5 |
| 2022 | ORCA-WHISPER: An Automatic Killer Whale Sound Type Generation Toolkit Using Deep Learning
Christian Bergler, Alexander Barnhill, Dominik Perrin, Manuel Schmitt, Andreas K. Maier, Elmar Nöth |
INTERSPEECH | 5 |
| 2022 | Cross-lingual Self-Supervised Speech Representations for Improved Dysarthric Speech RecognitionabstractState-of-the-art automatic speech recognition (ASR) systems perform well on healthy speech.However, the performance on impaired speech still remains an issue.The current study explores the usefulness of using Wav2Vec self-supervised speech representations as features for training an ASR system for dysarthric speech.Dysarthric speech recognition is particularly difficult as several aspects of speech such as articulation, prosody and phonation can be impaired.Specifically, we train an acoustic model with features extracted from Wav2Vec, Hubert, and the cross-lingual XLSR model.Results suggest that speech representations pretrained on large unlabelled data can improve word error rate (WER) performance.In particular, features from the multilingual model led to lower WERs than filterbanks (Fbank) or models trained on a single language.Improvements were observed in English speakers with cerebral palsy caused dysarthria (UASpeech corpus), Spanish speakers with Parkinsonian dysarthria (PC-GITA corpus) and Italian speakers with paralysis-based dysarthria (EasyCall corpus).Compared to using Fbank features, XLSR-based features reduced WERs by 6.8%, 22.0%, and 7.0% for the UASpeech, PC-GITA, and EasyCall corpus, respectively. Abner Hernandez, Paula Andrea Pérez-Toro, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 5 |
| 2022 | Alzheimer's Detection from English to Spanish Using Acoustic and Linguistic Embeddings
Paula Andrea Pérez-Toro, Philipp Klumpp, Abner Hernandez, Tomas Arias, Patricia Lillo, Andrea Slachevsky, Adolfo M. García, Maria Schuster, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 9 |
| 2022 | CoachLea: an Android Application to Evaluate the Speech Production and Perception of Children with Hearing Loss
P. Schäfer, Paula Andrea Pérez-Toro, Philipp Klumpp, Juan Rafael Orozco-Arroyave, Elmar Nöth, Andreas K. Maier, A. Abad, Maria Schuster, Tomás Arias-Vergara |
INTERSPEECH | 6 |
| 2022 | Disentangled Latent Speech Representation for Automatic Pathological Intelligibility AssessmentabstractSpeech intelligibility assessment plays an important role in the therapy of patients suffering from pathological speech disorders.Automatic and objective measures are desirable to assist therapists in their traditionally subjective and labor-intensive assessments.In this work, we investigate a novel approach for obtaining such a measure using the divergence in disentangled latent speech representations of a parallel utterance pair, obtained from a healthy reference and a pathological speaker.Experiments on an English database of Cerebral Palsy patients, using all available utterances per speaker, show high and significant correlation values (R = -0.9)with subjective intelligibility measures, while having only minimal deviation (±0.01) across four different reference speaker pairs.We also demonstrate the robustness of the proposed method (R = -0.89deviating ±0.02 over 1000 iterations) by considering a significantly smaller amount of utterances per speaker.Our results are among the first to show that disentangled speech representations can be used for automatic pathological speech intelligibility assessment, resulting in a reference speaker pair invariant method, applicable in scenarios with only few utterances available. Tobias Weise, Philipp Klumpp, Andreas K. Maier, Elmar Nöth, Björn Heismann, Maria Schuster, Seung-Hee Yang |
INTERSPEECH | 3 |
| 2022 | Deep Geometric Supervision Improves Spatial Generalization in Orthopedic Surgery Planning
Florian Kordon, Andreas K. Maier, Benedict Swartman, Maxim Privalov, Jan Siad El Barbari, Holger Kunze |
MICCAI (8) | 2 |
| 2022 | Fast Automatic Liver Tumor Radiofrequency Ablation Planning via Learned Physics Model
Felix Meister, Chloé Audigier, Tiziano Passerini, Èric Lluch, Viorel Mihalef, Andreas K. Maier, Tommaso Mansi |
MICCAI (8) | 6 |
| 2022 | A Spatiotemporal Model for Precise and Efficient Fully-Automatic 3D Motion Correction in OCT
Stefan B. Ploner, Jungeun Won, Lennart Husvogt, Katharina Breininger, Julia Schottenhamml, James G. Fujimoto, Andreas K. Maier |
MICCAI (2) | 8 |
| 2022 | Multi-modal Retinal Image Registration Using a Keypoint-Based Vessel Structure Aligning Network
Aline Sindel, Bettina Hohberger, Andreas K. Maier, Vincent Christlein |
MICCAI (6) | 3 |
| 2022 | Building Brains: Subvolume Recombination for Data Augmentation in Large Vessel Occlusion Detection
Florian Thamm, Oliver Taubmann, Markus Jürgens, Aleksandra Thamm, Felix Denzinger, Leonhard Rist, Hendrik Ditt, Andreas K. Maier |
MICCAI (3) | 8 |
| 2022 | A multi-sensor architecture combining human pose estimation and real-time location systems for workflow monitoring on hybrid operating suites
Vinicius Facco Rodrigues, Rodolfo Stoffel Antunes, Lucas Adams Seewald, Rodrigo Bazo, Eduardo Souza dos Reis, Uélison Jean Lopes dos Santos, Rodrigo da Rosa Righi, Luiz Gonzaga 0001, Cristiano André da Costa, Felipe L. Bertollo, Andreas K. Maier, Björn M. Eskofier, Tim Horz, Marcus Pfister, Rebecca Fahrig |
Future Gener. Comput. Syst. | 11 |
| 2022 | FlexParser - The adaptive log file parser for continuous results in a changing worldabstractAny modern system writes events into files, called log files. Those contain crucial information which are subject to various analyses. Examples range from cybersecurity, intrusion detection over usage analyses to trouble shooting. Before data analysis is possible, desired information needs to be extracted first out of the semi-structured log messages. State-of-the-art event parsing often assumes static log events. However, any modern system is updated consistently and with updates also log file structures can change. We call those changes "mutation" and study parsing performance for different mutation cases. Latest research discovers mutations using anomaly detection post mortem, however, does not cover actual continuous parsing. Thus, we propose a novel and flexible parser, called FlexParser, which can extract desired values despite gradual changes in the log messages. It implies basic text preprocessing followed by a supervised Deep Learning method. We train a stateful LSTM on parsing one event per data set. Statefulness enforces the model to learn log message structures across several examples. Our model was tested on seven different, publicly available log file data sets and various kinds of mutations. Exhibiting an average F1-Score of 0.98, it outperforms other Deep Learning methods as well as state-of-the-art unsupervised parsers. Nadine Rücker, Andreas K. Maier |
J. Softw. Evol. Process. | 2 |
| 2022 | Low Latency Speech Enhancement for Hearing Aids Using Deep FilteringabstractNoise reduction is an important feature supporting hearing aid (HA) users in their daily routines and is thus included in most commercially available devices. Latency requirements of HAs require short processing windows resulting in a poor frequency resolution in the whole processing chain including noise reduction. Previous studies have shown that deep neural network (DNN) based algorithms outperform conventional noise reduction algorithms especially for non-stationary noises. This study explores a DNN based noise reduction method using deep filtering targeted for wideband spectrograms given the employed HA filter bank. That is, we predict complex filter coefficients that are linearly applied to the noisy spectrum. We assess different filter sizes over time and frequency axis, and provide evidence for a superior performance over a complex ratio mask. Furthermore, we introduce a frequency response loss that operates on a per-frequency-band basis to fully utilize the deep filtering concept. We objectively demonstrate on-par performance with related state-of-the-art deep learning methods and show in a subjective user study that our method is perceptually preferred to existing HA noise reduction algorithms. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Pixelwise Distance Regression for Glacier Calving Front Detection and SegmentationabstractGlacier calving front position (CFP) is an important glaciological variable. Traditionally, delineating the CFPs has been carried out manually, which was subjective, tedious, and expensive. Automating this process is crucial for continuously monitoring the evolution and status of glaciers. Recently, deep learning approaches have been investigated for this application. However, the current methods get challenged by a severe class imbalance problem. In this work, we propose to mitigate the class imbalance between the calving front class and the noncalving front class by reformulating the segmentation problem into a pixelwise regression task. A convolutional neural network (CNN) gets optimized to predict the distance values to the glacier front for each pixel in the image. The resulting distance map localizes the CFP and is further postprocessed to extract the calving front line. We propose three postprocessing methods, one method based on statistical thresholding, a second method based on conditional random fields (CRFs), and finally the use of a second U-Net. The experimental results confirm that our approach significantly outperforms the state-of-the-art methods and produces accurate delineation. The second U-Net obtains the best performance results, resulting in an average improvement of about 21% Dice coefficient enhancement. AmirAbbas Davari, Christoph Baller, Thorsten Seehaus, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | On Mathews Correlation Coefficient and Improved Distance Map Loss for Automatic Glacier Calving Front Segmentation in SAR ImageryabstractThe vast majority of the outlet glaciers and ice streams of the polar ice sheets end in the ocean. Ice mass loss via calving of the glaciers into the ocean has increased over the last few decades. Information on the temporal variability of the calving front position provides fundamental information on the state of the glacier and ice stream, which can be exploited as calibration and validation data to enhance ice dynamics modeling. To identify the calving front position automatically, deep neural network-based semantic segmentation pipelines can be used to delineate the acquired SAR imagery. However, the extreme class imbalance is highly challenging for the accurate calving front segmentation in these images. Therefore, we propose the use of the Mathews correlation coefficient (MCC) as an early stopping criterion because of its symmetrical properties and its invariance towards class imbalance. Moreover, we propose an improvement to the distance map-based binary cross-entropy (BCE) loss function. The distance map adds context to the loss function about the important regions for segmentation and helps accounting for the imbalanced data. Using Mathews correlation coefficient as early stopping demonstrates an average 15% dice coefficient improvement compared to the commonly used BCE. The modified distance map loss further improves the segmentation performance by another 2%. These results are encouraging as they support the effectiveness of the proposed methods for segmentation problems suffering from extreme class imbalances. AmirAbbas Davari, Saahil Islam, Thorsten Seehaus, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | SmartPatch: Improving Handwritten Word Imitation with Patch Discriminators
Alexander Mattick, Martin Mayr, Mathias Seuret, Andreas K. Maier, Vincent Christlein |
ICDAR (1) | 4 |
| 2021 | ICDAR 2021 Competition on Historical Document Classification
Mathias Seuret, Anguelos Nicolaou, Dalia Rodríguez-Salas, Nikolaus Weichselbaumer, Dominique Stutzmann, Martin Mayr, Andreas K. Maier, Vincent Christlein |
ICDAR (4) | 7 |
| 2021 | Skinscan: Low-Cost 3D-Scanning for Dermatologic Diagnosis and DocumentationabstractThe utilization of computational photography becomes increasingly essential in the medical field. Today, imaging techniques for dermatology range from two-dimensional (2D) color imagery with a mobile device to professional clinical imaging systems measuring additional detailed three-dimensional (3D) data. The latter are commonly expensive and not accessible to a broad audience. In this work, we propose a novel system and software framework that relies only on low-cost (and even mobile) commodity devices present in every household to measure detailed 3D information of the human skin with a 3D-gradient-illumination-based method. We believe that our system has great potential for early-stage diagnosis and monitoring of skin diseases, especially in vastly populated or underdeveloped areas. Merlin A. Nau, Florian Schiffers, Andreas K. Maier, Jack Tumblin, Marc Walton, Aggelos K. Katsaggelos, Florian Willomitzer, Oliver Cossairt |
ICIP | 5 |
| 2021 | Craquelurenet: Matching The Crack Structure In Historical Paintings For Multi-Modal Image RegistrationabstractVisual light photography, infrared reflectography, ultraviolet fluorescence photography and x-radiography reveal even hidden compositional layers in paintings. To investigate the connections between these images, a multi-modal registration is essential. Due to varying image resolutions, modality dependent image content and depiction styles, registration poses a challenge. Historical paintings usually show crack structures called craquelure in the paint. Since craquelure is visible by all modalities, we extract craquelure features for our multi-modal registration method using a convolutional neural network. We jointly train our keypoint detector and descriptor using multi-task learning. We created a multi-modal dataset of historical paintings with keypoint pair annotations and class labels for craquelure detection and matching. Our method demonstrates the best registration performance on the multi-modal dataset in comparison to competing methods. Aline Sindel, Andreas K. Maier, Vincent Christlein |
ICIP | 2 |
| 2021 | Recurrence Plot Spacial Pyramid Pooling Network for Appliance Identification in Non-Intrusive Load MonitoringabstractParameter free Non-intrusive Load Monitoring (NILM) algorithms are a major step toward real-world NILM scenarios. The identification of appliances is the key element in NILM. The task consists of identification of the appliance category and its current state. In this paper, we present a parameter free appliance identification algorithm for NILM using a 2D representation of time series known as unthresholded Recurrence Plots (RP) for appliance category identification. One cycle of voltage and current (V-I trajectory) are transformed into a RP and classified using a Spacial Pyramid Pooling Convolutional Neural Network architecture. The performance of our approach is evaluated on the three public datasets COOLL, PLAID and WHITEDvl.1 and compared to previous publications. We show that compared to other approaches using our architecture no initial parameters have to be manually tuned for each specific dataset. Marc Wenninger, Sebastian P. Bayerl, Andreas K. Maier, Jochen Schmidt |
ICMLA | 3 |
| 2021 | Synthetic Glacier SAR Image Generation from Arbitrary Masks Using Pix2Pix AlgorithmabstractSupervised machine learning requires a large amount of labeled data to achieve proper test results. However, generating accurately labeled segmentation maps on remote sensing imagery, including images from synthetic aperture radar (SAR), is tedious and highly subjective. In this work, we propose to alleviate the issue of limited training data by generating synthetic SAR images with the pix2pix algorithm [1]. This algorithm uses conditional Generative Adversarial Networks (cGANs) to generate an artificial image while preserving the structure of the input. In our case, the input is a segmentation mask, from which a corresponding synthetic SAR image is generated. We present different models, perform a comparative study and demonstrate that this approach synthesizes convincing glaciers in SAR images with promising qualitative and quantitative results. Rosanna Dietrich-Sussner, AmirAbbas Davari, Thorsten Seehaus, Matthias H. Braun, Vincent Christlein, Andreas K. Maier, Christian Riess |
IGARSS | 6 |
| 2021 | Bayesian U-Net for Segmenting Glaciers in Sar ImageryabstractFluctuations of the glacier calving front have an important influence over the ice flow of whole glacier systems. It is therefore important to precisely monitor the position of the calving front. However, the manual delineation of SAR images is a difficult, laborious and subjective task. Convolutional neural networks have previously shown promising results in automating the glacier segmentation in SAR images, making them desirable for further exploration of their possibilities. In this work, we propose to compute uncertainty and use it in an Uncertainty Optimization regime as a novel two-stage process. By using dropout as a random sampling layer in a U-Net architecture, we create a probabilistic Bayesian Neural Network. With several forward passes we create a sampling distribution, which can estimate the model uncertainty for each pixel in the segmentation mask. The additional uncertainty map information can serve as a guideline for the experts in the manual annotation of the data. Furthermore, feeding the uncertainty map to the network leads to 95.24 % Dice similarity, which is an overall improvement in the segmentation performance compared to the state-of-the-art deterministic U-Net-based glacier segmentation pipelines. AmirAbbas Davari, Thorsten Seehaus, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IGARSS | 5 |
| 2021 | Glacier Calving Front Segmentation Using Attention U-NetabstractAn essential climate variable to determine the tidewater glacier status is the location of the calving front position and the separation of seasonal variability from long-term trends. Previous studies have proposed deep learning-based methods to semi-automatically delineate the calving fronts of tidewater glaciers. They used U-Net to segment the ice and non-ice regions and extracted the calving fronts in a post-processing step. In this work, we show a method to segment the glacier calving fronts from SAR images in an end-to-end fashion using Attention U-Net. The main objective is to investigate the attention mechanism in this application. Adding attention modules to the state-of-the-art U - N et network lets us analyze the learning process by extracting its attention maps. We use these maps as a tool to search for proper hyperparameters and loss functions in order to generate higher qualitative results. Our proposed attention U-Net performs comparably to the standard U-Net while providing additional insight into those regions on which the network learned to focus more. In the best case, the attention U-Net achieves a 1.5 % better Dice score compared to the canonical U-Net with a glacier front line prediction certainty of up to 237.12 meters. Michael Holzmann, AmirAbbas Davari, Thorsten Seehaus, Matthias H. Braun, Andreas K. Maier, Vincent Christlein |
IGARSS | 5 |
| 2021 | ORCA-SLANG: An Automatic Multi-Stage Semi-Supervised Deep Learning Framework for Large-Scale Killer Whale Call Type Identification
Christian Bergler, Manuel Schmitt, Andreas K. Maier, Helena Symonds, Paul Spong, Steven R. Ness, George Tzanetakis, Elmar Nöth |
Interspeech | 3 |
| 2021 | LACOPE: Latency-Constrained Pitch Estimation for Speech Enhancement
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
Interspeech | 4 |
| 2021 | Deep Iterative 2D/3D Registration
Srikrishna Jaganathan, Jian Wang 0009, Anja Borsdorf, Karthik Shetty, Andreas K. Maier |
MICCAI (4) | 5 |
| 2021 | Automatic Path Planning for Safe Guide Pin Insertion in PCL Reconstruction Surgery
Florian Kordon, Andreas K. Maier, Benedict Swartman, Maxim Privalov, Jan Siad El Barbari, Holger Kunze |
MICCAI (4) | 2 |
| 2021 | SyNCCT: Synthetic Non-contrast Images of the Brain from Single-Energy Computed Tomography Angiography
Florian Thamm, Oliver Taubmann, Felix Denzinger, Markus Jürgens, Hendrik Ditt, Andreas K. Maier |
MICCAI (7) | 6 |
| 2021 | Robust classification from noisy labels: Integrating additional knowledge for chest radiography abnormality assessment
Sebastian Gündel, Arnaud A. A. Setio, Florin C. Ghesu, Sasa Grbic, Bogdan Georgescu, Andreas K. Maier, Dorin Comaniciu |
Medical Image Anal. | 6 |
| 2021 | Cephalogram synthesis and landmark detection in dental cone-beam CT systems
Yixing Huang, Fuxin Fan, Christopher Syben, Philipp Roser, Leonid Mill, Andreas K. Maier |
Medical Image Anal. | 6 |
| 2021 | A global benchmark of algorithms for segmenting the left atrium from late gadolinium-enhanced cardiac magnetic resonance imaging
Zhaohan Xiong, Qing Xia 0002, Cheng Bian, Yefeng Zheng 0001, Sulaiman Vesal, Nishant Ravikumar, Andreas K. Maier, Xin Yang 0009, Pheng-Ann Heng, Dong Ni 0001, Caizi Li, Qianqian Tong 0001, Weixin Si, Élodie Puybareau, Younes Khoudli, Thierry Géraud, Jichao Zhao |
Medical Image Anal. | 9 |
| 2021 | Segmentation of photovoltaic module cells in uncalibrated electroluminescence imagesabstractAbstract High resolution electroluminescence (EL) images captured in the infrared spectrum allow to visually and non-destructively inspect the quality of photovoltaic (PV) modules. Currently, however, such a visual inspection requires trained experts to discern different kinds of defects, which is time-consuming and expensive. Automated segmentation of cells is therefore a key step in automating the visual inspection workflow. In this work, we propose a robust automated segmentation method for extraction of individual solar cells from EL images of PV modules. This enables controlled studies on large amounts of data to understanding the effects of module degradation over time—a process not yet fully understood. The proposed method infers in several steps a high-level solar module representation from low-level ridge edge features. An important step in the algorithm is to formulate the segmentation problem in terms of lens calibration by exploiting the plumbline constraint. We evaluate our method on a dataset of various solar modules types containing a total of 408 solar cells with various defects. Our method robustly solves this task with a median weighted Jaccard index of $$94.47\%$$ 94.47% and an $$F_1$$ F1 score of $$97.62\%$$ 97.62% , both indicating a high sensitivity and a high similarity between automatically segmented and ground truth solar cell masks. Sergiu Deitsch, Claudia Buerhop-Lutz, Evgenii Sovetkin, Ansgar Steland, Andreas K. Maier, Florian Gallwitz, Christian Riess |
Mach. Vis. Appl. | 5 |
| 2021 | Enhancing collaborative road scene reconstruction with unsupervised domain alignment
Moritz Venator, Selcuk Aklanoglu, Erich Bruns, Andreas K. Maier |
Mach. Vis. Appl. | 4 |
| 2021 | Quantifying the separability of data classes in neural networksabstractWe introduce the Generalized Discrimination Value (GDV) that measures, in a non-invasive manner, how well different data classes separate in each given layer of an artificial neural network. It turns out that, at the end of the training period, the GDV in each given layer L attains a highly reproducible value, irrespective of the initialization of the network's connection weights. In the case of multi-layer perceptrons trained with error backpropagation, we find that classification of highly complex data sets requires a temporal reduction of class separability, marked by a characteristic 'energy barrier' in the initial part of the GDV(L) curve. Even more surprisingly, for a given data set, the GDV(L) is running through a fixed 'master curve', independently from the total number of network layers. Finally, due to its invariance with respect to dimensionality, the GDV may serve as a useful tool to compare the internal representational dynamics of artificial neural networks with different architectures for neural architecture search or network compression; or even with brain activity in order to decide between different candidate models of brain function. Achim Schilling, Andreas K. Maier, Richard Gerum, Claus Metzner, Patrick Krauss |
Neural Networks | 2 |
| 2021 | Monocular multi-person pose estimation: A survey
Eduardo Souza dos Reis, Lucas Adams Seewald, Rodolfo Stoffel Antunes, Vinicius Facco Rodrigues, Rodrigo da Rosa Righi, Cristiano André da Costa, Luiz Gonzaga 0001, Björn M. Eskofier, Andreas K. Maier, Tim Horz, Rebecca Fahrig |
Pattern Recognit. | 9 |
| 2021 | Spatio-Temporal Multi-Task Learning for Cardiac MRI Left Ventricle QuantificationabstractQuantitative assessment of cardiac left ventricle (LV) morphology is essential to assess cardiac function and improve the diagnosis of different cardiovascular diseases. In current clinical practice, LV quantification depends on the measurement of myocardial shape indices, which is usually achieved by manual contouring of the endo- and epicardial. However, this process subjected to inter and intra-observer variability, and it is a time-consuming and tedious task. In this article, we propose a spatio-temporal multi-task learning approach to obtain a complete set of measurements quantifying cardiac LV morphology, regional-wall thickness (RWT), and additionally detecting the cardiac phase cycle (systole and diastole) for a given 3D Cine-magnetic resonance (MR) image sequence. We first segment cardiac LVs using an encoder-decoder network and then introduce a multitask framework to regress 11 LV indices and classify the cardiac phase, as parallel tasks during model optimization. The proposed deep learning model is based on the 3D spatio-temporal convolutions, which extract spatial and temporal features from MR images. We demonstrate the efficacy of the proposed method using cine-MR sequences of 145 subjects and comparing the performance with other state-of-the-art quantification methods. The proposed method obtained high prediction accuracy, with an average mean absolute error (MAE) of 129 mm2, 1.23 mm, 1.76 mm, Pearson correlation coefficient (PCC) of 96.4%, 87.2%, and 97.5% for LV and myocardium (Myo) cavity regions, 6 RWTs, 3 LV dimensions, and an error rate of 9.0% for phase classification. The experimental results highlight the robustness of the proposed method, despite varying degrees of cardiac morphology, image appearance, and low contrast in the cardiac MR sequences. Sulaiman Vesal, Mingxuan Gu, Andreas K. Maier, Nishant Ravikumar |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | PMS-GAN: Parallel Multi-Stream Generative Adversarial Network for Multi-Material Decomposition in Spectral Computed TomographyabstractSpectral computed tomography is able to provide quantitative information on the scanned object and enables material decomposition. Traditional projection-based material decomposition methods suffer from the nonlinearity of the imaging system, which limits the decomposition accuracy. Inspired by the generative adversarial network, we proposed a novel parallel multi-stream generative adversarial network (PMS-GAN) to perform projection-based multi-material decomposition in spectral computed tomography. By designing the differential map and incorporating the adversarial network into loss function, the decomposition accuracy was significantly improved with robust performance. The proposed network was quantitatively evaluated by both simulation and experimental study. The results show that PMS-GAN outperformed the reference methods with certain robustness. Compared with Pix2pix-GAN, PMS-GAN increased the structural similarity index by 172% on the contrast agent Ultravist370, 11% on bones, and 71% on bone marrow, respectively, in a simulated test scenario. In an experimental test scenario, 9% and 38% improvements of the structural similarity index on the biopsy needle and on a torso phantom were observed, respectively. The proposed network demonstrates its capability of multi-material decomposition and has certain potential toward clinical applications. Mufeng Geng, Zifeng Tian, Yunfei You, Ximeng Feng, Yan Xia 0002, Qiushi Ren, Xiangxi Meng 0001, Andreas K. Maier, Yanye Lu |
IEEE Trans. Medical Imaging | 10 |
| 2021 | Deep Learning-Based ECG-Free Cardiac Navigation for Multi-Dimensional and Motion-Resolved Continuous Magnetic Resonance ImagingabstractFor the clinical assessment of cardiac vitality, time-continuous tomographic imaging of the heart is used. To further detect e.g., pathological tissue, multiple imaging contrasts enable a thorough diagnosis using magnetic resonance imaging (MRI). For this purpose, time-continous and multi-contrast imaging protocols were proposed. The acquired signals are binned using navigation approaches for a motion-resolved reconstruction. Mostly, external sensors such as electrocardiograms (ECG) are used for navigation, leading to additional workflow efforts. Recent sensor-free approaches are based on pipelines requiring prior knowledge, e.g., typical heart rates. We present a sensor-free, deep learning-based navigation that diminishes the need for manual feature engineering or the necessity of prior knowledge compared to previous works. A classifier is trained to estimate the R-wave timepoints in the scan directly from the imaging data. Our approach is evaluated on 3-D protocols for continuous cardiac MRI, acquired in-vivo and free-breathing with single or multiple imaging contrasts. We achieve an accuracy of > 98% on previously unseen subjects, and a well comparable image quality with the state-of-the-art ECG-based reconstruction. Our method enables an ECG-free workflow for continuous cardiac scans with simultaneous anatomic and functional imaging with multiple contrasts. It can be potentially integrated without adapting the sampling scheme to other continuous sequences by using the imaging data for navigation and reconstruction. Elisabeth Hoppe, Jens Wetzl, Seung Su Yoon, Mario Bacher, Philipp Roser, Bernhard Stimpel, Alexander Preuhs, Andreas K. Maier |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Data Extrapolation From Learned Prior Images for Truncation Correction in Computed TomographyabstractData truncation is a common problem in computed tomography (CT). Truncation causes cupping artifacts inside the field-of-view (FOV) and anatomical structures missing outside the FOV. Deep learning has achieved impressive results in CT reconstruction from limited data. However, its robustness is still a concern for clinical applications. Although the image quality of learning-based compensation schemes may be inadequate for clinical diagnosis, they can provide prior information for more accurate extrapolation than conventional heuristic extrapolation methods. With extrapolated projection, a conventional image reconstruction algorithm can be applied to obtain a final reconstruction. In this work, a general plug-and-play (PnP) method for truncation correction is proposed based on this idea, where various deep learning methods and conventional reconstruction algorithms can be plugged in. Such a PnP method integrates data consistency for measured data and learned prior image information for truncated data. This shows to have better robustness and interpretability than deep learning only. To demonstrate the efficacy of the proposed PnP method, two state-of-the-art deep learning methods, FBPConvNet and Pix2pixGAN, are investigated for truncation correction in cone-beam CT in noise-free and noisy cases. Their robustness is evaluated by showing false negative and false positive lesion cases. With our proposed PnP method, false lesion structures are corrected for both deep learning methods. For FBPConvNet, the root-mean-square error (RMSE) inside the FOV can be improved from 92HU to around 30HU by PnP in the noisy case. Pix2pixGAN solely achieves better image quality than FBPConvNet solely for truncation correction in general. PnP further improves the RMSE inside the FOV from 42HU to around 27HU for Pix2pixGAN. The efficacy of PnP is also demonstrated on real clinical head data. Yixing Huang, Alexander Preuhs, Michael Manhart 0001, Günter Lauritsch, Andreas K. Maier |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Weakly Supervised Deep Learning-Based Optical Coherence Tomography AngiographyabstractOptical coherence tomography angiography (OCTA) is a promising imaging modality for microvasculature studies. Deep learning networks have been widely applied in the field of OCTA reconstruction, benefiting from its powerful mapping capability among images. However, these existing deep learning-based methods depend on high-quality labels, which are hard to acquire considering imaging hardware limitations and practical data acquisition conditions. In this article, we proposed an unprecedented weakly supervised deep learning-based pipeline for OCTA reconstruction task, in the absence of high-quality training labels. The proposed pipeline was investigated on an in vivo animal dataset and a human eye dataset by a cross-validation strategy. Compared with supervised learning approaches, the proposed approach demonstrated similar or even better performance in the OCTA reconstruction task. These investigations indicate that the proposed weakly supervised learning strategy is well capable of performing OCTA reconstruction, and has a certain potential towards clinical applications. Zhiyu Huang, Bin Qiu, Xiangxi Meng 0001, Yunfei You, Mufeng Geng, Gangjun Liu, Chuanqing Zhou, Andreas K. Maier, Qiushi Ren, Yanye Lu |
IEEE Trans. Medical Imaging | 11 |
| 2021 | X-Ray Scatter Estimation Using Deep Splines
Philipp Roser, Annette Birkhold, Alexander Preuhs, Christopher Syben, Lina Felsner, Elisabeth Hoppe, Norbert Strobel, Markus Kowarschik, Rebecca Fahrig, Andreas K. Maier |
IEEE Trans. Medical Imaging | 10 |
| 2021 | Adapt Everywhere: Unsupervised Adaptation of Point-Clouds and Entropy Minimization for Multi-Modal Cardiac Image SegmentationabstractDeep learning models are sensitive to domain shift phenomena. A model trained on images from one domain cannot generalise well when tested on images from a different domain, despite capturing similar anatomical structures. It is mainly because the data distribution between the two domains is different. Moreover, creating annotation for every new modality is a tedious and time-consuming task, which also suffers from high inter- and intra- observer variability. Unsupervised domain adaptation (UDA) methods intend to reduce the gap between source and target domains by leveraging source domain labelled data to generate labels for the target domain. However, current state-of-the-art (SOTA) UDA methods demonstrate degraded performance when there is insufficient data in source and target domains. In this paper, we present a novel UDA method for multi-modal cardiac image segmentation. The proposed method is based on adversarial learning and adapts network features between source and target domain in different spaces. The paper introduces an end-to-end framework that integrates: a) entropy minimization, b) output feature space alignment and c) a novel point-cloud shape adaptation based on the latent features learned by the segmentation model. We validated our method on two cardiac datasets by adapting from the annotated source domain, bSSFP-MRI (balanced Steady-State Free Procession-MRI), to the unannotated target domain, LGE-MRI (Late-gadolinium enhance-MRI), for the multi-sequence dataset; and from MRI (source) to CT (target) for the cross-modality dataset. The results highlighted that by enforcing adversarial learning in different parts of the network, the proposed method delivered promising performance, compared to other SOTA methods. Sulaiman Vesal, Mingxuan Gu, Ronak Kosti, Andreas K. Maier, Nishant Ravikumar |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Re-Ranking for Writer Identification and Writer Retrieval
Simon Jordan, Mathias Seuret, Pavel Král, Ladislav Lenc, Jirí Martínek, Barbara Wiermann, Tobias Schwinger, Andreas K. Maier, Vincent Christlein |
DAS | 8 |
| 2020 | The Notary in the Haystack - Countering Class Imbalance in Document Processing with CNNs
Martin Leipert, Georg Vogeler, Mathias Seuret, Andreas K. Maier, Vincent Christlein |
DAS | 4 |
| 2020 | The Effect of Data Augmentation on Classification of Atrial Fibrillation in Short Single-Lead ECG Signals Using Deep Neural NetworksabstractCardiovascular diseases are the most common cause of mortality worldwide. Detection of atrial fibrillation (AF) in the asymptomatic stage can help prevent strokes. It also improves clinical decision making through the delivery of suitable treatment such as, anticoagulant therapy, in a timely manner. The clinical significance of such early detection of AF in electrocardiogram (ECG) signals has inspired numerous studies in recent years, of which many aim to solve this task by leveraging machine learning algorithms. ECG datasets containing AF samples, however, usually suffer from severe class imbalance, which if unaccounted for, affects the performance of classification algorithms. Data augmentation is a popular solution to tackle this problem. In this study, we investigate the impact of various data augmentation algorithms, e.g., oversampling, Gaussian Mixture Models (GMMs) and Generative Adversarial Networks (GANs), on solving the class imbalance problem. These algorithms are quantitatively and qualitatively evaluated, compared and discussed in detail. The results show that deep learning-based AF signal classification methods benefit more from data augmentation using GANs and GMMs, than oversampling. Furthermore, the GAN results in circa 3% better AF classification accuracy in average while performing comparably to the GMM in terms of f1-score. Faezeh Nejati Hatamian, Nishant Ravikumar, Sulaiman Vesal, Felix P. Kemeth, Matthias Struck, Andreas K. Maier |
ICASSP | 6 |
| 2020 | CLCNET: Deep Learning-Based Noise Reduction for Hearing aids using Complex Linear CodingabstractNoise reduction is an important part of modern hearing aids and is included in most commercially available devices. Deep learning-based state-of-the-art algorithms, however, either do not consider real-time and frequency resolution constrains or result in poor quality under very noisy conditions.To improve monaural speech enhancement in noisy environments, we propose CLCNet, a framework based on complex valued linear coding. First, we define complex linear coding (CLC) motivated by linear predictive coding (LPC) that is applied in the complex frequency domain. Second, we propose a framework that incorporates complex spectrogram input and coefficient output. Third, we define a parametric normalization for complex valued spectrograms that complies with low-latency and on-line processing.Our CLCNet was evaluated on a mixture of the EUROM database and a real-world noise dataset recorded with hearing aids and compared to traditional real-valued Wiener-Filter gains. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Marc Aubreville, Andreas K. Maier |
ICASSP | 5 |
| 2020 | ICFHR 2020 Competition on Image Retrieval for Historical Handwritten FragmentsabstractThis competition succeeds upon a line of competitions for writer and style analysis of historical document images. In particular, we investigate the performance of large-scale retrieval of historical document fragments in terms of style and writer identification. The analysis of historic fragments is a difficult challenge commonly solved by trained humanists. In comparison to previous competitions, we make the results more meaningful by addressing the issue of sample granularity and moving from writer to page fragment retrieval. The two approaches, style and author identification, provide information on what kind of information each method makes better use of and indirectly contribute to the interpretability of the participating method. Therefore, we created a large dataset consisting of more than 120 000 fragments. Although the most teams submitted methods based on convolutional neural networks, the winning entry achieves an mAP below 40 %. Mathias Seuret, Anguelos Nicolaou, Andreas K. Maier, Vincent Christlein, Dominique Stutzmann |
ICFHR | 3 |
| 2020 | Art2Contour: Salient Contour Detection in Artworks Using Generative Adversarial NetworksabstractArtists or art workshops often reuse their motifs directly or in a slightly amended form. To allow a better comparison of these artworks, salient contours are extracted that reduce them to the most important lines or boundaries. For this task, we propose a generative adversarial network (GAN) based approach to learn the mapping from artwork images to contour drawings in a supervised manner. We introduce the combination of multiple regression task losses to encourage the learning of salient contours. For the evaluation, we created a dataset of high-resolution prints and paintings and corresponding annotated ground truth drawings. We show that our method visually and quantitatively outperforms competing methods in contour detection on prints and paintings. Aline Sindel, Andreas K. Maier, Vincent Christlein |
ICIP | 2 |
| 2020 | ORCA-CLEAN: A Deep Denoising Toolkit for Killer Whale CommunicationabstractIn bioacoustics, passive acoustic monitoring of animals living in the wild, both on land and underwater, leads to large data archives characterized by a strong imbalance between recorded animal sounds and ambient noises. Bioacoustic datasets suffer extremely from such large noise-variety, caused by a multitude of external influences and changing environmental conditions over years. This leads to significant deficiencies/problems concerning the analysis and interpretation of animal vocalizations by biologists and machine-learning algorithms. To counteract such huge noise diversity, it is essential to develop a denoising procedure enabling automated, efficient, and robust data enhancement. However, a fundamental problem is the lack of clean/denoised ground-truth samples. The current work is the first presenting a fully-automated deep denoising approach for bioacoustics, not requiring any clean ground-truth, together with one of the largest data archives recorded on killer whales (Orcinus Orca) – the Orchive. Therefor, an approach, originally developed for image restoration, known as Noise2Noise (N2N), was transferred to the field of bioacoustics, and extended by using automatic machine-generated binary masks as additional network attention mechanism. Besides a significant cross-domain signal enhancement, our previous results regarding supervised orca/noise segmentation and orca call type identification were outperformed by applying ORCACLEAN as additional data preprocessing/enhancement step Christian Bergler, Manuel Schmitt, Andreas K. Maier, Simeon Smeele, Volker Barth, Elmar Nöth |
INTERSPEECH | 3 |
| 2020 | Lightweight Online Noise Reduction on Embedded Devices Using Hierarchical Recurrent Neural NetworksabstractDeep-learning based noise reduction algorithms have proven their success especially for non-stationary noises, which makes it desirable to also use them for embedded devices like hearing aids (HAs). This, however, is currently not possible with state-of-the-art methods. They either require a lot of parameters and computational power and thus are only feasible using modern CPUs. Or they are not suitable for online processing, which requires constraints like low-latency by the filter bank and the algorithm itself. In this work, we propose a mask-based noise reduction approach. Using hierarchical recurrent neural networks, we are able to drastically reduce the number of neurons per layer while including temporal context via hierarchical connections. This allows us to optimize our model towards a minimum number of parameters and floating-point operations (FLOPs), while preserving noise reduction quality compared to previous work. Our smallest network contains only 5k parameters, which makes this algorithm applicable on embedded devices. We evaluate our model on a mixture of EUROM and a real-world noise database and report objective metrics on unseen noise. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Pascal Zobel, Andreas K. Maier |
INTERSPEECH | 5 |
| 2020 | Move Over There: One-Click Deformation Correction for Image Fusion During Endovascular Aortic Repair
Katharina Breininger, Marcus Pfister, Markus Kowarschik, Andreas K. Maier |
MICCAI (4) | 4 |
| 2020 | Automatic CAD-RADS Scoring Using Deep Learning
Felix Denzinger, Michael Wels, Katharina Breininger, Mehmet Akif Gülsün, Max Schöbinger, Florian André, Sebastian Buß 0001, Johannes Görich, Michael Sühling, Andreas K. Maier |
MICCAI (6) | 10 |
| 2020 | Contour-Based Bone Axis Detection for X-Ray Guided Surgery on the Knee
Florian Kordon, Andreas K. Maier, Benedict Swartman, Maxim Privalov, Jan Siad El Barbari, Holger Kunze |
MICCAI (6) | 2 |
| 2020 | Inertial Measurements for Motion Compensation in Weight-Bearing Cone-Beam CT of the Knee
Jennifer Maier, Marlies Nitschke, Jang Hwan Choi 0001, Garry Gold, Rebecca Fahrig, Björn M. Eskofier, Andreas K. Maier |
MICCAI (3) | 7 |
| 2020 | Are Fast Labeling Methods Reliable? A Case Study of Computer-Aided Expert Annotations on Microscopy Slides
Christian Marzahl, Christof Bertram, Marc Aubreville, Anne Petrick, Kristina Weiler, Agnes C. Gläsel, Marco Fragoso, Sophie Merz, Florian Bartenschlager, Judith Hoppe, Alina Langenhagen, Anne-Katherine Jasensky, Jörn Voigt, Robert Klopfleisch, Andreas K. Maier |
MICCAI (1) | 15 |
| 2020 | JBFnet - Low Dose CT Denoising by Trainable Joint Bilateral Filtering
Mayank Patwari, Ralf Gutjahr, Rainer Raupach, Andreas K. Maier |
MICCAI (2) | 4 |
| 2020 | Simultaneous Estimation of X-Ray Back-Scatter and Forward-Scatter Using Multi-task Learning
Philipp Roser, Xia Zhong, Annette Birkhold, Alexander Preuhs, Christopher Syben, Elisabeth Hoppe, Norbert Strobel, Markus Kowarschik, Rebecca Fahrig, Andreas K. Maier |
MICCAI (2) | 10 |
| 2020 | Automatic Detection of Free Intra-abdominal Air in Computed Tomography
Oliver Taubmann, Jingpeng Li 0004, Felix Denzinger, Eva Eibenberger, Felix C. Müller, Mathias W. Brejnebøl, Andreas K. Maier |
MICCAI (2) | 7 |
| 2020 | Automatic Plane Adjustment of Orthopedic Intraoperative Flat Panel Detector CT-Volumes
Celia Martín Vicario, Florian Kordon, Felix Denzinger, Markus Weiten, Sarina Thomas, Lisa Kausch, Jochen Franke, Holger Keil, Andreas K. Maier, Holger Kunze |
MICCAI (2) | 9 |
| 2020 | Dual-Mode Training with Style Control and Quality Enhancement for Road Image Domain AdaptationabstractDealing properly with different viewing conditions remains a key challenge for computer vision in autonomous driving. Domain adaptation has opened new possibilities for data augmentation, translating arbitrary road scene images into different environmental conditions. Although multimodal concepts have demonstrated the capability to separate content and style, we find that existing methods fail to reproduce scenes in the exact appearance given by a reference image. In this paper, we address the aforementioned problem by introducing a style alignment loss between output and reference image. We integrate this concept into a multimodal unsupervised image-to-image translation model with a novel dual-mode training process and additional adversarial losses. Focusing on road scene images, we evaluate our model in various aspects including visual quality and feature matching. Our experiments reveal that we are able to significantly improve both style alignment and image quality in different viewing conditions. Adapting concepts from neural style transfer, our new training approach allows to control the output of multimodal domain adaptation, making it possible to generate arbitrary scenes and viewing conditions for data augmentation. Moritz Venator, Fengyi Shen, Selcuk Aklanoglu, Erich Bruns, Klaus Diepold, Andreas K. Maier |
WACV | 6 |
| 2020 | Toward Bridging the Simulated-to-Real Gap: Benchmarking Super-Resolution on Real DataabstractCapturing ground truth data to benchmark super-resolution (SR) is challenging. Therefore, current quantitative studies are mainly evaluated on simulated data artificially sampled from ground truth images. We argue that such evaluations overestimate the actual performance of SR methods compared to their behavior on real images. Toward bridging this simulated-to-real gap, we introduce the Super-Resolution Erlangen (SupER) database, the first comprehensive laboratory SR database of all-real acquisitions with pixel-wise ground truth. It consists of more than 80k images of 14 scenes combining different facets: CMOS sensor noise, real sampling at four resolution levels, nine scene motion types, two photometric conditions, and lossy video coding at five levels. As such, the database exceeds existing benchmarks by an order of magnitude in quality and quantity. This paper also benchmarks 19 popular single-image and multi-frame algorithms on our data. The benchmark comprises a quantitative study by exploiting ground truth data and qualitative evaluations in a large-scale observer study. We also rigorously investigate agreements between both evaluations from a statistical perspective. One interesting result is that top-performing methods on simulated data may be surpassed by others on real data. Our insights can spur further algorithm development, and the publicy available dataset can foster future evaluations. Thomas Köhler 0004, Michel Bätz, Farzad Naderi, André Kaup, Andreas K. Maier, Christian Riess |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Appearance Learning for Image-Based Motion Estimation in TomographyabstractIn tomographic imaging, anatomical structures are reconstructed by applying a pseudo-inverse forward model to acquired signals. Geometric information within this process is usually depending on the system setting only, i.e., the scanner position or readout direction. Patient motion therefore corrupts the geometry alignment in the reconstruction process resulting in motion artifacts. We propose an appearance learning approach recognizing the structures of rigid motion independently from the scanned object. To this end, we train a siamese triplet network to predict the reprojection error (RPE) for the complete acquisition as well as an approximate distribution of the RPE along the single views from the reconstructed volume in a multi-task learning approach. The RPE measures the motion-induced geometric deviations independent of the object based on virtual marker positions, which are available during training. We train our network using 27 patients and deploy a 21-4-2 split for training, validation and testing. In average, we achieve a residual mean RPE of 0.013mm with an inter-patient standard deviation of 0.022mm. This is twice the accuracy compared to previously published results. In a motion estimation benchmark the proposed approach achieves superior results in comparison with two state-of-the-art measures in nine out of twelve experiments. The clinical applicability of the proposed method is demonstrated on a motion-affected clinical dataset. Alexander Preuhs, Michael Manhart 0001, Philipp Roser, Elisabeth Hoppe, Yixing Huang, Marios Nikos Psychogios, Markus Kowarschik, Andreas K. Maier |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Learning an Attention Model for Robust 2-D/3-D Registration Using Point-To-Plane CorrespondencesabstractMinimally invasive procedures rely on image guidance for navigation at the operation site to avoid large surgical incisions. X-ray images are often used for guidance, but important structures may be not well visible. These structures can be overlaid from pre-operative 3-D images and accurate alignment can be established using 2-D/3-D registration. Registration based on the point-to-plane correspondence model was recently proposed and shown to achieve state-of-the-art performance. However, registration may still fail in challenging cases due to a large portion of outliers. In this paper, we describe a learning-based correspondence weighting scheme to improve the registration performance. By learning an attention model, inlier correspondences get higher attention in the motion estimation while the outlier correspondences are suppressed. Instead of using per-correspondence labels, our objective function allows to train the model directly by minimizing the registration error. We demonstrate a highly increased robustness, e.g. increasing the success rate from 84.9% to 97.0% for spine registration. In contrast to previously proposed learning-based methods, we also achieve a high accuracy of around 0.5mm mean re-projection distance. In addition, our method requires a relatively small amount of training data, is able to learn from simulated data, and generalizes to images with additional structures which are not present during training. Furthermore, a single model can be trained for both, different views and different anatomical structures. Roman Schaffert, Jian Wang 0009, Peter Fischer 0001, Anja Borsdorf, Andreas K. Maier |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Robust Multi-View 2-D/3-D Registration Using Point-To-Plane Correspondence ModelabstractIn minimally invasive procedures, the clinician relies on image guidance to observe and navigate the operation site. In order to show structures which are not visible in the live X-ray images, such as vessels or planning annotations, X-ray images can be augmented with pre-operatively acquired images. Accurate image alignment is needed and can be provided by 2-D/3-D registration. In this paper, a multi-view registration method based on the point-to-plane correspondence model is proposed. The correspondence model is extended to be independent of the used camera coordinates and different multi-view registration schemes are introduced and compared. Evaluation is performed for a wide range of clinically relevant registration scenarios. We show for different applications that registration using correspondences from both views simultaneously provides accurate and robust registration, while the performance of the other schemes varies considerably. Our method also outperforms the state-of-the-art method for cerebral angiography registration, achieving a capture range of 18 mm and an accuracy of 0.22±0.07 mm. Furthermore, investigations on the minimum angle between the views are performed in order to provide accurate and robust registration, while minimizing the obstruction to the clinical workflow. We show that small angles around 30° are sufficient to provide reliable registration results. Roman Schaffert, Jian Wang 0009, Peter Fischer 0001, Andreas K. Maier, Anja Borsdorf |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Multi-Modal Deep Guided Filtering for Comprehensible Medical Image ProcessingabstractDeep learning-based image processing is capable of creating highly appealing results. However, it is still widely considered as a "blackbox" transformation. In medical imaging, this lack of comprehensibility of the results is a sensitive issue. The integration of known operators into the deep learning environment has proven to be advantageous for the comprehensibility and reliability of the computations. Consequently, we propose the use of the locally linear guided filter in combination with a learned guidance map for general purpose medical image processing. The output images are only processed by the guided filter while the guidance map can be trained to be task-optimal in an end-to-end fashion. We investigate the performance based on two popular tasks: image super resolution and denoising. The evaluation is conducted based on pairs of multi-modal magnetic resonance imaging and cross-modal computed tomography and magnetic resonance imaging datasets. For both tasks, the proposed approach is on par with state-of-the-art approaches. Additionally, we can show that the input image's content is almost unchanged after the processing which is not the case for conventional deep learning approaches. On top, the proposed pipeline offers increased robustness against degraded input as well as adversarial attacks. Bernhard Stimpel, Christopher Syben, Franziska Schirrmacher, Philip Hoelter, Arnd Dörfler, Andreas K. Maier |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Known Operator Learning Enables Constrained Projection Geometry Conversion: Parallel to Cone-Beam for Hybrid MR/X-Ray ImagingabstractX-ray imaging is a wide-spread real-time imaging technique. Magnetic Resonance Imaging (MRI) offers a multitude of contrasts that offer improved guidance to interventionalists. As such simultaneous real-time acquisition and overlay would be highly favorable for image-guided interventions, e.g., in stroke therapy. One major obstacle in this setting is the fundamentally different acquisition geometry. MRI k -space sampling is associated with parallel projection geometry, while the X-ray acquisition results in perspective distorted projections. The classical rebinning methods to overcome this limitation inherently suffers from a loss of resolution. To counter this problem, we present a novel rebinning algorithm for parallel to cone-beam conversion. We derive a rebinning formula that is then used to find an appropriate deep neural network architecture. Following the known operator learning paradigm, the novel algorithm is mapped to a neural network with differentiable projection operators enabling data-driven learning of the remaining unknown operators. The evaluation aims in two directions: First, we give a profound analysis of the different hypotheses to the unknown operator and investigate the influence of numerical training data. Second, we evaluate the performance of the proposed method against the classical rebinning approach. We demonstrate that the derived network achieves better results than the baseline method and that such operators can be trained with simulated data without losing their generality making them applicable to real data without the need for retraining or transfer learning. Christopher Syben, Bernhard Stimpel, Philipp Roser, Arnd Dörfler, Andreas K. Maier |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Evaluation of MRI to Ultrasound Registration Methods for Brain Shift Correction: The CuRIOUS2018 ChallengeabstractIn brain tumor surgery, the quality and safety of the procedure can be impacted by intra-operative tissue deformation, called brain shift. Brain shift can move the surgical targets and other vital structures such as blood vessels, thus invalidating the pre-surgical plan. Intra-operative ultrasound (iUS) is a convenient and cost-effective imaging tool to track brain shift and tumor resection. Accurate image registration techniques that update pre-surgical MRI based on iUS are crucial but challenging. The MICCAI Challenge 2018 for Correction of Brain shift with Intra-Operative UltraSound (CuRIOUS2018) provided a public platform to benchmark MRI-iUS registration algorithms on newly released clinical datasets. In this work, we present the data, setup, evaluation, and results of CuRIOUS 2018, which received 6 fully automated algorithms from leading academic and industrial research groups. All algorithms were first trained with the public RESECT database, and then ranked based on a test dataset of 10 additional cases with identical data curation and annotation protocols as the RESECT database. The article compares the results of all participating teams and discusses the insights gained from the challenge, as well as future work. Yiming Xiao 0001, Andreas K. Maier, Wolfgang Wein, Roozbeh Shams, Samuel Kadoury, David Drobny, Marc Modat, Ingerid Reinertsen, Hassan Rivaz, Matthieu Chabanas, Maryse Fortin, Inês Machado, Yangming Ou, Mattias P. Heinrich, Julia A. Schnabel, Xia Zhong |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Fast and Robust Detection of Solar Modules in Electroluminescence Images
Mathis Hoffmann, Bernd Doll, Florian Talkenberg, Christoph Brabec, Andreas K. Maier, Vincent Christlein |
CAIP (2) | 5 |
| 2019 | Segmentation, Classification, and Visualization of Orca Calls Using Deep LearningabstractAudiovisual media are increasingly used to study the communication and behavior of animal groups, e.g. by placing microphones in the animals habitat resulting in huge datasets with only a small amount of animal interactions. The Orcalab has recorded orca whales since 1973 using stationary underwater hydrophones and made it publicly available on the Orchive. There exist over 15 000 manually extracted orca/noise annotations and about 20 000 h unseen audio data. To analyze the behavior and communication of killer whales we need to interpret the different call types. In this work, we present a two-stage classification approach using the labeled call/noise files and a few labeled call-type files. Results indicate a reliable accuracy of 95.0 % for call segmentation and 87 % for classification of 12 call classes. We further visualize the learned orca call representations in the convolutional neural network (CNN) activations to explain the potential of CNN based recognition for bioaccousitc signals. Hendrik Schröter, Elmar Nöth, Andreas K. Maier, Rachael Cheng, Volker Barth, Christian Bergler |
ICASSP | 3 |
| 2019 | Estimating the Fundamental Matrix Without Point Correspondences With Application to Transmission ImagingabstractWe present a general method to estimate the fundamental matrix from a pair of images under perspective projection without the need for image point correspondences. Our method is particularly well-suited for transmission imaging, where state-of-the-art feature detection and matching approaches generally do not perform well. Estimation of the fundamental matrix plays a central role in auto-calibration methods for reflection imaging. Such methods are currently not applicable to transmission imaging. Furthermore, our method extends an existing technique proposed for reflection imaging which potentially avoids the outlier-prone feature matching step from an orthographic projection model to a perspective model. Our method exploits the idea that under a linear attenuation model line integrals along corresponding epipolar lines are equal if we compute their derivatives in orthogonal direction to their common epipolar plane. We use the fundamental matrix to parametrize this equality. Our method estimates the matrix by formulating a non-convex optimization problem, minimizing an error in our measurement of this equality. We believe this technique will enable the application of the large body of work on image-based camera pose estimation to transmission imaging leading to more accurate and more general motion compensation and auto-calibration algorithms, particularly in medical X-ray and Computed Tomography imaging. Tobias Würfl, André Aichert, Nicole Maass, Frank Dennerlein, Andreas K. Maier |
ICCV | 5 |
| 2019 | ICDAR 2019 Competition on Image Retrieval for Historical Handwritten DocumentsabstractThis competition investigates the performance of large-scale retrieval of historical document images based on writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing a total of 20 000 document images representing about 10 000 writers, divided in three types: writers of (i) manuscript books, (ii) letters, (iii) charters and legal documents. We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as writer retrieval. The most teams submitted traditional methods not using deep learning techniques. The competition results show that a combination of methods is outperforming single methods. Furthermore, letters are much more difficult to retrieve than manuscripts. Vincent Christlein, Anguelos Nicolaou, Mathias Seuret, Dominique Stutzmann, Andreas K. Maier |
ICDAR | 5 |
| 2019 | Deep Generalized Max PoolingabstractGlobal pooling layers are an essential part of Convolutional Neural Networks (CNN). They are used to aggregate activations of spatial locations to produce a fixed-size vector in several state-of-the-art CNNs. Global average pooling or global max pooling are commonly used for converting convolutional features of variable size images to a fix-sized embedding. However, both pooling layer types are computed spatially independent: each individual activation map is pooled and thus activations of different locations are pooled together. In contrast, we propose Deep Generalized Max Pooling that balances the contribution of all activations of a spatially coherent region by re-weighting all descriptors so that the impact of frequent and rare ones is equalized. We show that this layer is superior to both average and max pooling on the classification of Latin medieval manuscripts (CLAMM'16, CLAMM'17), as well as writer identification (Historical-WI'17). Vincent Christlein, Lukas Spranger, Mathias Seuret, Anguelos Nicolaou, Pavel Král, Andreas K. Maier |
ICDAR | 6 |
| 2019 | Weakly Supervised Segmentation of Cracks on Solar Cells Using Normalized Lp NormabstractPhotovoltaic is one of the most important renewable energy sources for dealing with world-wide steadily increasing energy consumption. This raises the demand for fast and scalable automatic quality management during production and operation. However, the detection and segmentation of cracks on electroluminescence (EL) images of mono- or polycrystalline solar modules is a challenging task. In this work, we propose a weakly supervised learning strategy that only uses image-level annotations to obtain a method that is capable of segmenting cracks on EL images of solar cells. We use a modified ResNet-50 to derive a segmentation from network activation maps. We use defect classification as a surrogate task to train the network. To this end, we apply normalized Lpnormalization to aggregate the activation maps into single scores for classification. In addition, we provide a study how different parameterizations of the normalized Lplayer affect the segmentation performance. This approach shows promising results for the given task. However, we think that the method has the potential to solve other weakly supervised segmentation problems as well. Martin Mayr, Mathis Hoffmann, Andreas K. Maier, Vincent Christlein |
ICIP | 3 |
| 2019 | Deep Learning for Orca Call Type Identification - A Fully Unsupervised Approach
Christian Bergler, Manuel Schmitt, Rachael Cheng, Andreas K. Maier, Volker Barth, Elmar Nöth |
INTERSPEECH | 4 |
| 2019 | Analysis by Adversarial Synthesis - A Novel Approach for Speech VocodingabstractClassical parametric speech coding techniques provide a compact representation for speech signals. This affords a very low transmission rate but with a reduced perceptual quality of the reconstructed signals. Recently, autoregressive deep generative models such as WaveNet and SampleRNN have been used as speech vocoders to scale up the perceptual quality of the reconstructed signals without increasing the coding rate. However, such models suffer from a very slow signal generation mechanism due to their sample-by-sample modelling approach. In this work, we introduce a new methodology for neural speech vocoding based on generative adversarial networks (GANs). A fake speech signal is generated from a very compressed representation of the glottal excitation using conditional GANs as a deep generative model. This fake speech is then refined using the LPC parameters of the original speech signal to obtain a natural reconstruction. The reconstructed speech waveforms based on this approach show a higher perceptual quality than the classical vocoder counterparts according to subjective and objective evaluation scores for a dataset of 30 male and female speakers. Moreover, the usage of GANs enables to generate signals in one-shot compared to autoregressive generative models. This makes GANs promising for exploration to implement high-quality neural vocoders. Ahmed Mustafa, Arijit Biswas, Christian Bergler, Julia Schottenhamml, Andreas K. Maier |
INTERSPEECH | 5 |
| 2019 | Coronary Artery Plaque Characterization from CCTA Scans Using Deep Learning and Radiomics
Felix Denzinger, Michael Wels, Nishant Ravikumar, Katharina Breininger, Anika Reidelshöfer, Joachim Eckert, Michael Sühling, Axel Schmermund, Andreas K. Maier |
MICCAI (4) | 9 |
| 2019 | A Divide-and-Conquer Approach Towards Understanding Deep Networks
Weilin Fu, Katharina Breininger, Roman Schaffert, Nishant Ravikumar, Andreas K. Maier |
MICCAI (1) | 5 |
| 2019 | RinQ Fingerprinting: Recurrence-Informed Quantile Networks for Magnetic Resonance Fingerprinting
Elisabeth Hoppe, Florian Thamm, Gregor Körzdörfer, Christopher Syben, Franziska Schirrmacher, Mathias Nittka, Josef Pfeuffer, Heiko Meyer, Andreas K. Maier |
MICCAI (3) | 9 |
| 2019 | Multi-task Localization and Segmentation for X-Ray Guided Planning in Knee Surgery
Florian Kordon, Peter Fischer 0001, Maxim Privalov, Benedict Swartman, Marc Schnetzke, Jochen Franke, Ruxandra Lasowski, Andreas K. Maier, Holger Kunze |
MICCAI (6) | 8 |
| 2019 | Learning to Avoid Poor Images: Towards Task-aware C-arm Cone-beam CT Trajectories
Jan-Nico Zaech, Cong Gao 0003, Bastian Bier, Russell H. Taylor, Andreas K. Maier, Nassir Navab, Mathias Unberath |
MICCAI (5) | 5 |
| 2019 | Toward analyzing mutual interference on infrared-enabled depth cameras
Lucas Adams Seewald, Vinicius Facco Rodrigues, Malte Ollenschläger, Rodolfo Stoffel Antunes, Cristiano André da Costa, Rodrigo da Rosa Righi, Luiz Gonzaga 0001, Andreas K. Maier, Björn M. Eskofier, Rebecca Fahrig |
Comput. Vis. Image Underst. | 8 |
| 2019 | Multi-Scale Deep Reinforcement Learning for Real-Time 3D-Landmark Detection in CT ScansabstractRobust and fast detection of anatomical structures is a prerequisite for both diagnostic and interventional medical image analysis. Current solutions for anatomy detection are typically based on machine learning techniques that exploit large annotated image databases in order to learn the appearance of the captured anatomy. These solutions are subject to several limitations, including the use of suboptimal feature engineering techniques and most importantly the use of computationally suboptimal search-schemes for anatomy detection. To address these issues, we propose a method that follows a new paradigm by reformulating the detection problem as a behavior learning task for an artificial agent. We couple the modeling of the anatomy appearance and the object search in a unified behavioral framework, using the capabilities of deep reinforcement learning and multi-scale image analysis. In other words, an artificial agent is trained not only to distinguish the target anatomical object from the rest of the body but also how to find the object by learning and following an optimal navigation path to the target object in the imaged volumetric space. We evaluated our approach on 1487 3D-CT volumes from 532 patients, totaling over 500,000 image slices and show that it significantly outperforms state-of-the-art solutions on detecting several anatomical structures with no failed cases from a clinical acceptance perspective, while also achieving a 20-30 percent higher detection accuracy. Most importantly, we improve the detection-speed of the reference methods by 2-3 orders of magnitude, achieving unmatched real-time performance on large 3D-CT scans. Florin C. Ghesu, Bogdan Georgescu, Yefeng Zheng 0001, Sasa Grbic, Andreas K. Maier, Joachim Hornegger, Dorin Comaniciu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Robust mixed one-bit compressive sensing
Xiaolin Huang, Yixing Huang, Lei Shi 0010, Andreas K. Maier, Ming Yan 0006 |
Signal Process. | 6 |
| 2018 | Learning to Recognize Abnormalities in Chest X-Rays with Location-Aware Dense Networks
Sebastian Gündel, Sasa Grbic, Bogdan Georgescu, Siqi Liu 0001, Andreas K. Maier, Dorin Comaniciu |
CIARP | 5 |
| 2018 | Encoding CNN Activations for Writer RecognitionabstractThe encoding of local features is an essential part for writer identification and writer retrieval. While CNN activations have already been used as local features in related works, the encoding of these features has attracted little attention so far. In this work, we compare the established VLAD encoding with triangulation embedding. We further investigate generalized max pooling as an alternative to sum pooling and the impact of decorrelation and Exemplar SVMs. With these techniques, we set new standards on two publicly available datasets (ICDAR13, KHATT). Vincent Christlein, Andreas K. Maier |
DAS | 2 |
| 2018 | Non-destructive Digitization of Soiled Historical Chinese Bamboo ScrollsabstractFor about 2000 years, no paper was used as a media in China but writings and drawings were captured on bamboo and wooden slips. Several slips were bound together with strips and rolled up to a scroll. The writings and drawings were either brushed or even carved into the wood. Those documents are very precious for culture inheritance and research, but due to aging processes, the discovered pieces are sometimes in a poor condition and also soiled. Because cleaning the slips is not only challenging but also writings could be erased, we developed a method to digitize such historical documents without the need of cleaning. We perform a 3-D X-ray micro-CT scan resulting in a 3-D volume of the complete document. With our approach, we were able to investigate the scroll without any manual labor (e.g. unwrapping or cleaning). We showed that the method also works for heavily soiled scrolls where nothing is readable with the naked eye. This can help conservators to store all writings before they may be erased by the cleaning process. Finally, we present a manual technique to virtually unwrap and post-process the documents resulting in a 2-D image of all bamboo slips. Daniel Stromer, Vincent Christlein, Andreas K. Maier, Patrick Zippert, Eric Helmecke, Tino Hausotte, Xiaolin Huang |
DAS | 3 |
| 2018 | Robust Blood Flow Velocity Estimation from 3D Rotational AngiographyabstractBlood flow velocity estimation techniques from 2D fluoroscopy and more recently rotational angiography images represent a topic of wide interest in various clinical research areas. In particular, they can be an important step towards patient-specific flow simulations. Additionally, it can be of diagnostic interest to evaluate volumetric blood flow in stenotic vessel segments; e.g. in a patient's brain. In this work, we present a robust optimization-based approach to estimate mean blood flow velocities from rotational digital subtraction angiography (DSA) images. Our method was extensively evaluated on 70 simulated datasets and 6 clinical datasets with MR phase contrast ground truth data. Our evaluation explores the limitations of image-based velocity estimation; i.e., measurements over short or small vessel segments. Overall, we were able to estimate the mean velocity with average errors as little as 4% for simulation studies, if the vessel segment is sufficiently long, and achieved results within the confines of the MR phase contrast ground truth data for our clinical data, with an average relative error to the centerline measurement of 9.5%±10.5%. The achieved accuracy enables patient-specific hemodynamic simulations and may also be of immediate diagnostic interest. Marco Bögel, Sonja Gehrisch, Thomas Redel, Christopher Rohkohl, Annette Birkhold, Philip Hoelter, Arnd Dörfler, Andreas K. Maier, Markus Kowarschik |
ICIP | 8 |
| 2018 | Hyper-Hue and EMAP on Hyperspectral Images for Supervised Layer Decomposition of Old Master DrawingsabstractOld master drawings were mostly created step by step in several layers using different materials. To art historians and restorers, examination of these layers brings various insights into the artistic work process and helps to answer questions about the object, its attribution and its authenticity. However, these layers typically overlap and are oftentimes difficult to differentiate with the unaided eye. For example, a common layer combination is red chalk under ink. In this work, we propose an image processing pipeline that operates on hyperspectral images to separate such layers. In particular, we propose to use two descriptors in hyperspectral historical document analysis, namely hyper-hue and extended multi-attribute profile (EMAP). We show that hyperspectral images enable better layer separation than RGB images, and that spectral focus stacking is an important preprocessing step towards that goal. Our comparative results with other features underline the efficacy of the three proposed improvements. AmirAbbas Davari, Nikolaos Sakaltras, Armin Häberle, Sulaiman Vesal, Vincent Christlein, Andreas K. Maier, Christian Riess |
ICIP | 6 |
| 2018 | Learning from a Handful Volumes: MRI Resolution Enhancement with Volumetric Super-Resolution ForestsabstractMagnetic resonance imaging (MRI) enables 3-D imaging of anatomical structures. However, the acquisition of MR volumes with high spatial resolution leads to long scan times. To this end, we propose volumetric super-resolution forests (VSRF) to enhance MRI resolution retrospectively. Our method learns a locally linear mapping between low-resolution and high-resolution volumetric image patches by employing random forest regression. We customize features suitable for volumetric MRI to train the random forest and propose a median tree ensemble for robust regression. VSRF out-performs state-of-the-art example-based super-resolution in terms of image quality and efficiency for model training and inference on different MRI datasets. It is also superior to unsupervised methods with just a handful or even a single volume to assemble training data. Aline Sindel, Katharina Breininger, Johannes Käßer, Andreas Hess 0001, Andreas K. Maier, Thomas Köhler 0004 |
ICIP | 5 |
| 2018 | Myocardial Scar Segmentation in LGE-MRI using Fractal Analysis and Random Forest ClassificationabstractLate-gadolinium enhanced magnetic resonance imaging (LGE-MRI) is the clinical gold standard to visualize myocardial scarring. The gadolinium based contrast agent accumulates in the damaged cells and leads to various enhancements in the LGE-MRI scan. The quantification of the scar tissue is very important for diagnosis, treatment planning, and guidance during the procedure. In clinical routine, the scar is often segmented manually. However, manual segmentation is prone to inter- and intra-observer variability and very time consuming. In this work a new texture based scar quantification is proposed. For texture characterization, segmentation based fractal analysis is used. First, the image is decomposed into a set of binary images by applying a two-threshold binary decomposition. Second, a set of features are extracted for each of the binary images, namely the fractal dimension, the mean gray value, and the size of the binary object. In addition, the local and global intensity of each patch is added to the feature vector. In the next step, the features are classified using a random forest classifier. The scar quantification is evaluated on 30 clinical LGE-MRI data sets. In addition, the results are compared to the x-fold standard deviation approach and the full-width-at-half-max method, which are implemented in a fully automatic manner. The proposed scar quantification achieved a mean Dice coefficient of 0.64±0.17 and outperforms the x-fold standard deviation approach. Tanja Kurzendorfer, Katharina Breininger, Stefan Steidl, Alexander Brost, Christoph Forman, Andreas K. Maier |
ICPR | 6 |
| 2018 | Precision Learning: Towards Use of Known Operators in Neural NetworksabstractIn this paper, we consider the use of prior knowledge within neural networks. In particular, we investigate the effect of a known transform within the mapping from input data space to the output domain. We demonstrate that use of known transforms is able to change maximal error bounds and that these are additive for the entire sequence of transforms. In order to explore the effect further, we consider the problem of X-ray material decomposition as an example to incorporate additional prior knowledge. We demonstrate that inclusion of a non-linear function known from the physical properties of the system is able to reduce prediction errors therewith improving prediction quality from SSIM values of 0.54 to 0.88. This approach is applicable to a wide set of applications in physics and signal processing that provide prior knowledge on such transforms. Also maximal error estimation and network understanding could be facilitated using this novel concept of precision learning. Andreas K. Maier, Frank Schebesch, Christopher Syben, Tobias Würfl, Stefan Steidl, Jang Hwan Choi 0001, Rebecca Fahrig |
ICPR | 1 |
| 2018 | Fast Sample Generation with Variational Bayesian for Limited Data Hyperspectral Image ClassificationabstractLabeling data for hyperspectral remote sensing image classification is a tedious and cost-intensive task. As a consequence, it is oftentimes necessary to perform classification when only very limited number of labeled training data is available. Several approaches have been proposed to address this problem. A recent proposal is to generate additional synthetic samples from a Gaussian Mixture Model for each class. One challenge with this approach lies in determining the number of components in the GMM. In this paper, we propose an approximation algorithm to select the number of components, namely Variational Bayesian (VB). The main advantage of VB is that it does not require multiple clustering computations in advance. Variational Bayesian not only greatly decreases the computational cost, but also generates comparable or better results in comparison to other methods. AmirAbbas Davari, Hasan Can Ozkan, Andreas K. Maier, Christian Riess |
IGARSS | 3 |
| 2018 | Intraoperative Brain Shift Compensation Using a Hybrid Mixture Model
Siming Bayer, Nishant Ravikumar, Maddalena Strumia, Xiaoguang Tong, Martin Ostermeier, Rebecca Fahrig, Andreas K. Maier |
MICCAI (4) | 8 |
| 2018 | X-ray-transform Invariant Anatomical Landmark Detection for Pelvic Trauma Surgery
Bastian Bier, Mathias Unberath, Jan-Nico Zaech, Javad Fotouhi, Mehran Armand, Greg Osgood, Nassir Navab, Andreas K. Maier |
MICCAI (4) | 8 |
| 2018 | Phase-Sensitive Region-of-Interest Computed Tomography
Lina Felsner, Martin Berger 0002, Sebastian Kaeppler, Johannes Bopp, Veronika Ludwig, Thomas Weber 0001, Georg Pelzer, Thilo Michel, Andreas K. Maier, Gisela Anton, Christian Riess |
MICCAI (1) | 9 |
| 2018 | Closing the Calibration Loop: An Inside-Out-Tracking Paradigm for Augmented Reality in Orthopedic Surgery
Jonas Hajek, Mathias Unberath, Javad Fotouhi, Bastian Bier, Sing Chun Lee, Greg Osgood, Andreas K. Maier, Mehran Armand, Nassir Navab |
MICCAI (4) | 7 |
| 2018 | Some Investigations on Robustness of Deep Learning in Limited Angle Tomography
Yixing Huang, Tobias Würfl, Katharina Breininger, Günter Lauritsch, Andreas K. Maier |
MICCAI (1) | 6 |
| 2018 | Double Your Views - Exploiting Symmetry in Transmission Imaging
Alexander Preuhs, Andreas K. Maier, Michael Manhart 0001, Javad Fotouhi, Nassir Navab, Mathias Unberath |
MICCAI (1) | 2 |
| 2018 | Adversarial and Perceptual Refinement for Compressed Sensing MRI Reconstruction
Maximilian Seitzer, Guang Yang 0006, Jo Schlemper, Ozan Oktay, Tobias Würfl, Vincent Christlein, Tom Wong, Raad Mohiaddin, David N. Firmin, Jennifer Keegan, Daniel Rueckert, Andreas K. Maier |
MICCAI (1) | 12 |
| 2018 | GMM-Based Synthetic Samples for Classification of Hyperspectral Images With Limited Training DataabstractThe amount of training data that is required to train a classifier scales with the dimensionality of the feature data. In hyperspectral remote sensing (HSRS), feature data can potentially become very high dimensional. However, the amount of training data is oftentimes limited. Thus, one of the core challenges in HSRS is how to perform multiclass classification using only relatively few training data points. In this letter, we address this issue by enriching the feature matrix with synthetically generated sample points. These synthetic data are sampled from a Gaussian mixture model (GMM) fitted to each class of the limited training data. Although the true distribution of features may not be perfectly modeled by the fitted GMM, we demonstrate that a moderate augmentation by these synthetic samples can effectively replace a part of the missing training samples. Doing so, the median gain in classification performance is 5% on two datasets. This performance gain is stable for variations in the number of added samples, which makes it easy to apply this method to real-world applications. AmirAbbas Davari, Erchan Aptoula, Berrin A. Yanikoglu, Andreas K. Maier, Christian Riess |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2018 | Towards intelligent robust detection of anatomical structures in incomplete volumetric data
Florin C. Ghesu, Bogdan Georgescu, Sasa Grbic, Andreas K. Maier, Joachim Hornegger, Dorin Comaniciu |
Medical Image Anal. | 4 |
| 2018 | Temporal and volumetric denoising via quantile sparse image prior
Franziska Schirrmacher, Thomas Köhler 0004, Jürgen Endres, Tobias Lindenberger, Lennart Husvogt, James G. Fujimoto, Joachim Hornegger, Arnd Dörfler, Philip Hoelter, Andreas K. Maier |
Medical Image Anal. | 10 |
| 2018 | Geometric primitive refinement for structured light cameras
Peter Fürsattel, Simon Placht, Andreas K. Maier, Christian Riess |
Mach. Vis. Appl. | 3 |
| 2018 | An MR-Based Model for Cardio-Respiratory Motion Compensation of Overlays in X-Ray FluoroscopyabstractIn X-ray fluoroscopy, static overlays are used to visualize soft tissue. We propose a system for cardiac and respiratory motion compensation of these overlays. It consists of a 3-D motion model created from real-time magnetic resonance (MR) imaging. Multiple sagittal slices are acquired and retrospectively stacked to consistent 3-D volumes. Slice stacking considers cardiac information derived from the ECG and respiratory information extracted from the images. Additionally, temporal smoothness of the stacking is enhanced. Motion is estimated from the MR volumes using deformable 3-D/3-D registration. The motion model itself is a linear direct correspondence model using the same surrogate signals as slice stacking. In X-ray fluoroscopy, only the surrogate signals need to be extracted to apply the motion model and animate the overlay in real time. For evaluation, points are manually annotated in oblique MR slices and in contrast-enhanced X-ray images. The 2-D Euclidean distance of these points is reduced from 3.85 to 2.75 mm in MR and from 3.0 to 1.8 mm in X-ray compared with the static baseline. Furthermore, the motion-compensated overlays are shown qualitatively as images and videos. Peter Fischer 0001, Anthony Faranesh, Thomas Pohl, Andreas K. Maier, Toby Rogers, Kanishka Ratnayaka, Robert Lederman, Joachim Hornegger |
IEEE Trans. Medical Imaging | 4 |
| 2018 | Prior-Free Respiratory Motion Estimation in Rotational AngiographyabstractRotational coronary angiography using C-arm angiography systems enables intra-procedural 3-D imaging that is considered beneficial for diagnostic assessment and interventional guidance. Despite previous efforts, rotational angiography was not yet successfully established in clinical practice for coronary artery procedures due to challenges associated with substantial intra-scan respiratory and cardiac motion. While gating handles cardiac motion during reconstruction, respiratory motion requires compensation. State-of-the-art algorithms rely on 3-D / 2-D registration that requires an uncompensated reconstruction of sufficient quality. To overcome this limitation, we investigate two prior-free respiratory motion estimation methods based on the optimization of: 1) epipolar consistency conditions (ECCs) and 2) a task-based auto-focus measure (AFM). The methods assess redundancies in projection images or impose favorable properties of 3-D space, respectively, and are used to estimate the respiratory motion of the coronary arteries within rotational angiograms. We evaluate our algorithms on the publicly available CAVAREV benchmark and on clinical data. We quantify reductions in error due to respiratory motion compensation using a dedicated reconstruction domain metric. Moreover, we study the improvements in image quality when using an analytic and a novel temporal total variation regularized algebraic reconstruction algorithm. We observed substantial improvement in all figures of merit compared with the uncompensated case. Improvements in image quality presented as a reduction of double edges, blurring, and noise. Benefits of the proposed corrections were notable even in cases suffering little corruption from respiratory motion, translating to an improvement in the vessel sharpness of (6.08 ± 4.46)% and (14.7 ± 8.80)% when the ECC-based and the AFM-based compensation were applied. On the CAVAREV data, our motion compensation approach exhibits an improvement of (27.6 ± 7.5)% and (97.0 ± 17.7)% when the ECC and AFM were used, respectively. At the time of writing, our method based on AFM is leading the CAVAREV scoreboard. Both motion estimation strategies are purely image-based and accurately estimate the displacements of the coronary arteries due to respiration. While current evidence suggests the superior performance of AFM, future work will further investigate the use of ECC in the context of angiography as they solely rely on geometric calibration and projection-domain images. Mathias Unberath, Oliver Taubmann, André Aichert, Stephan Achenbach, Andreas K. Maier |
IEEE Trans. Medical Imaging | 5 |
| 2018 | Deep Learning Computed Tomography: Learning Projection-Domain Weights From Image Domain in Limited Angle ProblemsabstractIn this paper, we present a new deep learning framework for 3-D tomographic reconstruction. To this end, we map filtered back-projection-type algorithms to neural networks. However, the back-projection cannot be implemented as a fully connected layer due to its memory requirements. To overcome this problem, we propose a new type of cone-beam back-projection layer, efficiently calculating the forward pass. We derive this layer's backward pass as a projection operation. Unlike most deep learning approaches for reconstruction, our new layer permits joint optimization of correction steps in volume and projection domain. Evaluation is performed numerically on a public data set in a limited angle setting showing a consistent improvement over analytical algorithms while keeping the same computational test-time complexity by design. In the region of interest, the peak signal-to-noise ratio has increased by 23%. In addition, we show that the learned algorithm can be interpreted using known concepts from cone beam reconstruction: the network is able to automatically learn strategies such as compensation weights and apodization windows. Tobias Würfl, Mathis Hoffmann, Vincent Christlein, Katharina Breininger, Yixing Huang, Mathias Unberath, Andreas K. Maier |
IEEE Trans. Medical Imaging | 7 |
| 2018 | Classification With Truncated $\ell _{1}$ Distance KernelabstractThis brief proposes a truncated distance (TL1) kernel, which results in a classifier that is nonlinear in the global region but is linear in each subregion. With this kernel, the subregion structure can be trained using all the training data and local linear classifiers can be established simultaneously. The TL1 kernel has good adaptiveness to nonlinearity and is suitable for problems which require different nonlinearities in different areas. Though the TL1 kernel is not positive semidefinite, some classical kernel learning methods are still applicable which means that the TL1 kernel can be directly used in standard toolboxes by replacing the kernel evaluation. In numerical experiments, the TL1 kernel with a pregiven parameter achieves similar or better performance than the radial basis function kernel with the parameter tuned by cross validation, implying the TL1 kernel a promising nonlinear kernel for classification tasks. Xiaolin Huang, Johan A. K. Suykens, Shuning Wang, Joachim Hornegger, Andreas K. Maier |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | GMM Supervectors for Limited Training Data in Hyperspectral Remote Sensing Image Classification
AmirAbbas Davari, Vincent Christlein, Sulaiman Vesal, Andreas K. Maier, Christian Riess |
CAIP (2) | 4 |
| 2017 | Unsupervised Feature Learning for Writer Identification and Writer RetrievalabstractDeep Convolutional Neural Networks (CNN) have shown great success in supervised classification tasks such as character classification or dating. Deep learning methods typically need a lot of annotated training data, which is not available in many scenarios. In these cases, traditional methods are often better than or equivalent to deep learning methods. In this paper, we propose a simple, yet effective, way to learn CNN activation features in an unsupervised manner. Therefore, we train a deep residual network using surrogate classes. The surrogate classes are created by clustering the training dataset, where each cluster index represents one surrogate class. The activations from the penultimate CNN layer serve as features for subsequent classification tasks. We evaluate the feature representations on two publicly available datasets. The focus lies on the ICDAR17 competition dataset on historical document writer identification (Historical-WI). We show that the activation features trained without supervision are superior to descriptors of state-of-the-art writer identification methods. Additionally, we achieve comparable results in the case of handwriting classification using the ICFHR16 competition dataset on historical Latin script types (CLaMM16). Vincent Christlein, Martin Gropp, Stefan Fiel, Andreas K. Maier |
ICDAR | 4 |
| 2017 | Browsing through Closed Books: Fully Automatic Book Page Extraction from a 3-D X-Ray CT VolumeabstractWhen digitizing or investigating historical documents, it is often the case that a document can not be opened, page-turned or touched anymore. Damages such as moisture or fire and aging processes disallow browsing through a book. To address these particular cases, our earlier work showed that Micro-CT X-ray scanners are able to image documents written with iron gall ink. A self-made book consisting of ten hand written pages was scanned and investigated without opening or page-turning. However, when analyzing the reconstruction results, we faced the problem of a proper automatic page segmentation and 2-D mapping within the volume in an acceptable time without losing information of the writings. The main problem is that the pages can be arbitrary deformed or squeezed together. In this paper, we present a fully automatic algorithm for the segmentation and extraction of book pages from the original 3-D volume. Our method delivers high quality results for our book model and can be easily adapted to other imaging modalities. We show that it performs well even for an extreme case with low resolution input data and wavy pages. To keep it simple for users, our algorithm works without any need of prior information or user interactions. Daniel Stromer, Vincent Christlein, Tobias Schön, Wolfgang Holub, Andreas K. Maier |
ICDAR | 5 |
| 2017 | Robust Multi-scale Anatomical Landmark Detection in Incomplete 3D-CT Data
Florin C. Ghesu, Bogdan Georgescu, Sasa Grbic, Andreas K. Maier, Joachim Hornegger, Dorin Comaniciu |
MICCAI (1) | 4 |
| 2017 | Robust Non-rigid Registration Through Agent-Based Action Learning
Julian Krebs, Tommaso Mansi, Hervé Delingette, Florin C. Ghesu, Shun Miao, Andreas K. Maier, Nicholas Ayache, Rui Liao, Ali Kamen |
MICCAI (1) | 7 |
| 2017 | QuaSI: Quantile Sparse Image Prior for Spatio-Temporal Denoising of Retinal OCT Data
Franziska Schirrmacher, Thomas Köhler 0004, Lennart Husvogt, James G. Fujimoto, Joachim Hornegger, Andreas K. Maier |
MICCAI (2) | 6 |
| 2017 | The 19th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2016)
Sébastien Ourselin, Mert R. Sabuncu, William M. Wells III, Leo Joskowicz, Gozde Unal, Andreas K. Maier |
Medical Image Anal. | 6 |
| 2017 | Writer Identification Using GMM Supervectors and Exemplar-SVMs
Vincent Christlein, David Bernecker, Florian Hönig, Andreas K. Maier, Elli Angelopoulou |
Pattern Recognit. | 4 |
| 2017 | Unsupervised Learning for Robust Respiratory Signal Estimation From X-Ray FluoroscopyabstractRespiratory signals are required for image gating and motion compensation in minimally invasive interventions. In X-ray fluoroscopy, extraction of a respiratory signal can be challenging due to characteristics of interventional imaging, in particular injection of contrast agent and automatic exposure control. We present a novel method for respiratory signal extraction based on dimensionality reduction that can tolerate these events. Images are divided into patches of multiple sizes. Low-dimensional embeddings are generated for each patch using illumination-invariant kernel PCA. Patches with respiratory information are selected automatically by agglomerative clustering. The signals from this respiratory cluster are combined robustly to a single respiratory signal. In the experiments, we evaluate our method on a variety of scenarios. If the diaphragm is visible, we track its superior-inferior motion as ground truth. Our method has a correlation coefficient of more than 91% with the ground truth irrespective of whether or not contrast agent injection or automatic exposure control occur. Additionally, we show that very similar signals are estimated from biplane sequences and from sequences without visible diaphragm. Since all these cases are handled automatically, the method is robust enough to be considered for use in a clinical setting. Peter Fischer 0001, Thomas Pohl, Anthony Faranesh, Andreas K. Maier, Joachim Hornegger |
IEEE Trans. Medical Imaging | 4 |
| 2017 | Dynamic 2-D/3-D Rigid Registration Framework Using Point-To-Plane Correspondence ModelabstractIn image-guided interventional procedures, live 2-D X-ray images can be augmented with preoperative 3-D computed tomography or MRI images to provide planning landmarks and enhanced spatial perception. An accurate alignment between the 3-D and 2-D images is a prerequisite for fusion applications. This paper presents a dynamic rigid 2-D/3-D registration framework, which measures the local 3-D-to-2-D misalignment and efficiently constrains the update of both planar and non-planar 3-D rigid transformations using a novel point-to-plane correspondence model. In the simulation evaluation, the proposed method achieved a mean 3-D accuracy of 0.07 mm for the head phantom and 0.05 mm for the thorax phantom using single-view X-ray images. In the evaluation on dynamic motion compensation, our method significantly increases the accuracy comparing with the baseline method. The proposed method is also evaluated on a publicly-available clinical angiogram data set with "gold-standard" registrations. The proposed method achieved a mean 3-D accuracy below 0.8 mm and a mean 2-D accuracy below 0.3 mm using single-view X-ray images. It outperformed the state-of-the-art methods in both accuracy and robustness in single-view registration. The proposed method is intuitive, generic, and suitable for both initial and dynamic registration scenarios. Jian Wang 0009, Roman Schaffert, Anja Borsdorf, Benno Heigl, Xiaolin Huang, Joachim Hornegger, Andreas K. Maier |
IEEE Trans. Medical Imaging | 7 |
| 2016 | Confidence-aware Levenberg-Marquardt optimization for joint motion estimation and super-resolutionabstractMotion estimation across low-resolution frames and the reconstruction of high-resolution images are two coupled subproblems of multi-frame super-resolution. This paper introduces a new joint optimization approach for motion estimation and image reconstruction to address this interdependence. Our method is formulated via non-linear least squares optimization and combines two principles of robust super-resolution. First, to enhance the robustness of the joint estimation, we propose a confidence-aware energy minimization framework augmented with sparse regularization. Second, we develop a tailor-made Levenberg-Marquardt iteration scheme to jointly estimate motion parameters and the high-resolution image along with the corresponding model confidence parameters. Our experiments on simulated and real images confirm that the proposed approach outperforms decoupled motion estimation and image reconstruction as well as related state-of-the-art joint estimation algorithms. Cosmin Bercea, Andreas K. Maier, Thomas Köhler 0004 |
ICIP | 2 |
| 2016 | Vesselness for text detection in historical document imagesabstractText detection is typically the first step for any text processing such as hand-written text recognition, layout analysis, line detection, or writer identification. This paper describes a new method to detect text in images, particularly in historical document images. For a robust detection, we propose the use of the vesselness filter as a new preprocessing step for text detection. We show, that this step improves the detection rate significantly. At the locations segmented by this filter, SIFT keypoints are detected which are spatially clustered. Overlapping windows from these clusters are subsequently VLAD encoded and classified in text and non-text. We evaluate this approach on a newly created database, where we achieve an F1-score of 92%. Additionally, we demonstrate the effectiveness of this method for line segmentation. Simon Hofmann, Martin Gropp, David Bernecker, Christopher Pollin, Andreas K. Maier, Vincent Christlein |
ICIP | 5 |
| 2016 | Joint Estimation of Cardiac Motion and T_1^* Maps for Magnetic Resonance Late Gadolinium Enhancement ImagingabstractIn the diagnosis of myocardial infarction, magnetic resonance imaging can provide information about myocardial contractility and tissue characterization, including viability. In current clinical practice, separate scans are required for each aspect. A recently proposed method showed how the same information can be extracted from a single, short scan of \(4\,\text {s}\) , but made strong assumptions about the underlying cardiac motion. We propose a fixed-point iteration scheme that retains the benefits of their approach while lifting its limitations, making it robust to cardiac arrhythmia. We compare our method to the state of the art using phantom data as well as data from 11 patients and show a consistent improvement of all evaluation criteria, e. g. the end-diastolic Dice coefficient of an arrythmic case improves from \(86\,\%\) (state-of-the-art method) to \(94\,\%\) (proposed method). These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Jens Wetzl, Aurélien F. Stalder, Michaela Schmidt, Yigit H. Akgök, Christoph Tillmanns, Felix Lugauer, Christoph Forman, Joachim Hornegger, Andreas K. Maier |
MICCAI (3) | 9 |
| 2016 | Deep Learning Computed Tomography
Tobias Würfl, Florin C. Ghesu, Vincent Christlein, Andreas K. Maier |
MICCAI (3) | 4 |
| 2016 | Automatic detection and analysis of photovoltaic modules in aerial infrared imageryabstractDrone-based aerial thermography has become a convenient quality assessment tool for the precise localization of defective modules and cells in large photovoltaic-power plants. However, manual evaluation of aerial infrared recordings can be extremely time-consuming. Therefore, we propose an approach for automatic detection and analysis of photovoltaic modules in aerial infrared images. Significant temperature abnormalities such as hot spots and hot areas can be identified using our processing pipeline. To identify such defects, we first detect the individual modules in infrared images, and then use statistical tests to detect the defective modules. A quantitative evaluation of the detection and analysis pipeline on real-world, infrared recordings shows the applicability of our approach. Sergiu Deitsch, Manuel Dalsass, Ludwig Winkler, Tobias Wurzner, Christoph Brabec, Andreas K. Maier, Florian Gallwitz |
WACV | 6 |
| 2016 | OCPAD - Occluded checkerboard pattern detectorabstractMany camera calibration techniques require the detection of a pattern with known geometry, e.g., a checkerboard. Typically, the pattern must be fully contained in the field of view. This brings several limitations, one of which is that lens distortion can not reliably be estimated in outer image regions. This paper presents the occluded checkerboard pattern detector (OCPAD) to find checkerboards, even in a) low-resolution images, b) images with high lens distortion and if c) the pattern is partly occluded or not completely within the field of view. We exploit that checkerboards can easily be represented by a graph. We use graph matching to find the largest partial checkerboard in the image. Our detector complements a state-of-the-art calibration algorithm. Quantitatively, detection rates are considerably improved over the state-of-the-art. Additionally, estimation of lens distortion is greatly improved at outer image regions. Here, the reprojection error is improved by up to 50%. Peter Fürsattel, Sergiu Deitsch, Simon Placht, Michael Balda, Andreas K. Maier, Christian Riess |
WACV | 5 |
| 2016 | Editorial for the Special Issue on MICCAI 2015
Nassir Navab, Alejandro F. Frangi, William M. Wells III, Andreas K. Maier |
Medical Image Anal. | 4 |
| 2016 | Electrophysiology Catheter Detection and Reconstruction From Two Views in Fluoroscopic ImagesabstractElectrophysiology (EP) studies and catheter ablation have become important treatment options for several types of cardiac arrhythmias. We present a novel image-based approach for automatic detection and 3-D reconstruction of EP catheters where the physician marks the catheter to be reconstructed by a single click in each image. The result can be used to provide 3-D information for enhanced navigation throughout EP procedures. Our approach involves two X-ray projections acquired from different angles, and it is based on two steps: First, we detect the catheter in each view after manual initialization using a graph-search method. Then, the detection results are used to reconstruct a full 3-D model of the catheter based on automatically determined point pairs for triangulation. An evaluation on 176 different clinical fluoroscopic images yielded a detection rate of 83.4%. For measuring the error, we used the coupling distance which is a more accurate quality measure than the average point-wise distance to a reference. For successful outcomes, the 2-D detection error was 1.7 mm ±1.2 mm. Using successfully detected catheters for reconstruction, we obtained a reconstruction error of 1.8 mm ±1.1 mm on phantom data. On clinical data, our method yielded a reconstruction error of 2.2 mm ±2.2 mm. Matthias Hoffmann, Alexander Brost, Martin Koch 0002, Felix Bourier, Andreas K. Maier, Klaus Kurzidim, Norbert Strobel, Joachim Hornegger |
IEEE Trans. Medical Imaging | 5 |
| 2016 | Cryo-Balloon Catheter Localization Based on a Support-Vector-Machine ApproachabstractCryo-balloon catheters have attracted an increasing amount of interest in the medical community as they can reduce patient risk during left atrial pulmonary vein ablation procedures. As cryo-balloon catheters are not equipped with electrodes, they cannot be localized automatically by electro-anatomical mapping systems. As a consequence, X-ray fluoroscopy has remained an important means for guidance during the procedure. Most recently, image guidance methods for fluoroscopy-based procedures have been proposed, but they provide only limited support for cryo-balloon catheters and require significant user interaction. To improve this situation, we propose a novel method for automatic cryo-balloon catheter detection in fluoroscopic images by detecting the cryo-balloon catheter's built-in X-ray marker. Our approach is based on a blob detection algorithm to find possible X-ray marker candidates. Several of these candidates are then excluded using prior knowledge. For the remaining candidates, several catheter specific features are introduced. They are processed using a machine learning approach to arrive at the final X-ray marker position. Our method was evaluated on 75 biplane fluoroscopy images from 40 patients, from two sites, acquired with a biplane angiography system. The method yielded a success rate of 99.0% in plane A and 90.6% in plane B, respectively. The detection achieved an accuracy of 1.00 mm±0.82 mm in plane A and 1.13 mm±0.24 mm in plane B. The localization in 3-D was associated with an average error of 0.36 mm±0.86 mm. Tanja Kurzendorfer, Philip Walter Mewes, Andreas K. Maier, Norbert Strobel, Alexander Brost |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Fully Automated Data-Driven Respiratory Signal Extraction From SPECT Images Using Laplacian EigenmapsabstractWe propose a data-driven method for extracting a respiratory surrogate signal from SPECT list-mode data. The approach is based on dimensionality reduction with Laplacian Eigenmaps. By setting a scale parameter adaptively and adding a series of post-processing steps to correct polarity and normalization between projections, we enable fully-automatic operation and deliver a respiratory surrogate signal for the entire SPECT acquisition. We validated the method using 67 patient scans from three acquisition types (myocardial perfusion, liver shunt diagnostic, lung inhalation/perfusion) and an Anzai pressure belt as a gold standard. The proposed method achieved a mean correlation against the Anzai of 0.81 ± 0.17 (median 0.89). In a subsequent analysis, we characterize the performance of the method with respect to count rates and describe a predictor for identifying scans with insufficient statistics. To the best of our knowledge, this is the first large validation of a data-driven respiratory signal extraction method published thus far for SPECT, and our results compare well with those reported in the literature for such techniques applied to other modalities such as MR and PET. James C. Sanders, Philipp Ritt, Torsten Kuwert, Alexander Hans Vija, Andreas K. Maier |
IEEE Trans. Medical Imaging | 5 |
| 2015 | Epipolar Consistency in Fluoroscopy for Image-Based TrackingabstractGeometry and physics of absorption imaging impose certain constraints on X-ray projections. Recently, the Epipolar Consistency Conditions (ECC) have been introduced and applied to motion correction in flat-detector computed tomography (CT). They are based on redundant information in transmission images along epipolar lines. Unlike other consistency conditions for CT scans, they act directly on an arbitrary pair of X-ray images. This paper proposes an application of ECC to 3D patient tracking in interventional radiology. We evaluate the proposed method against 2D-3D registration with a previously acquired CT. Our experiments on synthetic data based on a patient CT and phantom data from an interventional C-arm demonstrate that our method is able to compensate online for rotations of up to ±10° and translations of ±25 mm between consecutive frames. We successfully track rotations of as much as 45° over 45 images. The outstanding property of the approach is that no 3D scan is required for tracking a 3D object in space. We show, that small rotations of about 3° in space and translations of about 50 mm can be tracked based on just two reference X-ray images. Since the proposed approach works directly on X-ray images, it exceeds regular 2D-3D registration with a CT in an order of magnitude in computational speed. We conclude that ECC are a simple and effective new tool for pre-aligment and online patient tracking for fluoroscopic sequences. André Aichert, Jian Wang 0009, Roman Schaffert, Arnd Dörfler, Joachim Hornegger, Andreas K. Maier |
BMVC | 6 |
| 2015 | A Unified Bayesian Approach to Multi-Frame Super-Resolution and Single-Image Upsampling in Multi-Sensor Imaging
Thomas Köhler 0004, Johannes Jordan, Andreas K. Maier, Joachim Hornegger |
BMVC | 3 |
| 2015 | Surrogate-Driven Estimation of Respiratory Motion and Layers in X-Ray Fluoroscopy
Peter Fischer 0001, Thomas Pohl, Andreas K. Maier, Joachim Hornegger |
MICCAI (1) | 3 |
| 2015 | Robust Spectral Denoising for Water-Fat Separation in Magnetic Resonance Imaging
Felix Lugauer, Marcel Dominik Nickel, Jens Wetzl, Stephan A. R. Kannengiesser, Andreas K. Maier, Joachim Hornegger |
MICCAI (2) | 5 |
| 2015 | Estimate, Compensate, Iterate: Joint Motion Estimation and Compensation in 4-D Cardiac C-arm Computed Tomography
Oliver Taubmann, Günter Lauritsch, Andreas K. Maier, Rebecca Fahrig, Joachim Hornegger |
MICCAI (2) | 3 |
| 2015 | Adaption of 3D Models to 2D X-Ray Images during Endovascular Abdominal Aneurysm Repair
Daniel Toth 0001, Marcus Pfister, Andreas K. Maier, Markus Kowarschik, Joachim Hornegger |
MICCAI (1) | 3 |
| 2015 | Epipolar Consistency in Transmission ImagingabstractThis paper presents the derivation of the Epipolar Consistency Conditions (ECC) between two X-ray images from the Beer-Lambert law of X-ray attenuation and the Epipolar Geometry of two pinhole cameras, using Grangeat's theorem. We motivate the use of Oriented Projective Geometry to express redundant line integrals in projection images and define a consistency metric, which can be used, for instance, to estimate patient motion directly from a set of X-ray images. We describe in detail the mathematical tools to implement an algorithm to compute the Epipolar Consistency Metric and investigate its properties with detailed random studies on both artificial and real FD-CT data. A set of six reference projections of the CT scan of a fish were used to evaluate accuracy and precision of compensating for random disturbances of the ground truth projection matrix using an optimization of the consistency metric. In addition, we use three X-ray images of a pumpkin to prove applicability to real data. We conclude, that the metric might have potential in applications related to the estimation of projection geometry. By expression of redundancy between two arbitrary projection views, we in fact support any device or acquisition trajectory which uses a cone-beam geometry. We discuss certain geometric situations, where the ECC provide the ability to correct 3D motion, without the need for 3D reconstruction. André Aichert, Martin Berger 0002, Jian Wang 0009, Nicole Maass, Arnd Dörfler, Joachim Hornegger, Andreas K. Maier |
IEEE Trans. Medical Imaging | 7 |
| 2015 | A Gauss-Seidel Iteration Scheme for Reference-Free 3-D Histological Image ReconstructionabstractThree-dimensional (3-D) reconstruction of histological slice sequences offers great benefits in the investigation of different morphologies. It features very high-resolution which is still unmatched by in vivo 3-D imaging modalities, and tissue staining further enhances visibility and contrast. One important step during reconstruction is the reversal of slice deformations introduced during histological slice preparation, a process also called image unwarping. Most methods use an external reference, or rely on conservative stopping criteria during the unwarping optimization to prevent straightening of naturally curved morphology. Our approach shows that the problem of unwarping is based on the superposition of low-frequency anatomy and high-frequency errors. We present an iterative scheme that transfers the ideas of the Gauss-Seidel method to image stacks to separate the anatomy from the deformation. In particular, the scheme is universally applicable without restriction to a specific unwarping method, and uses no external reference. The deformation artifacts are effectively reduced in the resulting histology volumes, while the natural curvature of the anatomy is preserved. The validity of our method is shown on synthetic data, simulated histology data using a CT data set and real histology data. In the case of the simulated histology where the ground truth was known, the mean Target Registration Error (TRE) between the unwarped and original volume could be reduced to less than 1 pixel on average after six iterations of our proposed method. Simone Gaffling, Volker Daum, Stefan Steidl, Andreas K. Maier, Harald Köstler, Joachim Hornegger |
IEEE Trans. Medical Imaging | 4 |
| 2015 | Multi-Dimensional Flow-Preserving Compressed Sensing (MuFloCoS) for Time-Resolved Velocity-Encoded Phase Contrast MRIabstract4-D time-resolved velocity-encoded phase-contrast MRI (4-D PCI) is a fully non-invasive technique to assess hemodynamics in vivo with a broad range of potential applications in multiple cardiovascular diseases. It is capable of providing quantitative flow values and anatomical information simultaneously. The long acquisition time, however, still inhibits its wider clinical use. Acceleration is achieved at present using parallel MRI (pMRI) techniques which can lead to substantial loss of image quality for higher acceleration factors. Both the high-dimensionality and the significant degree of spatio-temporal correlation in 4-D PCI render it ideally suited for recently proposed compressed sensing (CS) techniques. We propose the Multi-Dimensional Flow-preserving Compressed Sensing (MuFloCoS) method to exploit these properties. A multi-dimensional iterative reconstruction is combined with an interleaved sampling pattern (I-VT), an adaptive masked and weighted temporal regularization (TMW) and fully automatically obtained vessel-masks. The performance of the novel method was analyzed concerning image quality, feasibility of acceleration factors up to 15, quantitative flow values and diagnostic accuracy in phantom experiments and an in vivo carotid study with 18 volunteers. Comparison with iterative state-of-the-art methods revealed significant improvements using the new method, the temporal normalized root mean square error of the peak velocity was reduced by 45.32% for the novel MuFloCoS method with acceleration factor 9. The method was furthermore applied to two patient cases with diagnosed high-grade stenosis of the ICA, which confirmed the performance of MuFloCoS to produce valuable results in the presence of pathological findings in 56 s instead of over 8 min (full sampling). Jana Hutter, Peter Schmitt, Marc Saake, Axel Stubinger, Robert Grimm 0002, Christoph Forman, Andreas Greiser, Joachim Hornegger, Andreas K. Maier |
IEEE Trans. Medical Imaging | 9 |
| 2015 | Axially Extended-Volume C-Arm CT Using a Reverse Helical Trajectory in the Interventional RoomabstractC-arm computed tomography (CT) is an innovative technique that enables a C-arm system to generate 3-D images from a set of 2-D X-ray projections. This technique can reduce treatment-related complications and may improve interventional efficacy and safety. However, state-of-the-art C-arm systems rely on a circular short scan for data acquisition, which limits coverage in the axial direction. This limitation was reported as a problem in hepatic vascular interventions. To solve this problem, as well as to further extend the value of C-arm CT, axially extended-volume C-arm CT is needed. For example, such an extension would enable imaging the full aorta, the peripheral arteries or the spine in the interventional room, which is currently not feasible. In this paper, we demonstrate that performing long object imaging using a reverse helix is feasible in the interventional room. This demonstration involved developing a novel calibration method, assessing geometric repeatability, implementing a reconstruction method that applies to real reverse helical data, and quantitatively evaluating image quality. Our results show that: 1) the reverse helical trajectory can be implemented and reliably repeated on a multiaxis C-arm system; and 2) a long volume can be reconstructed with satisfactory image quality using reverse helical data. Zhicong Yu, Andreas K. Maier, Günter Lauritsch, Florian Vogt, Manfred Schonborn, Christoph Kohler, Joachim Hornegger, Frédéric Noo |
IEEE Trans. Medical Imaging | 2 |
| 2014 | Erlangen-CLP: A Large Annotated Corpus of Speech from Children with Cleft Lip and Palate
Tobias Bocklet, Andreas K. Maier, Korbinian Riedhammer, Ulrich Eysholdt, Elmar Nöth |
LREC | 2 |
| 2014 | Signal Decomposition for X-ray Dark-Field Imaging
Sebastian Kaeppler, Florian Bayer, Thomas Weber 0001, Andreas K. Maier, Gisela Anton, Joachim Hornegger, Matthias W. Beckmann, Peter A. Fasching, Arndt Hartmann, Felix Heindl, Thilo Michel, Gueluemser Oezguel, Georg Pelzer, Claudia Rauh, Jens Rieger, Rüdiger Schulz-Wendtland, Michael Uder, David Wachter, Evelyn Wenkel, Christian Riess |
MICCAI (1) | 4 |
| 2014 | Unsupervised Unstained Cell Detection by SIFT Keypoint Clustering and Self-labeling Algorithm
Firas Mualla, Simon Schöll, Björn Sommerfeldt, Andreas K. Maier, Stefan Steidl, Rainer Buchholz, Joachim Hornegger |
MICCAI (3) | 4 |
| 2014 | Towards Clinical Application of a Laplace Operator-Based Region of Interest Reconstruction Algorithm in C-Arm CTabstractIt is known that a reduction of the field-of-view in 3-D X-ray imaging is proportional to a reduction in radiation dose. The resulting truncation, however, is incompatible with conventional reconstruction algorithms. Recently, a novel method for region of interest reconstruction that uses neither prior knowledge nor extrapolation has been published, named approximated truncation robust algorithm for computed tomography (ATRACT). It is based on a decomposition of the standard ramp filter into a 2-D Laplace filtering and a 2-D Radon-based residual filtering step. In this paper, we present two variants of the original ATRACT. One is based on expressing the residual filter as an efficient 2-D convolution with an analytically derived kernel. The second variant is to apply ATRACT in 1-D to further reduce computational complexity. The proposed algorithms were evaluated by using a reconstruction benchmark, as well as two clinical data sets. The results are encouraging since the proposed algorithms achieve a speed-up factor of up to 245 compared to the 2-D Radon-based ATRACT. Reconstructions of high accuracy are obtained, e.g., even real-data reconstruction in the presence of severe truncation achieve a relative root mean square error of as little as 0.92% with respect to nontruncated data. Yan Xia 0002, Hannes G. Hofmann, Frank Dennerlein, Kerstin Müller 0002, Chris Schwemmer, Sebastian Bauer 0001, Gouthami Chintalapani, Ponraj Chinnadurai, Joachim Hornegger, Andreas K. Maier |
IEEE Trans. Medical Imaging | 10 |
| 2013 | Free-Breathing Whole-Heart Coronary MRA: Motion Compensation Integrated into 3D Cartesian Compressed Sensing Reconstruction
Christoph Forman, Robert Grimm 0002, Jana Hutter, Andreas K. Maier, Joachim Hornegger, Michael O. Zenge |
MICCAI (2) | 4 |
| 2013 | Low-Rank and Sparse Matrix Decomposition for Compressed Sensing Reconstruction of Magnetic Resonance 4D Phase Contrast Blood Flow Imaging (LoSDeCoS 4D-PCI)
Jana Hutter, Peter Schmitt, Gunhild E. Aandal, Andreas Greiser, Christoph Forman, Robert Grimm 0002, Joachim Hornegger, Andreas K. Maier |
MICCAI (1) | 8 |
| 2013 | Dynamic Iterative Reconstruction for Interventional 4-D C-Arm CT Perfusion ImagingabstractTissue perfusion measurement using C-arm angiography systems capable of CT-like imaging (C-arm CT) is a novel technique with potentially high benefit for catheter guided treatment of stroke in the interventional suite. However, perfusion C-arm CT (PCCT) is challenging: the slow C-arm rotation speed only allows measuring samples of contrast time attenuation curves (TACs) every 5-6 s if reconstruction algorithms for static data are used. Furthermore, the peak values of the TACs in brain tissue typically lie in a range of 5-30 HU, thus perfusion imaging is very sensitive to noise. We present a dynamic, iterative reconstruction (DIR) approach to reconstruct TACs described by a weighted sum of basis functions. To reduce noise, a regularization technique based on joint bilateral filtering (JBF) is introduced. We evaluated the algorithm with a digital dynamic brain phantom and with data from six canine stroke models. With our dynamic approach, we achieve an average Pearson correlation (PC) of the PCCT canine blood flow maps to co-registered perfusion CT maps of 0.73. This PC is just as high as the PC achieved in a recent PCCT study, which required repeated injections and acquisitions. Michael Manhart 0001, Markus Kowarschik, Andreas Fieselmann, Yu Deuerling-Zheng, Kevin Royalty, Andreas K. Maier, Joachim Hornegger |
IEEE Trans. Medical Imaging | 6 |
| 2013 | Automatic Cell Detection in Bright-Field Microscope Images Using SIFT, Random Forests, and Hierarchical ClusteringabstractWe present a novel machine learning-based system for unstained cell detection in bright-field microscope images. The system is fully automatic since it requires no manual parameter tuning. It is also highly invariant with respect to illumination conditions and to the size and orientation of cells. Images from two adherent cell lines and one suspension cell line were used in the evaluation for a total number of more than 3500 cells. Besides real images, simulated images were also used in the evaluation. The detection error was between approximately zero and 15.5% which is a significantly superior performance compared to baseline approaches. Firas Mualla, Simon Schöll, Björn Sommerfeldt, Andreas K. Maier, Joachim Hornegger |
IEEE Trans. Medical Imaging | 4 |
| 2010 | Improvement of a speech recognizer for standardized medical assessment of children's speech by integration of prior knowledgeabstractSpeech recognition of children is a more difficult task than speech recognition of adults. This problem is amplified for children with articulation disorders like cleft lip and palate (CLP). In this work we improved our automatic speech recognition system by integrating prior knowledge. Prior knowledge focuses on two different aspects: A test-dependent language modeling and an age-dependent acoustic modeling. These two approaches are merged at the end to different test- and age-dependent recognizers. We evaluated our system on a dataset of 35 children with CLP. Significant improvements could be found on this dataset. With our baseline system we achieved a negative word accuarcy (WA) of -11.0%. By an extended language modeling we achieved 27.5%. The age-dependent recognition system gains a huge improvement and achieves aWA of 42.6%. With the significant improvements in WA it is possible to perform an automatic detection and identification of specific words. Thus, we took the first step towards a speech assessment on word and subword level. Tobias Bocklet, Andreas K. Maier, Ulrich Eysholdt, Elmar Nöth |
SLT | 2 |
| 2009 | Immersive Painting
Stefan Soutschek, Florian Hönig, Andreas K. Maier, Stefan Steidl, Michael Stürmer, Hellmut Erzigkeit, Joachim Hornegger, Johannes Kornhuber |
ArtsIT | 3 |
| 2009 | A language-independent feature set for the automatic evaluation of prosodyabstractIn second language learning, the correct use of prosody plays a vital role.Therefore, an automatic method to evaluate the naturalness of the prosody of a speaker is desirable.We present a novel method to model prosody independently of the text and thus independently of the language as well.For this purpose, the voiced and unvoiced speech segments are extracted and a 187-dimensional feature vector is computed for each voiced segment.This approach is compared to word based prosodic features on a German text passage.Both are confronted with the perceptive evaluation of two native speakers of German.The word-based feature set yielded correlations of up to 0.92 while the text-independent feature set yielded 0.88.This is in the same range as the inter-rater correlation with 0.88.Furthermore, the text-independent features were computed for a Japanese translation of the passage which was also rated by two native speakers of Japanese.Again, the correlation between the automatic system and the human perception of the naturalness was high with 0.83 and not significantly lower than the inter-rater correlation of 0.92. Andreas K. Maier, Florian Hönig, Viktor Zeißler, Anton Batliner, Erik Körner, Nobuyuki Yamanaka, Peter Ackermann, Elmar Nöth |
INTERSPEECH | 1 |
| 2009 | A microphone-independent visualization technique for speech disordersabstractIn this paper we introduce a novel method for the visualization of speech disorders. We demonstrate the method with disordered speech and a control group. However, both groups were recorded using two different microphones. The projection of the patient data using a single microphone yields significant correlations between the coordinates on the map and certain criteria of the disorder which were perceptually rated. However, projection of data from multiple microphones reduces this correlation. Usually, the acoustical mismatch between the microphones is greater than the mismatch between the speakers, i.e., not the disorders but the microphones form clusters in the visualization. Based on an extension of the Sammon mapping, we are able to create a map which projects the same speakers onto the same position even if multiple microphones are used. Furthermore, our method also restores the correlation between the map coordinates and the perceptual assessment. Index Terms: visualization, robustness, speech processing. Andreas K. Maier, Stefan Wenhardt, Tino Haderlein, Maria Schuster, Elmar Nöth |
INTERSPEECH | 1 |
| 2009 | Intelligibility assessment in children with cleft lip and palate in Italian and GermanabstractCurrent research has shown that the speech intelligibility in children with cleft lip and palate (CLP) can be estimated automatically using speech recognition methods. On German CLP data high and significant correlations between human ratings and the recognition accuracy of a speech recognition system were already reported. In this paper we investigate whether the approach is also suitable for other languages. Therefore, we compare the correlations obtained on German data with the correlations on Italian data. A high and significant correlation (r=0.76; p 0.05). Index Terms: speech recognition, speech intelligibility, cleft lip and palate. Marcello Scipioni, Matteo Gerosa, Diego Giuliani, Elmar Nöth, Andreas K. Maier |
INTERSPEECH | 5 |
| 2009 | PEAKS - A system for the automatic evaluation of voice and speech disorders
Andreas K. Maier, Tino Haderlein, Ulrich Eysholdt, Frank Rosanowski, Anton Batliner, Maria Schuster, Elmar Nöth |
Speech Commun. | 1 |
| 2008 | Age and gender recognition for telephone applications based on GMM supervectors and support vector machinesabstractThis paper compares two approaches of automatic age and gender classification with 7 classes. The first approach are Gaussian mixture models (GMMs) with universal background models (UBMs), which is well known for the task of speaker identification/verification. The training is performed by the EM algorithm or MAP adaptation respectively. For the second approach for each speaker of the test and training set a GMM model is trained. The means of each model are extracted and concatenated, which results in a GMM supervector for each speaker. These supervectors are then used in a support vector machine (SVM). Three different kernels were employed for the SVM approach: a polynomial kernel (with different polynomials), an RBF kernel and a linear GMM distance kernel, based on the KL divergence. With the SVM approach we improved the recognition rate to 74% (p < 0.001) and are in the same range as humans. Tobias Bocklet, Andreas K. Maier, Josef G. Bauer, Felix Burkhardt, Elmar Nöth |
ICASSP | 2 |
| 2008 | Automatic evaluation of characteristic speech disorders in children with cleft lip and palateabstractAbstract This paper discusses the automatic evaluation of speech of chil-dren with cleft lip and palate (CLP). CLP speech shows specialcharacteristics such as hypernasality, backing, and weakeningof plosives. In total ve criteria were subjectively assessed byan experienced speech expert on the phone level. This subjec-tive evaluation was used as a gold standard to train a classi-cation system. The automatic system achieves recognition re-sults on frame, phone, and word level of up to 75.8% CL. Onspeaker level signicant and high correlations between the sub-jective evaluation and the automatic system of up to 0.89 areobtained. Index Terms : pathologic speech, speech assessment, pronun-ciation scoring, children’s speech 1. Introduction Cleft Lip and Palate (CLP) is the most common malformationof the head. It constitutes almost two-thirds of the major facialdefects and almost 80% of all orofacial clefts [1]. Its prevalencediffers in different populations from 1 in 400 to 500 newborns inAsians to 1 in 1500 to 2000 in African Americans. The preva-lence in Caucasians is 1 in 750 to 900 births [2, 3].In clinical practice, articulation disorders are mainly eval-uated by subjective tools. The simplest method is the audi-tive perception, mostly performed by a speech therapist. Pre-vious studies have shown that experience is an important fac-tor that inuences the subjective estimation of speech disorderswhich leads to inaccurate evaluation by persons with only fewyears of experience as speech therapist [4]. Until now, objectivemeans exist only for quantitative measurements of nasal emis-sions [5, 6, 7] and for the detection of secondary voice disorders[8]. But other specic articulation disorders in CLP cannot besufciently quantied.In this paper, we present a new technical procedure for themeasurement and evaluation of specic speech disorders andcompare the results obtained with subjective ratings of an expe-rienced speech therapist. Andreas K. Maier, Florian Hönig, Christian Hacker, Maria Schuster, Elmar Nöth |
INTERSPEECH | 1 |
| 2007 | Towards robust automatic evaluation of pathologic telephone speechabstractFor many aspects of speech therapy an objective evaluation of the intelligibility of a patient's speech is needed. We investigate the evaluation of the intelligibility of speech by means of automatic speech recognition. Previous studies have shown that measures like word accuracy are consistent with human experts' ratings. To ease the patient's burden, it is highly desirable to conduct the assessment via phone. However, the telephone channel influences the quality of the speech signal which negatively affects the results. To reduce inaccuracies, we propose a combination of two speech recognizers. Experiments on two sets of pathological speech show that the combination results in consistent improvements in the correlation between the automatic evaluation and the ratings by human experts. Furthermore, the approach leads to reductions of 10% and 25% of the maximum error of the intelligibility measure. Korbinian Riedhammer, Georg Stemmer, Tino Haderlein, Maria Schuster, Frank Rosanowski, Elmar Nöth, Andreas K. Maier |
ASRU | 7 |
| 2007 | Boosting of Prosodic and Pronunciation Features to Detect Mispronunciations of Non-Native ChildrenabstractCommercial products that support L2-learners with computer assisted pronunciation training usually focus per exercise only on one possible pronunciation mistake that is typical for speakers of the respective L1 group. Acoustic models for words with wrong pronunciation are added to the system. In the present paper a more general approach with features that have proved to be widely independent of the learners' mother tongue is proposed. It is able to take various possible mistakes into consideration all at once. High dimensional feature vectors that encode prosodic varieties and differences of reference and recognized sentences are analyzed. With the ADABOOST algorithm those features are found, which contain the most important information to assess German children learning English. With 35 features 89 % of the agreement of experts is achieved. Christian Hacker, Tobias Cincarek, Andreas K. Maier, Andre Heßler, Elmar Nöth |
ICASSP (4) | 3 |
| 2007 | Towards More Reality in the Recognition of Emotional SpeechabstractAs automatic emotion recognition based on speech matures, new challenges can be faced. We therefore address the major aspects in view of potential applications in the field, to benchmark today's emotion recognition systems and bridge the gap between commercial interest and current performances: acted vs. spontaneous speech, realistic emotions, noise and microphone conditions, and speaker independence. Three different data-sets are used: the Berlin Emotional Speech Database, the Danish Emotional Speech Database, and the spontaneous AIBO Emotion Corpus. By using different feature types such as word- or turn-based statistics, manual versus forced alignment, and optimization techniques we show how to best cope with this demanding task and how noise addition or different microphone positions affect emotion recognition. Björn W. Schuller, Dino Seppi, Anton Batliner, Andreas K. Maier, Stefan Steidl |
ICASSP (4) | 4 |
| 2007 | Automatic scoring of the intelligibility in patients with cancer of the oral cavityabstractAfter surgical treatment of cancer of the oral cavity patients often suffer from functional restrictions such as speech disorders.In this paper we present a novel approach to assess the outcome of the treatment w.r.t. the intelligibility of the patient using the result of an automatic speech recognition system.The word recognition rate was taken as intelligibility score.Compared to four speech experts this method yields results that are as good as the best speech expert compared to the other experts.The correlation between our system and the mean opinion of the experts is .92.Furthermore we show that our system has better performance than the average expert and is more reliable. Andreas K. Maier, Maria Schuster, Anton Batliner, Elmar Nöth, Emeka Nkenke |
INTERSPEECH | 1 |