VLDB 2026 Research / reviewers in the wild / expert
Dorit Merhof
dblp:63/3515
· DBLP profile ↗
67ranked-venue papers
3as first author
30since 2021 · last 2026
0000-0002-1672-2185ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 2 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 7 since 2021Systems, architecture and hardware · 12 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-level fusion network for two-stage polyp segmentation via integrity learning
Junzhuo Liu 0001, Dorit Merhof |
Image Vis. Comput. | 2 |
| 2026 | Spatial transcriptomics expression prediction from histopathology based on cross-modal mask reconstruction and contrastive learningabstractSpatial transcriptomics is a technology that captures gene expression at different spatial locations, widely used in tumor microenvironment analysis and molecular profiling of histopathology, providing valuable insights into resolving gene expression and clinical diagnosis of cancer. Due to the high cost of data acquisition, large-scale spatial transcriptomics data remain challenging to obtain. In this study, we develop a contrastive learning-based deep learning method to predict spatially resolved gene expression from the whole-slide images (WSIs). Unlike existing end-to-end prediction frameworks, our method leverages multi-modal contrastive learning to establish a correspondence between histopathological morphology and spatial gene expression in the feature space. By computing cross-modal feature similarity, our method generates spatially resolved gene expression directly from WSIs. Furthermore, to enhance the standard contrastive learning paradigm, a cross-modal masked reconstruction is designed as a pretext task, enabling feature-level fusion between modalities. Notably, our method does not rely on large-scale pretraining datasets or abstract semantic representations from either modality, making it particularly effective for scenarios with limited spatial transcriptomics data. Evaluation across six different disease datasets demonstrates that, compared to existing studies, our method improves Pearson Correlation Coefficient (PCC) in the prediction of highly expressed genes, highly variable genes, and marker genes by 6.27 %, 6.11 %, and 11.26 % respectively. Further analysis indicates that our method preserves gene-gene correlations and applies to datasets with limited samples. Additionally, our method exhibits potential in cancer tissue localization based on biomarker expression. The code repository for this work is available at https://github.com/ngfufdrdh/CMRCNet. Junzhuo Liu 0001, Markus Eckstein, Friedrich Feuerhake, Dorit Merhof |
Medical Image Anal. | 5 |
| 2025 | SL2 A-INR: Single-Layer Learnable Activation for Implicit Neural Representation
Reza Rezaeian, Moein Heidari, Reza Azad, Dorit Merhof, Hamid Soltanian-Zadeh, Ilker Hacihaliloglu |
ICCV | 4 |
| 2025 | HoloPointNet: A Deep Learning Framework for Efficient 3D Point Cloud Holography
Ankit Amrutkar, Ahmet Nazlioglu, Björn Kampa, Volkmar Schulz, Johannes Stegmaier, Markus Rothermel, Dorit Merhof |
MICCAI (11) | 7 |
| 2025 | CENet: Context Enhancement Network for Medical Image Segmentation
Afshin Bozorgpour, Sina Ghorbani Kolahi, Reza Azad, Ilker Hacihaliloglu, Dorit Merhof |
MICCAI (1) | 5 |
| 2025 | Top-Down Attention-Based Multiple Instance Learning for Whole Slide Image Analysis
Daniel Reisenbüchler, Ruining Deng, Christian Matek, Friedrich Feuerhake, Dorit Merhof |
MICCAI (1) | 5 |
| 2025 | LHU-Net: A Lean Hybrid U-Net for Cost-Efficient, High-Performance Volumetric Segmentation
Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof |
MICCAI (14) | 5 |
| 2025 | Frequency-Domain Refinement of Vision Transformers for Robust Medical Image Segmentation Under DegradationabstractMedical image segmentation is crucial for precise diagnosis, treatment planning, and disease monitoring in clinical settings. While convolutional neural networks (CNNs) have achieved remarkable success, they struggle with modeling long-range dependencies. Vision Transformers (ViTs) address this limitation by leveraging self-attention mechanisms to capture global contextual information. However, ViTs often fall short in local feature description, which is crucial for precise segmentation. To address this issue, we reformulate self-attention in the frequency domain to enhance both local and global feature representation. Our approach, the Enhanced Wave Vision Transformer (EW-ViT), incorporates wavelet decomposition within the self-attention block to adaptively refine feature representation in low and high-frequency components. We also introduce the Prompt-Guided High-Frequency Refiner (PGHFR) module to handle image degradation, which mainly affects high-frequency components. This module uses implicit prompts to encode degradation-specific information and adjust high-frequency representations accordingly. Additionally, we apply a contrastive learning strategy to maintain feature consistency and ensure robustness against noise, leading to state-of-the-art (SOTA) performance in medical image segmentation, especially under various conditions of degradation. Source code is available at GitHub. Sanaz Karimijafarbigloo, Sina Ghorbani Kolahi, Reza Azad, Ulas Bagci, Dorit Merhof |
WACV | 5 |
| 2025 | Addressing Missing Modality Challenges in MRI Images: A Comprehensive ReviewabstractMagnetic resonance imaging (MRI) is one of the most prevalent imaging modalities used for diagnosis, treatment planning, and outcome control in various medical conditions. MRI sequences provide physicians with the ability to view and monitor tissues at multiple contrasts within a single scan and serve as input for automated systems to perform downstream tasks. However, in clinical practice, there is usually no concise set of identically acquired sequences for a whole group of patients. As a consequence, medical professionals and automated systems both face difficulties due to the lack of complementary information from such missing sequences. This problem is well known in computer vision, particularly in medical image processing tasks such as tumor segmentation, tissue classification, and image generation. With the aim of helping researchers, this literature review examines a significant number of recent approaches that attempt to mitigate these problems. Basic techniques such as early synthesis methods, as well as later approaches that deploy deep learning, such as common latent space models, knowledge distillation networks, mutual information maximization, and generative adversarial networks (GANs) are examined in detail. We investigate the novelty, strengths, and weaknesses of the aforementioned strategies. Moreover, using a case study on the segmentation task, our survey offers quantitative benchmarks to further analyze the effectiveness of these methods for addressing the missing modalities challenge. Furthermore, a discussion offers possible future research directions. Reza Azad, Mohammad Dehghanmanshadi, Nika Khosravi, Julien Cohen-Adad, Dorit Merhof |
Comput. Vis. Media | 5 |
| 2025 | TransCeption: Enhancing Medical Image Segmentation with an Inception-Like Transformer Design for Efficient Feature FusionabstractWhile CNN-based methods have been the cornerstone of medical image segmentation due to their promising performance and robustness, they suffer from limitations in capturing long-range dependencies. Transformer-based approaches are currently prevailing since they enlarge the receptive field to model global contextual correlations. To further extract rich representations, some extensions of U-Net employ multi-scale feature extraction and fusion modules to obtain improved performance. Inspired by this idea, we propose TransCeption for medical image segmentation, a pure transformer-based U-shaped network incorporating an inception-like module in the encoder and adopting a contextual bridge for better feature fusion. The design proposed in this work is based on three core principles. (i) The patch merging module in the encoder is redesigned to use ResInception Patch Merging (RIPM). The MultiBranch (MB) transformer has the same number of branches as the outputs of RIPM. Combining the two modules enables the model to capture a multi-scale representation within a single stage. (ii) We apply an Intra-stage Feature Fusion (IFF) module following the MB transformer to enhance the aggregation of feature maps from all branches and particularly focus on the interaction between the different channels at all scales. (iii) In contrast to a bridge that only contains token-wise self-attention, we propose a Dual Transformer Bridge that also includes channel-wise self-attention to exploit correlations between scales at different stages from a dual perspective. Extensive experiments on multi-organ and skin lesion segmentation tasks show the superiority of TransCeption to previous work. The code is publicly available on GitHub. Reza Azad, Yiwei Jia, Ehsan Khodapanah Aghdam, Julien Cohen-Adad, Dorit Merhof |
Comput. Vis. Media | 5 |
| 2025 | MedScale-Former: Self-guided multiscale transformer for medical image segmentation
Sanaz Karimijafarbigloo, Reza Azad, Amirhossein Kazerouni, Dorit Merhof |
Medical Image Anal. | 4 |
| 2025 | Continual learning in medical image analysis: A comprehensive review of recent advancements and future prospectsabstractMedical image analysis has witnessed remarkable advancements, even surpassing human-level performance in recent years, driven by the rapid development of advanced deep-learning algorithms. However, when the inference dataset slightly differs from what the model has seen during one-time training, the model performance is greatly compromised. The situation requires restarting the training process using both the old and the new data, which is computationally costly, does not align with the human learning process, and imposes storage constraints and privacy concerns. Alternatively, continual learning has emerged as a crucial approach for developing unified and sustainable deep models to deal with new classes, tasks, and the drifting nature of data in non-stationary environments for various application areas. Continual learning techniques enable models to adapt and accumulate knowledge over time, which is essential for maintaining performance on evolving datasets and novel tasks. Owing to its popularity and promising performance, it is an active and emerging research topic in the medical field and hence demands a survey and taxonomy to clarify the current research landscape of continual learning in medical image analysis. This systematic review paper provides a comprehensive overview of the state-of-the-art in continual learning techniques applied to medical image analysis. We present an extensive survey of existing research, covering topics including catastrophic forgetting, data drifts, stability, and plasticity requirements. Further, an in-depth discussion of key components of a continual learning framework, such as continual learning scenarios, techniques, evaluation schemes, and metrics, is provided. Continual learning techniques encompass various categories, including rehearsal, regularization, architectural, and hybrid strategies. We assess the popularity and applicability of continual learning categories in various medical sub-fields like radiology and histopathology. Our exploration considers unique challenges in the medical domain, including costly data annotation, temporal drift, and the crucial need for benchmarking datasets to ensure consistent model evaluation. The paper also addresses current challenges and looks ahead to potential future research directions. Pratibha Kumari 0001, Joohi Chauhan, Afshin Bozorgpour, Boqiang Huang, Reza Azad, Dorit Merhof |
Medical Image Anal. | 6 |
| 2024 | MSA2Net: Multi-scale Adaptive Attention-guided Network for Medical Image Segmentation
Sina Ghorbani Kolahi, Seyed Kamal Chaharsooghi, Toktam Khatibi, Afshin Bozorgpour, Reza Azad, Moein Heidari, Ilker Hacihaliloglu, Dorit Merhof |
BMVC | 8 |
| 2024 | Continual Domain Incremental Learning for Privacy-Aware Digital Pathology
Pratibha Kumari 0001, Daniel Reisenbüchler, Lucas Luttner, Nadine S. Schaadt, Friedrich Feuerhake, Dorit Merhof |
MICCAI (12) | 6 |
| 2024 | Unsupervised Latent Stain Adaptation for Computational Pathology
Daniel Reisenbüchler, Lucas Luttner, Nadine S. Schaadt, Friedrich Feuerhake, Dorit Merhof |
MICCAI (11) | 5 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 30 |
| 2024 | Beyond Self-Attention: Deformable Large Kernel Attention for Medical Image SegmentationabstractMedical image segmentation has seen significant improvements with transformer models, which excel in grasping far-reaching contexts and global contextual information. However, the increasing computational demands of these models, proportional to the squared token count, limit their depth and resolution capabilities. Most current methods process D volumetric image data slice-by-slice (called pseudo 3D), missing crucial inter-slice information and thus reducing the model’s overall performance. To address these challenges, we introduce the concept of Deformable Large Kernel Attention (D-LKA Attention), a streamlined attention mechanism employing large convolution kernels to fully appreciate volumetric context. This mechanism operates within a receptive field akin to self-attention while sidestepping the computational overhead. Additionally, our proposed attention mechanism benefits from deformable convolutions to flexibly warp the sampling grid, enabling the model to adapt appropriately to diverse data patterns. We designed both 2D and 3D adaptations of the D-LKA Attention, with the latter excelling in cross-depth data understanding. Together, these components shape our novel hierarchical Vision Transformer architecture, the D-LKA Net. Evaluations of our model against leading methods on popular medical segmentation datasets (Synapse, NIH Pancreas, and Skin lesion) demonstrate its superior performance. Our code is publicly available at GitHub. Reza Azad, Leon Niggemeier, Michael Huttemann, Amirhossein Kazerouni, Ehsan Khodapanah Aghdam, Yury Velichko, Ulas Bagci, Dorit Merhof |
WACV | 8 |
| 2024 | INCODE: Implicit Neural Conditioning with Prior Knowledge EmbeddingsabstractImplicit Neural Representations (INRs) have revolutionized signal representation by leveraging neural networks to provide continuous and smooth representations of complex data. However, existing INRs face limitations in capturing fine-grained details, handling noise, and adapting to diverse signal types. To address these challenges, we introduce INCODE, a novel approach that enhances the control of the sinusoidal-based activation function in INRs using deep prior knowledge. INCODE comprises a harmonizer network and a composer network, where the harmonizer network dynamically adjusts key parameters of the activation function. Through a task-specific pre-trained model, INCODE adapts the task-specific parameters to optimize the representation process. Our approach not only excels in representation, but also extends its prowess to tackle complex tasks such as audio, image, and 3D shape reconstructions, as well as intricate challenges such as neural radiance fields (NeRFs), and inverse problems, including denoising, super-resolution, inpainting, and CT reconstruction. Through comprehensive experiments, INCODE demonstrates its superiority in terms of robustness, accuracy, quality, and convergence rate, broadening the scope of signal representation. Please visit the project’s website for details on the proposed method and access to the code. Amirhossein Kazerouni, Reza Azad, Alireza Hosseini, Dorit Merhof, Ulas Bagci |
WACV | 4 |
| 2024 | Advances in medical image analysis with vision Transformers: A comprehensive review
Reza Azad, Amirhossein Kazerouni, Moein Heidari, Ehsan Khodapanah Aghdam, Amirali Molaei, Yiwei Jia, Abin Jose, Rijo Roy, Dorit Merhof |
Medical Image Anal. | 9 |
| 2024 | Medical Image Segmentation Review: The Success of U-NetabstractAutomatic medical image segmentation is a crucial topic in the medical domain and successively a critical counterpart in the computer-aided diagnosis paradigm. U-Net is the most widespread image segmentation architecture due to its flexibility, optimized modular design, and success in all medical image modalities. Over the years, the U-Net model has received tremendous attention from academic and industrial researchers who have extended it to address the scale and complexity created by medical tasks. These extensions are commonly related to enhancing the U-Net's backbone, bottleneck, or skip connections, or including representation learning, or combining it with a Transformer architecture, or even addressing probabilistic prediction of the segmentation map. Having a compendium of different previously proposed U-Net variants makes it easier for machine learning researchers to identify relevant research questions and understand the challenges of the biological tasks that challenge the model. In this work, we discuss the practical aspects of the U-Net model and organize each variant model into a taxonomy. Moreover, to measure the performance of these strategies in a clinical application, we propose fair evaluations of some unique and famous designs on well-known datasets. Furthermore, we provide a comprehensive implementation library with trained models. In addition, for ease of future studies, we created an online list of U-Net papers with their possible official implementation. Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland, Yiwei Jia, Atlas Haddadi Avval, Afshin Bozorgpour, Sanaz Karimijafarbigloo, Joseph Paul Cohen, Ehsan Adeli-Mosabbeb, Dorit Merhof |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2023 | Laplacian-Former: Overcoming the Limitations of Vision Transformers in Local Texture Detection
Reza Azad, Amirhossein Kazerouni, Babak Azad, Ehsan Khodapanah Aghdam, Yury Velichko, Ulas Bagci, Dorit Merhof |
MICCAI (3) | 7 |
| 2023 | HiFormer: Hierarchical Multi-scale Representations Using Transformers for Medical Image SegmentationabstractConvolutional neural networks (CNNs) have been the consensus for medical image segmentation tasks. However, they suffer from the limitation in modeling long-range dependencies and spatial correlations due to the nature of convolution operation. Although transformers were first developed to address this issue, they fail to capture low-level features. In contrast, it is demonstrated that both local and global features are crucial for dense prediction, such as segmenting in challenging contexts. In this paper, we propose HiFormer, a novel method that efficiently bridges a CNN and a transformer for medical image segmentation. Specifically, we design two multi-scale feature representations using the seminal Swin Transformer module and a CNN-based encoder. To secure a fine fusion of global and local features obtained from the two aforementioned representations, we propose a Double-Level Fusion (DLF) module in the skip connection of the encoder-decoder structure. Extensive experiments on various medical image segmentation datasets demonstrate the effectiveness of HiFormer over other CNN-based, transformer-based, and hybrid methods in terms of computational complexity, quantitative and qualitative results. Our code is publicly available at GitHub. Moein Heidari, Amirhossein Kazerouni, Milad Soltany Kadarvish, Reza Azad, Ehsan Khodapanah Aghdam, Julien Cohen-Adad, Dorit Merhof |
WACV | 7 |
| 2023 | The Liver Tumor Segmentation Benchmark (LiTS)abstractIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094. Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze |
Medical Image Anal. | 75 |
| 2023 | SegPC-2021: A challenge & dataset on segmentation of Multiple Myeloma plasma cells from microscopic images
Anubha Gupta, Shiv Gehlot, Shubham Goswami, Sachin Motwani, Álvaro García-Faura, Dejan Stepec, Tomaz Martincic, Reza Azad, Dorit Merhof, Afshin Bozorgpour, Babak Azad, Alaa Sulaiman, Deepanshu Pandey, Pradyumna Gupta, Sumit Bhattacharya, Aman Sinha 0002, Xinyun Qiu, Yoonbeom Park, Dae-Hong Lee, Joon Sik Park, KwangYeol Lee, Jaehyung Ye |
Medical Image Anal. | 10 |
| 2023 | Diffusion models in medical imaging: A comprehensive survey
Amirhossein Kazerouni, Ehsan Khodapanah Aghdam, Moein Heidari, Reza Azad, Mohsen Fayyaz, Ilker Hacihaliloglu, Dorit Merhof |
Medical Image Anal. | 7 |
| 2022 | Instance Segmentation of Dense and Overlapping Objects via Layering
Long Chen 0014, Yuli Wu 0001, Dorit Merhof |
BMVC | 3 |
| 2022 | Automatic Embedding Interventions for the Classification of Hematopoietic CellsabstractThe classification of hematopoietic cells is the most essential step in automating the analysis of human bone marrow samples. However, the complex structure of cell classes as well as class imbalance make this a challenging task, even for neural networks. Based on projective latent interventions, we propose automatic interventions that iteratively update a learned embedding with suitable transformations that shift different cell types apart and contract samples of the same type together. We present different ways of applying these: either directly on a higher-dimensional embedding or in a parametric version in two dimensions. We analyze the hyper-parameters and evaluate the proposed approach on a challenging dataset of hematopoietic cells. The results show an improvement of up to 3 percentage points for the classification F-score. Philipp Gräbel, Julian Thull, Martina Crysandt, Barbara Mara Klinkhammer, Peter Boor, Tim H. Brümmendorf, Dorit Merhof |
ICPR | 7 |
| 2021 | Anomaly Detection for the Automated Visual Inspection of PET Preform ClosuresabstractRecent advances in supervised Machine Learning have enabled the automated visual inspection of increasingly complex products. The economic feasibility of developed methods is, however, limited by their requirement for large amounts of labeled training data. As an alternative, Anomaly Detection (AD) methods have been proposed, which do not rely on large-scale datasets and often require defect-free images only. In our work, we benchmark and adapt proposed state-of-the-art AD as well as Anomaly Segmentation (AS) methods to the Polyethylene Terephthalate (PET) preform closure inspection task in a high-throughput setting. We furthermore compare AD/AS methods with fully supervised baselines and show that, in the low-data regime, comparable AS performance can be achieved. Moreover, fully supervised baselines are outperformed with respect to AD by our adapted methods. We also conduct ablation studies to quantify the effects of proposed modifications, and show that they reduce memory footprint by a factor of 3 while maintaining AD/AS performance. Together, this demonstrates the suitability of AD/AS methods for the automated visual inspection of complex objects such as transparent PET preform closures in high-throughput settings. Oliver Rippel, Peter Haumering, Johannes Brauers, Dorit Merhof |
ETFA | 4 |
| 2021 | Leveraging pre-trained Segmentation Networks for Anomaly SegmentationabstractEmploying representations generated by large-scale training in a transfer-learning setting achieves state-of-the-art anomaly segmentation results when applied to the visual inspection task. Current approaches, however, focus exclusively on features of pre-trained classification networks, which are known to posess lower spatial resolution than segmentation or object detection networks. In our work, we investigate whether features extracted from pre-trained segmentation networks can be used to further improve anomaly segmentation performance in the transfer-learning setting. To this end, we apply state-of-the-art transfer-learning methods to encoder-decoder based segmentation networks. Results show that the encoders of pre-trained segmentation networks yield improved anomaly segmentation performance compared to their pre-trained classification counterparts. However, no consistent improvements can be observed yet regarding the decoders of the pre-trained segmentation networks. Together, this demonstrates that pre-trained segmentation networks can be used to further improve transfer-learned anomaly segmentation performance and that additional research is required to fully unleash their potential. Oliver Rippel, Dorit Merhof |
ETFA | 2 |
| 2021 | Estimating the Probability Density Function of New Fabrics for Fabric Anomaly Detection
Oliver Rippel, Maximilian Müller, Andreas Münkel, Thomas Gries, Dorit Merhof |
ICPRAM | 5 |
| 2020 | Accurate Stitch Position Identification of Sewn Threads in TextilesabstractSeam stitches made by a computer program on a CNC sewing machine may be shifted from their desired positions due to the elasticity of thread and substrate material. The mismatch may lead to a visually inferior appearance of the resulting seam pattern compared to the one desired by the manufacturer.This paper presents a software to accurately identify the distortion vector for each individual stitch. The proposed software consists of a thread pixel detection in the scanned image of the textile based upon the Frangi Filter and the Expectation Maximization algorithm. A Subsequent registration with the desired model finds an accurate assignment of each model stitch position and the corresponding thread pixel. This allows the computation of individual distortion vectors facilitating the automatic reprogramming of corrected stitch positions in the CNC sewing program to minimize the final distortion. Dieter Geller, Tarek Stiebel, Oliver Rippel, Jonas Osburg, Volker Lutz, Thomas Gries, Dorit Merhof |
ETFA | 7 |
| 2020 | GAN-based Defect Synthesis for Anomaly Detection in FabricsabstractImage-based quality control aims at detecting anomalies in products and is a crucial part of the production process. Challenges arise from the complexity and variety of products, defects, and the rarity of defect occurrence. Quality control thus still relies heavily on manual inspection. Supervised, data driven approaches have greatly improved performance, but suffer from a major drawback: They require large amounts of annotated training data, limiting their economic viability.In this work, we overcome this drawback by leveraging the consistency of defect appearance across fabrics to transfer knowledge about anomalies from one fabric to another. We realize this by adapting the image-to-image translation framework, introducing guidance by means of a segmentation map. We evaluate both image quality as well as usability of generated defects. For the image quality, both classical (L1, L2 and Structured Similarity Measurement (SSIM)) and perceptual (Learned Perceptual Image Patch Similarity (LPIPS)) metrics indicate efficacy of the presented approach, i.e. defect synthesis only occurs in the targeted regions, and the background fabric pattern remains largely unchanged. To evaluate usability of generated defects, we train pseudo-supervised models on synthesized defects of fabrics unseen during training. A comparison with semi-supervised, autoencoder based approaches demonstrates the suitability of our approach, yielding average Area Under the Receiver Operating Characteristic Curve (AUROC) values of 0.81 for pseudo-supervised vs. 0.69 for semi-supervised settings. Oliver Rippel, Maximilian Müller, Dorit Merhof |
ETFA | 3 |
| 2020 | Modeling the Distribution of Normal Data in Pre-Trained Deep Features for Anomaly DetectionabstractAnomaly Detection (AD) in images is a fundamental computer vision problem and refers to identifying images and/or image substructures that deviate significantly from the norm. Popular AD algorithms commonly try to learn a model of normality from scratch using task specific datasets, but are limited to semi-supervised approaches employing mostly normal data due to the inaccessibility of anomalies on a large scale combined with the ambiguous nature of anomaly appearance. We follow an alternative approach and demonstrate that deep feature representations learned by discriminative models on large natural image datasets are well suited to describe normality and detect even subtle anomalies in a transfer learning setting. Our model of normality is established by fitting a multivariate Gaussian (MVG) to deep feature representations of classification networks trained on ImageNet using normal data only. By subsequently applying the Mahalanobis distance as the anomaly score we outperform the current state of the art on the public MVTec AD dataset, achieving an Area Under the Receiver Operating Characteristic curve of 95.8 ± 1.2% (mean ± SEM) over all 15 classes. We further investigate why the learned representations are discriminative to the AD task using Principal Component Analysis. We find that the principal components containing little variance in normal data are the ones crucial for discriminating between normal and anomalous instances. This gives a possible explanation to the often subpar performance of AD approaches trained from scratch using normal data only. By selectively fitting a MVG to these most relevant components only, we are able to further reduce model complexity while retaining AD performance. We also investigate setting the working point by selecting acceptable False Positive Rate thresholds based on the MVG assumption. Code is publicly available at https://github.com/ORippler/gaussian-ad-mvtec. Oliver Rippel, Patrick Mertens, Dorit Merhof |
ICPR | 3 |
| 2019 | Interpolating Maps between Neural Response Spaces for Chemosensing with Fruit Fly Antenna SensorsabstractThe odorant receptor neurons on the fruit fly antenna are highly sensitive to a broad range of chemicals. A compound signal of receptor activity on the antenna can be read out in real time with functional neuroimaging, and individual receptor responses to hundreds of odorants are available in a database. Utilizing the fruit fly antenna as chemosensor enables applications ranging from biomarker detection to identification of unknown chemicals in samples. Here, we propose to connect neural response spaces, mapping odorant responses from one fly to another and to database space. A map is defined exactly for reference odorants common to both subject and target space, while the map for the remaining odorants is estimated based on radial basis function interpolation. On a data set with chemically diverse odorants, mapping to another antenna allows identifying unlabelled subject space odorants by the proximity of their mapped position to labelled odorants in target space. Furthermore, mapping from antenna to database space predicts the individual receptor responses significantly better than a random baseline model, suggesting that receptor responses can be inferred from the compound antenna signal given a sufficiently dense net of reference odorants to support the map. Martin Strauch, Karl Krüger, Latha Mukunda, Alja Lüdke, C. Giovanni Galizia, Dorit Merhof |
BIBE | 6 |
| 2019 | Instance Segmentation of Biomedical Images with an Object-Aware Embedding Learned with Local Constraints
Long Chen 0014, Martin Strauch, Dorit Merhof |
MICCAI (1) | 3 |
| 2019 | Combined Learning for Similar Tasks with Domain-Switching Networks
Daniel Bug, Dennis Eschweiler, Justus Schock, Leon Weninger, Friedrich Feuerhake, Julia Schüler, Johannes Stegmaier, Dorit Merhof |
MICCAI (5) | 9 |
| 2019 | GAN-Based Image Enrichment in Digital Pathology Boosts Segmentation Accuracy
Laxmi Gupta, Barbara Mara Klinkhammer, Peter Boor, Dorit Merhof, Michael Gadermayr |
MICCAI (1) | 4 |
| 2019 | Multi Scale Curriculum CNN for Context-Aware Breast MRI Malignancy Classification
Christoph Haarburger, Michael Baumgartner 0001, Daniel Truhn, Mirjam Broeckmann, Hannah Schneider, Simone Schrading, Christiane Kuhl, Dorit Merhof |
MICCAI (4) | 8 |
| 2019 | Generative Adversarial Networks for Facilitating Stain-Independent Supervised and Unsupervised Segmentation: A Study on Kidney HistologyabstractA major challenge in the field of segmentation in digital pathology is given by the high effort for manual data annotations in combination with many sources introducing variability in the image domain. This requires methods that are able to cope with variability without requiring to annotate a large amount of samples for each characteristic. In this paper, we develop approaches based on adversarial models for image-to-image translation relying on unpaired training. Specifically, we propose approaches for stain-independent supervised segmentation relying on image-to-image translation for obtaining an intermediate representation. Furthermore, we develop a fully-unsupervised segmentation approach exploiting image-to-image translation to convert from the image to the label domain. Finally, both approaches are combined to obtain optimum performance in unsupervised segmentation independent of the characteristics of the underlying stain. Experiments on patches showing kidney histology proof that stain-translation can be performed highly effectively and can be used for domain adaptation to obtain independence of the underlying stain. It is even capable of facilitating the underlying segmentation task, thereby boosting the accuracy if an appropriate intermediate stain is selected. Combining domain adaptation with unsupervised segmentation finally showed the most significant improvements. Michael Gadermayr, Laxmi Gupta, Vitus Appel, Peter Boor, Barbara Mara Klinkhammer, Dorit Merhof |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Recovering a Chemotopic Feature Space from a Group of Fruit Fly Antenna ChemosensorsabstractThe ensemble of odorant receptors on the antenna of the fruit fly Drosophila melanogaster acts as an encoder for chemical molecules. Chemically similar odorants elicit activity in similar subsets of the receptors, spanning a so-called chemotopic feature space that enables chemical similarity search. A compound signal of receptor activity can be read out by calcium imaging of the antenna, yet without revealing corresponding receptors on different antennae. Employing Canonical Correlation Analysis (CCA) for multiple sets, we show that a consensus feature space can nevertheless be recovered from a group of variable antenna sensors that all respond to a common sequence of odorants. In the chemotopic consensus feature space, properties of novel odorants can be inferred, demonstrating how fruit fly antenna chemosensors may be employed as an alternative to electronic noses. Martin Strauch, Latha Mukunda, Alja Lüdke, C. Giovanni Galizia, Dorit Merhof |
BIBE | 5 |
| 2018 | An Inspection System for Multi-Label Polymer ClassificationabstractWaste treatment, especially treatment of plastic waste, is arguably one of the biggest challenges that humanity faces in context of preserving the environment besides global warming. This work presents a visual inspection system for plastic classification and proposes a classification algorithm that is based on near-infrared spectroscopy and convolutional neural networks. The method allows for a highly accurate classification of several main polymer types while being robust against image disturbances occurring in a real world scenario. Most importantly, it is able to cope with layers of multiple materials. This work therefore offers for the very first time a solution to multi-material classification in the context of plastic recycling. Since the manual creation and annotation of layered materials is a cumbersome task due to the manifold of possible combinations, it is also shown how the creation of artificial data can greatly facilitate the ground truth generation. Tarek Stiebel, Marcel Bosling, Aljoscha Steffens, Thomas Pretz, Dorit Merhof |
ETFA | 5 |
| 2018 | Correction of Thread Position Mismatch in High Precision CNC SewingabstractDue to the elasticity of thread and tissue material, the stitches made by a computer program on a CNC sewing machine are shifted from their desired positions. The mismatch leads to a visually inferior appearance of the resulting thread pattern compared to the one intended by the manufacturer. This paper introduces a camera based inspection system that helps to find and evaluate the distortion vector for each individual stitch. These vectors can facilitate the automatic reprogramming of corrected stitch positions in the CNC sewing program to minimize the final distortion. An image processing pipeline is proposed, consisting of thread detection and model-based registration, finding an assignment of thread pixels and model stitch positions. The corresponding distortion vectors are computed for every individual stitch. Tarek Stiebel, Dieter Geller, Dorit Merhof |
ETFA | 3 |
| 2018 | Gradual Domain Adaptation for Segmenting Whole Slide Images Showing Pathological Variability
Michael Gadermayr, Dennis Eschweiler, Barbara Mara Klinkhammer, Peter Boor, Dorit Merhof |
ICISP | 5 |
| 2018 | Fully Automatic Faulty Weft Thread Detection using a Camera System and Feature-based Pattern Recognition
Marcin Kopaczka, Marco Saggiomo, Moritz Güttler, Thomas Gries, Dorit Merhof |
ICPRAM | 5 |
| 2018 | Which Way Round? A Study on the Performance of Stain-Translation for Segmenting Arbitrarily Dyed Histological Images
Michael Gadermayr, Vitus Appel, Barbara Mara Klinkhammer, Peter Boor, Dorit Merhof |
MICCAI (2) | 5 |
| 2017 | Face Detection in Thermal Infrared Images: A Comparison of Algorithm- and Machine-Learning-Based Approaches
Marcin Kopaczka, Jan Nestler, Dorit Merhof |
ACIVS | 3 |
| 2017 | An unsupervised learning approach for tracking mice in an enclosed areaabstractBACKGROUND: In neuroscience research, mouse models are valuable tools to understand the genetic mechanisms that advance evidence-based discovery. In this context, large-scale studies emphasize the need for automated high-throughput systems providing a reproducible behavioral assessment of mutant mice with only a minimum level of manual intervention. Basic element of such systems is a robust tracking algorithm. However, common tracking algorithms are either limited by too specific model assumptions or have to be trained in an elaborate preprocessing step, which drastically limits their applicability for behavioral analysis. RESULTS: We present an unsupervised learning procedure that is basically built as a two-stage process to track mice in an enclosed area using shape matching and deformable segmentation models. The system is validated by comparing the tracking results with previously manually labeled landmarks in three setups with different environment, contrast and lighting conditions. Furthermore, we demonstrate that the system is able to automatically detect non-social and social behavior of interacting mice. The system demonstrates a high level of tracking accuracy and clearly outperforms the MiceProfiler, a recently proposed tracking software, which serves as benchmark for our experiments. CONCLUSIONS: The proposed method shows promising potential to automate behavioral screening of mice and other animals. Therefore, it could substantially increase the experimental throughput in behavioral assessment automation. Jakob Unger, Mike Mansour, Marcin Kopaczka, Nina Gronloh, Marc Spehr, Dorit Merhof |
BMC Bioinform. | 6 |
| 2016 | Automated enhancement and detection of stripe defects in large circular weft knitted fabricsabstractStripes are periodic defects that are difficult to detect during production even by experienced human inspectors. Therefore, we introduce an image processing method for automatically detecting stripe defects in circularly knitted fabric. We show how a barely visible defect can be optically enhanced to improve manual assessment as well as how descriptor-based image processing and machine learning can be used to allow automated stripe detection. Image enhancement is performed by applying gabor and matched filters to histogram-equalized fabric images. Subsequently, we extract image information with different descriptors (LBP, GLCM, HOG) and feed these into random forest and SVM classifiers. The full pipeline is validated by training and testing it on three sets of fabric produced with different knitting machines and parameter settings. Results show that the proposed enhancement combined with a statistics-based descriptor such as GLCM or HOG allows to train both tested classifiers with good classification rates of up to 98.9%. Marcin Kopaczka, Hanry Ham, Kristina Simonis, Raphael Kolk, Dorit Merhof |
ETFA | 5 |
| 2016 | Have we underestimated the power of image undistortion?abstractAn indispensable pre-processing step for image-based 3D reconstruction or photogrammetry tasks is undistortion, which is employed for correcting the non-linear projection of the surface points of objects onto the image plane due to lens distortion. To deal with this undesirable artifact, state-of-the-art methods focus on directly recovering the 2D coordinates of image points from distorted images, where a mapping function rather than the real undistortion model is obtained. Contrary to the standard undistortion, we explicitly consider the imaging process of a finite projective camera: the projection of a 3D point onto the image plane is determined not only by the lens, but also by the intrinsic parameters of the camera. After an appropriate transformation of the intrinsic matrix for parameter reduction, it is shown that an exact recovery of the real undistorted coordinates can be achieved for most of the commonly used undistortion functions. Interestingly, a simultaneous estimation of the two intrinsic parameters, aspect ratio a and skew factor s, is also embedded in the undistortion process, which demonstrates the underestimated power of image undistortion in a conventional way. As a benefit of considering the physical model instead of fitting a mapping function, improved undistortion performance can be achieved. Wei Li 0052, Wenting Huang, Matthias Breier, Dorit Merhof |
ICIP | 4 |
| 2016 | Color filter arrays revisited - Evaluation of Bayer pattern interpolation for industrial applicationsabstractModern industrial cameras mainly use the Bayer pattern as color filter array (CFA). However, this filtering limits the resolution of the color space. As interpolation methods cannot reconstruct the original image perfectly, they have to be optimized for a specific application. Therefore, the interpolation should match the purpose of the image processing system. Many standard algorithms are optimized for the subjective impression of a human observer which is not necessarily ideal for industrial applications. This paper introduces a new methodology for evaluating interpolation algorithms in industrial applications. For this purpose, a new dataset based on printed circuit boards as representatives for objects in industrial application is introduced. It shows many distinct features which are typical for industrial scenarios such as high contrast object edges, highly reflective materials and low variety of surface colors. Furthermore, two new error measures based on edge accuracy are presented which are tailored to many measurement tasks which employ edges. It can be shown in an evaluation of common CFA interpolation algorithms that these new error measures are better suited to identify the best interpolation algorithm for retaining edge accuracy than conventional error measures. Matthias Breier, Constantin Haas, Wei Li 0052, Dorit Merhof |
INDIN | 4 |
| 2015 | Skew correction and line extraction in binarized printed text imagesabstractSkew correction and text line extraction are essential steps for optical character recognition (OCR) applications. For this purpose, numerous approaches were developed, which conduct the analysis primarily in document images. However, they often suffer from limited detection range and application-specific parameter tuning. Inspired by the intrinsic properties of printed text, a novel subregion-based approach is proposed in this paper, which is applicable for generic printed text images and no parameter tuning is required. Guided by the spacing between text lines, the detection of a skew angle between ±90° is feasible. As verified by the experimental results, the proposed approach is robust to diverse skew directions and significantly improves the state-of-the-art OCR performance. Wei Li 0052, Matthias Breier, Dorit Merhof |
ICIP | 3 |
| 2015 | Detection of elevated regions in surface images from laser beam melting processesabstractLaser Beam Melting (LBM) is a promising Additive Manufacturing technology that allows the layer-based production of complex metallic components suitable for industrial applications. Widespread application of LBM is hindered by a lack of quality management and process control. Elevated regions in produced layers pose a major risk to process stability as collisions between the powder coating mechanism and the part may occur, which cause damages to either one or even both. We train a classifier-based detector for elevated regions in laser exposure result images. For this purpose we acquire two high resolution layer images: one after laser exposure and another one after powder deposition for the next layer. Ground truth labels for critical regions are obtained from analysis of the latter, where elevated regions are not covered by powder. We compute dense descriptors (HOG, DAISY, LBP) on the surface image after laser exposure and compare their predictive power. The top five descriptor configurations are used to optimize parameters of Random Forest, Support Vector Machine and Stochastic Gradient Descent (SGD) classifiers. We validate the detectors with optimized parameters using cross-validation on 281 images from three build jobs. Using a DAISY descriptor with a SGD classifier we achieve a F1-score of 0.670. The presented method enables detection of elevated regions before powder coating is performed and can be extended to other surface inspection tasks in LBM layer images. Detection results can be used to assess LBM process parameters with respect to process stability during process design and for quality management in production. Joschka zur Jacobsmühlen, Stefan Kleszczynski, Gerd Witt, Dorit Merhof |
IECON | 4 |
| 2015 | Discrimination of cell cycle phases in PCNA-immunolabeled cellsabstractBACKGROUND: Protein function in eukaryotic cells is often controlled in a cell cycle-dependent manner. Therefore, the correct assignment of cellular phenotypes to cell cycle phases is a crucial task in cell biology research. Nuclear proteins whose localization varies during the cell cycle are valuable and frequently used markers of cell cycle progression. Proliferating cell nuclear antigen (PCNA) is a protein which is involved in DNA replication and has cell cycle dependent properties. In this work, we present a tool to identify cell cycle phases and in particular, sub-stages of the DNA replication phase (S-phase) based on the characteristic patterns of PCNA distribution. Single time point images of PCNA-immunolabeled cells are acquired using confocal and widefield fluorescence microscopy. In order to discriminate different cell cycle phases, an optimized processing pipeline is proposed. For this purpose, we provide an in-depth analysis and selection of appropriate features for classification, an in-depth evaluation of different classification algorithms, as well as a comparative analysis of classification performance achieved with confocal versus widefield microscopy images. RESULTS: We show that the proposed processing chain is capable of automatically classifying cell cycle phases in PCNA-immunolabeled cells from single time point images, independently of the technique of image acquisition. Comparison of confocal and widefield images showed that for the proposed approach, the overall classification accuracy is slightly higher for confocal microscopy images. CONCLUSION: Overall, automated identification of cell cycle phases and in particular, sub-stages of the DNA replication phase (S-phase) based on the characteristic patterns of PCNA distribution, is feasible for both confocal and widefield images. Felix Schönenberger, Anja Deutzmann, Elisa Ferrando-May, Dorit Merhof |
BMC Bioinform. | 4 |
| 2015 | Blind weave detection for woven fabrics
Dorian Schneider, Dorit Merhof |
Pattern Anal. Appl. | 2 |
| 2015 | Interactive tracking of insect posture
Minmin Shen, Chen Li 0022, Wei Huang 0013, Paul Szyszka, Kimiaki Shirahama, Marcin Grzegorzek, Dorit Merhof, Oliver Deussen |
Pattern Recognit. | 7 |
| 2014 | Robustness analysis of imaging system for inspection of laser beam melting systemsabstractLaser Beam Melting (LBM) is an additive manufacturing process, which enables the layer-based production of complex parts from metal powder, i.e. “3D printing” with metal. In previous publications, we presented a high-resolution imaging system for inspection of LBM processes, which uses a high-resolution camera to acquire images of each powder layer and laser exposure result. The external camera position necessitates perspective correction, which is based on calibration markers which are “drawn” onto the powder by the LBM system's laser in the first layer. As movements of the powder deposition mechanism cause vibrations, the orientation of the camera may be changed, which would invalidate the calibration results and lead to imprecise measurements or segmentations. To evaluate the effect of these disturbances, we placed calibrations markers in multiple layers and determined the position offset using template matching. We analyze the relative marker drift in three LBM processes and determine the spatial acquisition error. The maximum distance is 4.91 pixels (156.1 μm on the part), while most detected markers deviate by less than 1.5 pixels (46 μm). Compared to the pixel size of 20 μm to 32μm, these deviations are significant and require a repeated calibration in higher layers for valid high-resolution image-based measurements. Joschka zur Jacobsmühlen, Stefan Kleszczynski, Gerd Witt, Dorit Merhof |
ETFA | 4 |
| 2014 | Interactive Framework for Insect Tracking with Active LearningabstractExtracting motion trajectories of insects is an important prerequisite in many behavioral studies. Despite great efforts to design efficient automatic tracking algorithms, tracking errors are unavoidable. In this paper, we propose general principles that help to minimize the human effort required for accurate multi-target tracking in the form of applications that can track the antennae and mouthparts of a honey bee based on a set of low frame rate videos. This interactive framework estimates which key frames will require user correction, i.e. those that are used for user correction, which are used for 1) incrementally learning an object classifier and 2) data association based tracking. To this framework we apply a standard classification algorithm (i.e. naive Bayesian classification) and an association optimization algorithm (i.e. Hungarian algorithm). The precision of tracking results by our framework on real-world video data is above 98%. Minmin Shen, Wei Huang 0013, Paul Szyszka, C. Giovanni Galizia, Dorit Merhof |
ICPR | 5 |
| 2014 | Text recognition for information retrieval in images of printed circuit boardsabstractIn order to achieve an efficient and environment-friendly recycling of printed circuit boards (PCBs), a comprehensive analysis of their material composition is essential. Besides sophisticated chemical and physical methods for a direct material analysis, an indirect method based on information retrieval provides a less costly and more efficient alternative. During the process of information retrieval, PCBs and their components need to be recognized based on their appearance and the corresponding text information. Their material composition is then available through a pre-established database. Therefore, a practical text recognition is necessary for a successful data analysis prior to PCB recycling. Our paper is focusing on two key aspects of text recognition: binarization and final recognition of text objects using optical character recognition (OCR) engines. For binarization of text contents, a novel local thresholding method using an adaptive window size along with background estimation is presented. Several state-of-the-art algorithms and the proposed method were evaluated for comparing their binarization performance on text objects in PCB images. With respect to a data set containing manually created references, our novel method provides superior results. Furthermore, in contrast to previous work on text recognition, an additional evaluation of available open source OCR engines was conducted to asses technical limitations of OCR applications. We show that the quality of text recognition can be significantly improved if the binarization approach accounts for these technical limitations of OCR software. The presented method and results are expected to provide improved OCR performance also in other applications. Wei Li 0052, Stefan Neullens, Matthias Breier, Marcel Bosling, Thomas Pretz, Dorit Merhof |
IECON | 6 |
| 2014 | Rotation estimation for printed circuit board recyclingabstractElectronic devices are nowadays an integral part of everyday life. The number of discarded electronic items has grown significantly over the last years. Due to the amount of precious materials used in the manufacturing of these devices recycling of electronic devices is becoming more and more important. Currently, the processes to regain some of these precious materials such as gold, copper, scarce elements etc. are not able to differentiate electronic waste according to its material composition. To enhance these processes, as much information as possible needs to be retrieved per electronic waste item. In particular, information used for the classification of the processed printed circuit boards (PCBs) is important as PCBs are extensively used in electronic devices. One key aspect of this classification process is the estimation of the orientation of the PCBs in this process (e.g. on a conveyor belt). In this paper typical properties of PCBs with respect to orientation estimation are introduced and three different orientation estimation algorithms are evaluated and discussed. These approaches comprise the Hough transform, structure tensors and orientation estimation via Fourier transform. Finally, all presented algorithms are evaluated based on images of PCBs with known orientation. The results indicate, that the Hough transform yields the best results for completely visible PCBs while structure tensors performed best on partially visible PCBs. Matthias Breier, Wei Li 0052, Marcel Bosling, Thomas Pretz, Dorit Merhof |
IECON | 5 |
| 2014 | A traverse inspection system for high precision visual on-loom fabric defect detection
Dorian Schneider, Timm Holtermann, Dorit Merhof |
Mach. Vis. Appl. | 3 |
| 2013 | Automatic framework for tracking honeybee's antennae and mouthparts from low framerate videoabstractAutomatic tracking of the movement of bee's antennae and mouthparts is necessary for studying associative learning of individuals. However, the problem of tracking them is challenging: First, the different classes of objects possess similar appearance and are close to each other. Second, tracking gaps are often present, due to the low frame-rate of the acquired video and the fast motion of the objects. Most existing insect tracking approaches have been developed for slow moving objects, and are not suitable for this application. In this paper, a novel Bayesian framework is proposed to automatically track bees' antennae and their mouthparts. This framework incorporates information about their kinematics, shape, order and temporal correlation between neighboring frames. Experimental evaluation demonstrates the effectiveness and efficiency of the proposed framework. Minmin Shen, Paul Szyszka, C. Giovanni Galizia, Dorit Merhof |
ICIP | 4 |
| 2013 | The looks of an odour - Visualising neural odour response patterns in real timeabstractBACKGROUND: Calcium imaging in insects reveals the neural response to odours, both at the receptor level on the antenna and in the antennal lobe, the first stage of olfactory information processing in the brain. Changes of intracellular calcium concentration in response to odour presentations can be observed by employing calcium-sensitive, fluorescent dyes. The response pattern across all recorded units is characteristic for the odour. METHOD: Previously, extraction of odour response patterns from calcium imaging movies was performed offline, after the experiment. We developed software to extract and to visualise odour response patterns in real time. An adaptive algorithm in combination with an implementation for the graphics processing unit enables fast processing of movie streams. Relying on correlations between pixels in the temporal domain, the calcium imaging movie can be segmented into regions that correspond to the neural units. RESULTS: We applied our software to calcium imaging data recorded from the antennal lobe of the honeybee Apis mellifera and from the antenna of the fruit fly Drosophila melanogaster. Evaluation on reference data showed results comparable to those obtained by previous offline methods while computation time was significantly lower. Demonstrating practical applicability, we employed the software in a real-time experiment, performing segmentation of glomeruli--the functional units of the honeybee antennal lobe--and visualisation of glomerular activity patterns. CONCLUSIONS: Real-time visualisation of odour response patterns expands the experimental repertoire targeted at understanding information processing in the honeybee antennal lobe. In interactive experiments, glomeruli can be selected for manipulation based on their present or past activity, or based on their anatomical position. Apart from supporting neurobiology, the software allows for utilising the insect antenna as a chemosensor, e.g. to detect or to classify odours. Martin Strauch, Clemens Müthing, Marc P. Broeg, Paul Szyszka, Daniel Münch, Thomas Laudes, Oliver Deussen, Cosmas Galizia, Dorit Merhof |
BMC Bioinform. | 9 |
| 2012 | HiTSEE KNIME: a visualization tool for hit selection and analysis in high-throughput screening experiments for the KNIME platformabstractWe present HiTSEE (High-Throughput Screening Exploration Environment), a visualization tool for the analysis of large chemical screens used to examine biochemical processes. The tool supports the investigation of structure-activity relationships (SAR analysis) and, through a flexible interaction mechanism, the navigation of large chemical spaces. Our approach is based on the projection of one or a few molecules of interest and the expansion around their neighborhood and allows for the exploration of large chemical libraries without the need to create an all encompassing overview of the whole library. We describe the requirements we collected during our collaboration with biologists and chemists, the design rationale behind the tool, and two case studies on different datasets. The described integration (HiTSEE KNIME) into the KNIME platform allows additional flexibility in adopting our approach to a wide range of different biochemical problems and enables other research groups to use HiTSEE. Hendrik Strobelt, Enrico Bertini, Joachim Braun, Oliver Deussen, Ulrich Groth, Thomas U. Mayer, Dorit Merhof |
BMC Bioinform. | 7 |
| 2007 | Correction of susceptibility artifacts in diffusion tensor data using non-linear registration
Dorit Merhof, Grzegorz Soza, Andreas Stadlbauer, Günther Greiner, Christopher Nimsky |
Medical Image Anal. | 1 |
| 2006 | Fast and Accurate Connectivity Analysis Between Functional Regions Based on DT-MRI
Dorit Merhof, Mirco Richter, Frank Enders, Peter Hastreiter, Oliver Ganslandt, Michael Buchfelder, Christopher Nimsky, Günther Greiner |
MICCAI (2) | 1 |
| 2006 | Hybrid Visualization for White Matter Tracts using Triangle Strips and Point SpritesabstractDiffusion tensor imaging is of high value in neurosurgery, providing information about the location of white matter tracts in the human brain. For their reconstruction, streamline techniques commonly referred to as fiber tracking model the underlying fiber structures and have therefore gained interest. To meet the requirements of surgical planning and to overcome the visual limitations of line representations, a new real-time visualization approach of high visual quality is introduced. For this purpose, textured triangle strips and point sprites are combined in a hybrid strategy employing GPU programming. The triangle strips follow the fiber streamlines and are textured to obtain a tube-like appearance. A vertex program is used to orient the triangle strips towards the camera. In order to avoid triangle flipping in case of fiber segments where the viewing and segment direction are parallel, a correct visual representation is achieved in these areas by chains of point sprites. As a result, a high quality visualization similar to tubes is provided allowing for interactive multimodal inspection. Overall, the presented approach is faster than existing techniques of similar visualization quality and at the same time allows for real-time rendering of dense bundles encompassing a high number of fibers, which is of high importance for diagnosis and surgical planning. Dorit Merhof, Markus Sonntag, Frank Enders, Christopher Nimsky, Peter Hastreiter, Günther Greiner |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2005 | Visualization of White Matter Tracts with Wrapped StreamlinesabstractDiffusion tensor imaging is a magnetic resonance imaging method which has gained increasing importance in neuroscience and especially in neurosurgery. It acquires diffusion properties represented by a symmetric 2nd order tensor for each voxel in the gathered dataset. From the medical point of view, the data is of special interest due lo different diffusion characteristics of varying brain tissue allowing conclusions about the underlying structures such as while matter tracts. An obvious way to visualize this data is to focus on the anisotropic areas using the major eigenvector for tractography and rendering lines for visualization of the simulation results. Our approach extends this technique to avoid line representation since lines lead 10 very complex illustrations and furthermore are mistakable. Instead, we generate surfaces wrapping bundles of lines. Thereby, a more intuitive representation of different tracts is achieved. Frank Enders, Natascha Sauber, Dorit Merhof, Peter Hastreiter, Christopher Nimsky, Marc Stamminger |
IEEE Visualization | 3 |