Gerald Schaefer

dblp:34/6798 · DBLP profile ↗
← Back
151ranked-venue papers
28as first author
41since 2021 · last 2026
0000-0003-1292-7674ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 10 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 15 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 20 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 ACID-Style: An Adaptive Condition Injection Diffusion Model for Arbitrary Style Transfer
abstract
Arbitrary style transfer (AST), a popular AI-powered photo editing function, aims to strike an optimal balance between content and style injection from two images in order to generate a novel high-fidelity stylised image. Recently, diffusion models have been applied to AST due to their high generation quality as well as flexibility to embed conditions. However, these models are still not satisfactory and may exhibit inferior performance compared to non-diffusion based methods. This is due to the diffusion process not being purposely designed for AST, leading to suboptimal solutions to trade-off content preservation and style embedding. In this paper, we propose ACID-Style, a novel adaptive condition injection diffusion-based AST framework for improved content/style feature injection to address this research challenge. Using two lightweight adapters, a content and a style injection module, and an adaptive injection mechanism, our approach is able to fully exploit a pre-trained stable diffusion model for AST-specific adaptation and our diffusion model thus learns the most effective timing for content and style injection in the diffusion sampling process. Comprehensive evaluations demonstrate that our method achieves superior style transfer performance, both quantitatively and qualitatively, compared to other state-of-the-art style transfer methods.
Ting Yang 0009, Siyu Yang 0005, Xiyao Liu 0001, Songtao Wu, Gerald Schaefer, Kuanhong Xu, Hui Fang 0003
AAAI5
2026 A gated recurrent unit-based soft actor-critic approach with social force model crowd simulation for improved mobile robot path planning
Dezhen Zhang, Guoxu Wang, Gerald Schaefer, Hui Fang 0003
Eng. Appl. Artif. Intell.3
2026 An adaptive multimodal semantic knowledge enhanced framework for sarcasm detection
Jing Dong 0009, Yu Sui, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang
Expert Syst. Appl.5
2026 Fine-grained face personalisation using a text-guided multi-attribute embedded diffusion model
Jing Dong 0009, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang
Expert Syst. Appl.5
2026 Convergence analysis of the human mental search algorithm by a Markov model
Seyed Jalaleddin Mousavirad, Hossein Ebrahimpour-Komleh, Gerald Schaefer
Knowl. Based Syst.3
2026 Refining pseudo-labels through iterative mix-up for weakly supervised semantic segmentation
abstract
Weakly supervised semantic segmentation (WSSS) aims to provide accurate pixel-level annotation based on only weak guidance, primarily derived from image-level labels. Recent WSSS methods exploit pseudo-labels generated from improved class activation maps (CAMs) to train a fine-grained classification model for semantic segmentation. However, these pseudo-labels are unreliable because they tend to either miss parts of the objects or include irrelevant regions due to weak guidance from individual images. In this paper, we propose a simple yet effective iterative mix-up strategy, Pseudo-Label-based Mix (PL-Mix), that refines pseudo-labels iteratively, thereby further enhancing WSSS performance. During each iteration, we migrate object regions from pseudo-labels produced in previous steps and render them with new contexts in a mix-up fashion. Due to model consistency enforcement across varied backgrounds and new combinations of multiple objects from enriched image samples, these pseudo-labels progressively become more accurate and reliable. Further enhanced by a masking strategy and a CAM-based earth mover’s distance loss, we achieve state-of-the-art performance on the PASCAL VOC2012 and MS COCO2014 benchmark datasets.
Yifan Wang 0008, Kunhao Yuan, Gerald Schaefer, Xiyao Liu 0001, Linglin Jing, Kehua Guo, James Z. Wang 0001, Hui Fang 0003
Pattern Recognit.3
2026 Robust and Diversified Image Steganography Without Embedding Through a Disentanglement Autoencoder
abstract
Image Steganography without Embedding (SWE) is an emerging data hiding paradigm. Instead of embedding a secret message into a container image, SWE synthesises a novel image by using the secret message as a latent code. Current SWE methods have achieved high synthesis quality and strong resistance to steganalysis tools. However, it remains challenging to apply the SWE due to two reasons: (i) lack of synthesis diversity and (ii) recovery of secret messages under malicious image attacks. In this paper, we present a novel SWE framework with a disentanglement autoencoder to tackle the above challenges. Specifically, the autoencoder disentangles an image into a structure and texture representation. Then, we exploit the stability of the structure representation to improve secret message recovery reliability, while increasing synthesis diversity by randomising texture representations and employing a chaotic system for structure randomisation to enhance its security. To further achieve a robust message recovery under malicious attacks, an adversarial learning strategy is introduced into our framework, which guarantees high recovery accuracy. Our method outperforms other state-of-the-art SWE methods in terms of synthesis quality, synthesis diversity and secret message recovery accuracy under various image attacks. The source code is publicly available athttps://github.com/Lemok00/RDI-SWE.
Xiyao Liu 0001, Ziping Ma 0002, Jian Zhang 0048, Gerald Schaefer, Kehua Guo, Yuesheng Zhu, Shichao Zhang 0001
IEEE Trans. Dependable Secur. Comput.4
2026 High-Capacity Generative Image Steganography Approach for Hiding Multiple Secret Images
abstract
Current high-capacity image steganography methods face challenges in balancing hidden capacity, imperceptibility, and recovery quality. Existing embedding-based image-in-image steganography approaches tend to produce detectable artifacts when hiding multiple images, whereas existing generative methods struggle to conceal full-sized secret images and often generate unrealistic stego images. To address these issues, this paper proposes a novel generative steganography approach that hides multiple secret images in a single realistic generated image. Our main contributions include a meticulously designed autoencoder that compresses and injects secret images into the shallow layer of the generator to increase hidden capacity, a three-stage optimization strategy for stable training to enhance the recovery quality of secret images, and an automatic image selection procedure which explores the advantage of generation diversity to enhance the imperceptibility of stego images. Experimental results demonstrate that our method outperforms embedding-based approaches by achieving higher recovered image quality with a PSNR value of 30.45 dB when concealing four images while maintaining stronger resistance against steganalysis tools, with an accuracy of 50%. Against generative approaches, our method achieves a higher hidden capacity while preserving a superior visual quality of stego images, with a FID of 6.97, surpassing the suboptimal method's FID of 22.72.
Xiyao Liu 0001, Lian Zhong, Xiangui Kang, Gerald Schaefer, Da Huang 0002, Ziping Ma 0002
IEEE Trans. Multim.4
2025 Recoverable Facial Identity Protection via Adaptive Makeup Transfer Adversarial Attacks
abstract
Unauthorised face recognition (FR) systems have posed significant threats to digital identity and privacy protection. To alleviate the risk of compromised identities, recent makeup transfer-based attack methods embed adversarial signals in order to confuse unauthorised FR systems. However, their major weakness is that they set up a fixed image unrelated to both the protected and the makeup reference images as the confusion identity, which in turn has a negative impact on both attack success rate and visual quality of transferred photos. In addition, the generated images cannot be recognised by authorised FR systems once attacks are triggered. To address these challenges, in this paper, we propose a Recoverable Makeup Transferred Generative Adversarial Network (RMT-GAN) which has the distinctive feature of improving its image-transfer quality by selecting a suitable transfer reference photo as the target identity. Moreover, our method offers a solution to recover the protected photos to their original counterparts that can be recognised by authorised systems. Experimental results demonstrate that our method provides significantly improved attack success rates while maintaining higher visual quality compared to state-of-the-art makeup transfer-based adversarial attack methods. Our code and supplementary materials are available on Github.
Xiyao Liu 0001, Junxing Ma, Xinda Wang 0006, Qianyu Lin, Jian Zhang 0048, Gerald Schaefer, Cagatay Turkay, Hui Fang 0003
AAAI6
2025 PruneClust-DE: A Novel Dual-Strategy Clustering-based Differential Evolution Algorithm for Neural Network Training
abstract
Training artificial neural networks is a fundamental step in developing machine learning models, as it determines their ability to learn and generalise from data. While gradient-based methods such as stochastic gradient descent and its variants dominate training approaches, they are susceptible to issues like sensitivity to initialisation and convergence to local optima. To address these challenges, gradient-free metaheuristic algorithms, such as differential evolution (DE), are promising alternatives due to their ability to effectively explore complex optimisation landscapes. In this paper, we propose a novel DE-based algorithm, PruneClust-DE, for training multilayer neural networks. Our approach introduces two key strategies: (1) clustering-based interpolation, which partitions the population into clusters, identifies centroids, and generates new candidate solutions by interpolating between cluster centroids to balance exploration and exploitation, and (2) fitness-based pruning, a mechanism that retains only the fittest individuals after introducing new candidates, ensuring a constant yet high-quality population. We validate our proposed algorithm across diverse datasets and compare its performance with other state-of-the-art methods, demonstrating its superiority in achieving robust results.
Seyed Jalaleddin Mousavirad, Mattias O'Nils, Gerald Schaefer, Diego Oliva 0001
SMC3
2025 C2L-DE-Lite: A Lightweight Solution to Clustering Complexity in Differential Evolution for Neural Network Training *
abstract
Determining optimal weights and biases for neural networks is a critical task. While gradient-based methods are widely used for training, they are sensitive to initialisation and susceptible to local optima. Population-based metaheuristics, such as differential evolution (DE), can offer a reliable alternative. Recently, clustering-based DE approaches have been proposed to further improve this process. However, they suffer from increased complexity, particularly with growing network sizes, leading to longer computation times. In this paper, we introduce strategies to reduce the time complexity of clustering-based DE, including clustering in the objective space, a two-tier clustering period, and one-step k-means clustering. We select one of the recent training algorithms, C2L-DE, as a representative method to incorporate our proposed strategies, leading to a lightweight version, C2L-DE-Lite. We show that C2L-DE-Lite decreases the complexity from $O\left. {\left({\sqrt {{N_{pop}}} \cdot{N_{pop}}\cdot} \right.d\cdot I}\right)$, where Npopis the population size, d is the dimensionality, and I is the number of iterations, to $O\left({\frac{{{N_{pop}}\cdot\sqrt {{N_{pop}}} }}{{CP}}}\right)$, where CP is the clustering period. This means that the complexity remains constant for increasing sizes of networks. Extensive experiments demonstrate that while significantly reducing time complexity, C2L-DE-Lite maintains similar performance levels.
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Diego Oliva 0001, Mattias O'Nils
SMC2
2025 A memory-based conditional neural process for video instance segmentation
abstract
Video instance segmentation (VIS) is an evolving research topic in computer vision that aims to simultaneously detect, segment, and track semantic objects across multiple video frames. However, existing VIS methods are typically unaware of the reliability of the training samples from insufficient and imbalanced datasets, leading to suboptimal performance. To address this challenge, we propose a memory-based conditional neural process (MemCNP) module to exploit the strengths of both memory networks and the CNP model which handles heterogeneous latent space distributions for reliable modelling with insufficient data. Our MemCNP utilises predicted uncertainty to regularise VIS predictions as well as to identify reliable samples for effective training. Notably, our MemCNP is model-agnostic and can thus be seamlessly integrated into various VIS models to improve their performance. Extensive experiments on the YouTube-VIS and OVIS datasets demonstrate the effectiveness of MemCNP regardless of the underlying model architecture. • A memory-based conditional neural process. • Reliability modelling for object detection. • Uncertainty-based dynamic training sample selection. • Contrastive instance tracking.
Kunhao Yuan, Gerald Schaefer, Yukun Lai, Xiyao Liu 0001, Hui Fang 0003
Neurocomputing2
2025 A dual-aligned knowledge self-distillation framework for visible-infrared cross-modal person re-identification
abstract
• Dual alignment knowledge self-distillation to better capture modality-invariant/specific features for VI-ReID • Temperature-modulated alignment and confidence-based selective masking to enhance model reliability. • CutSwap augmentation to improve model robustness against intra-class variations and modality discrepancies. • State-of-the-art performance on SYSU-MM01 and RegDB benchmarks. Visible-infrared person re-identification (VI-ReID) significantly enhances identity retrieval across different illumination conditions by matching visible and infrared modalities. However, existing contrastive-learning-based approaches predominantly focus on cross-modal feature alignment, thus undermining model reliability in complex scenarios. To address this challenge, we introduce a Dual Alignment Knowledge Distillation (DAKD) framework that leverages comprehensive self-distillation at both instance and class levels. Our framework incorporates a temperature-modulated alignment strategy, capturing rich modality-invariant generalities as well as modality-specific discriminative details. Additionally, we propose a confidence-based selective masking mechanism that guides the distillation towards confident and informative teacher predictions. To further enhance robustness against modality discrepancies and intra-class variations, we develop a dedicated augmentation technique, CutSwap, which exchanges image channels to simulate realistic cross-modality variations. Extensive experiments on the benchmark SYSU-MM01 and RegDB datasets demonstrate superior performance compared to other state-of-the-art methods, achieving rank-1 accuracies of 76.31% and 94.83%, respectively and validating the efficacy of DAKD in maintaining robust cross-modal alignment while preserving essential identity-specific discriminative information.
Siyuan Deng, Kunhao Yuan, Gerald Schaefer, Shihua Zhou, George Vogiatzis, Yifan Wang 0008, Hui Fang 0003
Knowl. Based Syst.3
2025 A sequential mixing fusion network for enhanced feature representations in multimodal sentiment analysis
Qiang Zhang 0008, Jing Dong 0009, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015
Knowl. Based Syst.5
2025 Class activation map guided level sets for weakly supervised semantic segmentation
Yifan Wang 0008, Gerald Schaefer, Xiyao Liu 0001, Jing Dong 0009, Linglin Jing, Xianghua Xie, Hui Fang 0003
Pattern Recognit.2
2025 Attack-Defending Contrastive Learning for Volumetric Medical Image Zero-Watermarking
abstract
Zero-watermarking is an emerging distortion-free copyright protection method for volumetric medical images. However, achieving both robustness against various malicious attacks and distinguishability between individual images remains challenging. In this article, we propose a novel attack-defending contrastive learning zero-watermarking (ADCL-ZW) scheme to tackle the above challenge using deep learning-based representations. In our approach, we design an attack-defending data enrichment mechanism to enhance the watermarking robustness by generating a large number of image samples under various watermarking attacks. Subsequently, features for both watermarking distinguishability and robustness are enhanced through application of a contrastive loss. In particular, we implement a dual-stream Siamese network architecture to effectively handle both signal attacks and geometric attacks in order to enhance the watermarking performance. Experimental results demonstrate that ADCL-ZW achieves stronger watermarking robustness and a better tradeoff between watermarking robustness and distinguishability compared with state-of-the art zero-watermarking methods. One of the highlighted metrics is that the false-negative rate of ADCL-ZW achieves 0.01 when a fixed false-positive rate is set to 1%, which is more than 13.3 times better than the benchmark methods.
Xiyao Liu 0001, Cundian Yang, Hui Fang 0003, Gerald Schaefer, Jian Zhang 0048, Yuesheng Zhu, Shichao Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2024 HPL-ESS: Hybrid Pseudo-Labeling for Unsupervised Event-based Semantic Segmentation
abstract
Event-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions, which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data, previous approaches rely on event-to-image reconstruction to obtain pseudo labels for training. However, this will inevitably introduce noise, and learning from noisy pseudo labels, especially when generated from a single source, may reinforce the errors. This drawback is also called confirmation bias in pseudo-labeling. In this paper, we propose a novel hybrid pseudo-labeling framework for unsupervised event-based semantic segmentation, HPL-ESS, to alleviate the influence of noisy pseudo labels. Specifically, we first employ a plain unsupervised domain adaptation framework as our baseline, which can generate a set of pseudo labels through self-training. Then, we incorporate offline event-to-image re-construction into the framework, and obtain another set of pseudo labels by predicting segmentation maps on the re-constructed images. A noisy label learning strategy is designed to mix the two sets of pseudo labels and enhance the quality. Moreover, we propose a soft prototypical alignment (SPA) module to further improve the consistency of target domain features. Extensive experiments show that the proposed method outperforms existing state-of-the-art methods by a large margin on benchmarks (e.g., +5.88% accuracy, +10.32% mIoU on DSEC-Semantic dataset), and even surpasses several supervised methods.
Linglin Jing, Zhigang Wang 0002, Xu Yan 0005, Dong Wang 0028, Gerald Schaefer, Hui Fang 0003, Bin Zhao 0001, Xuelong Li 0001
CVPR7
2024 Towards Compact Reversible Image Representations for Neural Style Transfer
Xiyao Liu 0001, Siyu Yang 0005, Jian Zhang 0048, Gerald Schaefer, Jiya Li, Xunli Fan, Songtao Wu, Hui Fang 0003
ECCV (66)4
2024 Multi-Strategy Adversarial Learning for Robust Face Forgery Detection Under Heterogeneous and Composite Attacks
abstract
Face forgery detection has recently progressed to address the threat from image synthesis technology, although robust face forgery detection under heterogeneous attacks remains challenging. When forgers leverage image post-processing techniques to manipulate forged photos, recent detection methods exhibit significant performance degradation. In this work, we propose a novel multi-strategy adversarial learning (MAL) method to extract salient features in order to achieve more reliable forgery detection under attacks. In particular, our MAL framework creates a large number of positive and negative sample pairs by designing a composite attack generation module with supervised contrastive training to ensure the attack robustness. In addition, we exploit two intuitive strategies, hard sample selection and region consistency, to enhance the contrastive losses for further strengthened feature reliability. Extensive experimental results demonstrate our proposed method to outperform recent state-of-the-art face forgery detection methods in terms of overall accuracy under various single and composite attacks.
Xiyao Liu 0001, Fengkai Dong, Jian Zhang 0048, Gerald Schaefer, Hui Fang 0003
ICME7
2024 Memory-facilitated Joint-space Shift Adaptation in Traffic Forecasting
abstract
Traffic forecasting, crucial for intelligent transport systems, faces significant challenges from distribution shifts due to the dynamic nature of traffic patterns. Although normalisation approaches have been proposed to address distribution shifts in other time-series forecasting tasks such as predicting electricity consumption load prediction or influenza-like illness patient number estimation, they fall short in handling the complex spatial and temporal shifts in traffic data. In this paper, we propose a novel memory-facilitated joint-space shift adaptation framework, ST-Align, to address this problem in traffic forecasting. ST-Align comprises two key components targeting the input and latent space, respectively: a memory-based data alignment module in the input space, and an end-to-end memory network structure dedicated to alignment within the latent space. This joint-space design enables our ST-Align framework to effectively capture and adapt to dynamic distribution shifts in both spatial and temporal dimensions, thus enhancing model performance. Extensive experiments on various real-world datasets and prediction backbones convincingly demonstrate the robustness and generalisability of our method.
He Haitao, Gerald Schaefer, Zhigang Ji, Yifan Wang 0008, Hui Fang 0003
IJCNN3
2024 Self-supervised memory learning for scene text image super-resolution
Kehua Guo, Xiangyuan Zhu, Gerald Schaefer, Rui Ding 0017, Hui Fang 0003
Expert Syst. Appl.3
2024 Multi-scale feature fusion for single image novel view synthesis
abstract
Single image novel view synthesis allows the generation of target images with different views from a single input image. Pixel generation methods are one of the main approaches for novel view synthesis, with previous methods typically using the input image to infer the target image in the new view. However, only features from input images in the source view might not be sufficient to generate a good target image, especially when only a single input image is available. In this paper, we fuse features from an input and a warped image to collaboratively generate pixels in the new view, with the warped image as an intermediate output generated by projecting pixels of the input image onto the target view via an estimated depth. Since the estimated depth and the generated warped image are not perfect, errors will be introduced when generating target pixels. To alleviate these and to ensure better channel information between the features from input and warped image, channel attention blocks are employed. In addition, in order to use skip connections for better novel view synthesis results, encoder features in different layers from the input image are transformed to the target view via multi-resolution depths. Here, instead of downsampling a single full-resolution depth to several lower-resolution depths, we adopt a multi-scale depth estimation network to predict multiple depths at different resolutions. Experimental results on benchmark datasets show that our method gives excellent view synthesis results and outperforms other state-of-the-art novel view synthesis methods.
Gerald Schaefer, Qinggang Meng
Neurocomputing2
2024 A Memory-augmented Conditional Neural Process model for traffic prediction
abstract
This paper presents the first neural process-based model for traffic prediction, the Memory-augmented Conditional Neural Process (MemCNP). Spatio-temporal traffic prediction involves predicting future traffic patterns based on historical traffic data and the road network structure. This problem remains a challenge due to the dynamic and heterogeneous nature of urban traffic. Existing models often struggle to capture these complexities, particularly in data-limited scenarios. To address these limitations, our model presents a novel framework for uncertainty estimation based on the conditional neural process, and further incorporates a memory network module designed to acquire a representative contextual reference, thereby improving model performance under complex data distributions. By integrating the conditional neural process and the memory network, MemCNP enables the learning of the most representative contexts through iterative updates, enhancing the model’s generalisability. This allows our model to be applicable beyond car traffic, effectively handling diverse real-world traffic scenarios, including urban non-motorised traffic such as cycling, which is essential for advancing more sustainable transportation systems. This is demonstrated by comprehensive experimental results on six benchmark datasets (PeMS04, PeMS07, PeMS08, NYCTaxi, CHIBike, and T-Drive) against existing state-of-the-art traffic prediction models, where MemCNP demonstrates superior performance. Additionally, through ablation and reliability studies, we provide a comprehensive analysis of the model’s effectiveness. • The first neural process-based model, MemCNP, for traffic prediction with limited data. • MemCNP introduces a novel framework for uncertainty estimation. • A novel memory network module acquires a representative contextual reference. • MemCNP is effective across diverse scenarios, including non-motorised traffic.
He Haitao, Kunhao Yuan, Gerald Schaefer, Zhigang Ji, Hui Fang 0003
Knowl. Based Syst.4
2024 Watermarking in Secure Federated Learning: A Verification Framework Based on Client-Side Backdooring
abstract
Federated learning (FL) allows multiple participants to collaboratively build deep learning (DL) models without directly sharing data. Consequently, the issue of copyright protection in FL becomes important since unreliable participants may gain access to the jointly trained model. Application of homomorphic encryption (HE) in a secure FL framework prevents the central server from accessing plaintext models. Thus, it is no longer feasible to embed the watermark at the central server using existing watermarking schemes. In this article, we propose a novel client-side FL watermarking scheme to tackle the copyright protection issue in secure FL with HE. To the best of our knowledge, it is the first scheme to embed the watermark to models under a secure FL environment. We design a black-box watermarking scheme based on client-side backdooring to embed a pre-designed trigger set into an FL model by a gradient-enhanced embedding method. Additionally, we propose a trigger set construction mechanism to ensure that the watermark cannot be forged. Experimental results demonstrate that our proposed scheme delivers outstanding protection performance and robustness against various watermark removal attacks and ambiguity attack.
Shuo Shao 0002, Yue Yang 0007, Xiyao Liu 0001, Ximeng Liu, Zhihua Xia, Gerald Schaefer, Hui Fang 0003
ACM Trans. Intell. Syst. Technol.7
2023 Gradient-Based Graph Attention for Scene Text Image Super-resolution
abstract
Scene text image super-resolution (STISR) in the wild has been shown to be beneficial to support improved vision-based text recognition from low-resolution imagery. An intuitive way to enhance STISR performance is to explore the well-structured and repetitive layout characteristics of text and exploit these as prior knowledge to guide model convergence. In this paper, we propose a novel gradient-based graph attention method to embed patch-wise text layout contexts into image feature representations for high-resolution text image reconstruction in an implicit and elegant manner. We introduce a non-local group-wise attention module to extract text features which are then enhanced by a cascaded channel attention module and a novel gradient-based graph attention module in order to obtain more effective representations by exploring correlations of regional and local patch-wise text layout properties. Extensive experiments on the benchmark TextZoom dataset convincingly demonstrate that our method supports excellent text recognition and outperforms the current state-of-the-art in STISR. The source code is available at https://github.com/xyzhu1/TSAN.
Xiangyuan Zhu, Kehua Guo, Hui Fang 0003, Rui Ding 0017, Zheng Wu 0004, Gerald Schaefer
AAAI6
2023 Centroid-Based Differential Evolution with Composite Trial Vector Generation Strategies for Neural Network Training
Sahar Rahmani, Seyed Jalaleddin Mousavirad, Mohammed El-Abd, Gerald Schaefer, Diego Oliva 0001
EvoApplications@EvoStar4
2023 A Novel Class Activation Map for Visual Explanations in Multi-Object Scenes
abstract
Class activation maps (CAMs) have emerged as a popular technique to improve model interpretability of deep learning-based models. While existing CAM methods are able to extract salient semantic regions to provide high-confidence pseudo-labels for downstream tasks such as semantic segmentation, they are less effective when dealing with multi-object scenes. In this paper, we design a multi-channel weight assignment scheme that learns from both positive and negative regions to yield an improved CAM model for images comprising multiple objects. We demonstrate the effectiveness of our proposed method on two new data sets, a cat-and-dog dataset and a PASCAL VOC 2012-based multi-object dataset, and show it to compare favourably with other state-of-the-art CAM methods, outperforming them in terms of both mIoU and inter-object activation ratio (IAR), a new evaluation measure proposed to evaluate CAM performance in multi-object scenes.
Yifan Wang 0008, Siyuan Deng, Kunhao Yuan, Gerald Schaefer, Xiyao Liu 0001, Hui Fang 0003
ICIP4
2023 Robust Steganography without Embedding Based on Secure Container Synthesis and Iterative Message Recovery
abstract
Synthesis-based steganography without embedding (SWE) methods transform secret messages to container images synthesised by generative networks, which eliminates distortions of container images and thus can fundamentally resist typical steganalysis tools. However, existing methods suffer from weak message recovery robustness, synthesis fidelity, and the risk of message leakage. To address these problems, we propose a novel robust steganography without embedding method in this paper. In particular, we design a secure weight modulation-based generator by introducing secure factors to hide secret messages in synthesised container images. In this manner, the synthesised results are modulated by secure factors and thus the secret messages are inaccessible when using fake factors, thus reducing the risk of message leakage. Furthermore, we design a difference predictor via the reconstruction of tampered container images together with an adversarial training strategy to iteratively update the estimation of hidden messages. This ensures robustness of recovering hidden messages, while degradation of synthesis fidelity is reduced since the generator is not included in the adversarial training. Extensive experimental results convincingly demonstrate that our proposed method is effective in avoiding message leakage and superior to other existing methods in terms of recovery robustness and synthesis fidelity.
Ziping Ma 0002, Yuesheng Zhu, Guibo Luo, Xiyao Liu 0001, Gerald Schaefer, Hui Fang 0003
IJCAI5
2023 How effective are current population-based metaheuristic algorithms for variance-based multi-level image thresholding?
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Huiyu Zhou 0001, Mahshid Helali Moghadam
Knowl. Based Syst.2
2023 A multi-strategy contrastive learning framework for weakly supervised semantic segmentation
abstract
Weakly supervised semantic segmentation (WSSS) has gained significant popularity as it relies only on weak labels such as image level annotations rather than the pixel level annotations required by supervised semantic segmentation (SSS) methods. Despite drastically reduced annotation costs, typical feature representations learned from WSSS are only representative of some salient parts of objects and less reliable compared to SSS due to the weak guidance during training. In this paper, we propose a novel Multi-Strategy Contrastive Learning (MuSCLe) framework to obtain enhanced feature representations and improve WSSS performance by exploiting similarity and dissimilarity of contrastive sample pairs at image, region, pixel and object boundary levels. Extensive experiments demonstrate the effectiveness of our method and show that MuSCLe outperforms current state-of-the-art methods on the widely used PASCAL VOC 2012 dataset.
Kunhao Yuan, Gerald Schaefer, Yukun Lai, Yifan Wang 0008, Xiyao Liu 0001, Hui Fang 0003
Pattern Recognit.2
2022 An Improved DE Algorithm to Optimise the Learning Process of a BERT-based Plagiarism Detection Model
abstract
Plagiarism detection is a challenging task, aiming to identify similar items in two documents. In this paper, we present a novel approach to automatic plagiarism detection that combines BERT (bidirectional encoder representations from transformers) word embedding, attention mechanism-based long short-term memory (LSTM) networks, and an improved differential evolution (DE) algorithm for weight initialisation. BERT is used to pretrain deep bidirectional representations in all layers, while the pre-trained BERT model can be fine-tuned with only one extra output layer without significant changes in architecture. Deep learning algorithms often use the random weighting method for initialisation, followed by gradient-based optimisation algorithms such as back-propagation for training, making them susceptible to getting trapped in local optima. To address this, population- based metaheuristic algorithms such as DE can be used. We propose an improved DE algorithm with a clustering-based mutation operator, where first a winning cluster of candidate solutions is identified and a new updating strategy is then applied to include new candidate solutions in the current population. The proposed DE algorithm is used in LSTM, attention mechanism, and feed- forward neural networks to yield the initial seeds for subsequent gradient-based optimisation. We compare our proposed model with conventional and population-based approaches on three datasets (SNLI, MSRP and SemEval2014) and demonstrate it to give superior plagiarism detection performance.
Seyed Vahid Moravvej, Seyed Jalaleddin Mousavirad, Diego Oliva 0001, Gerald Schaefer, Zahra Sobhaninia
CEC4
2022 Image Disentanglement Autoencoder for Steganography without Embedding
abstract
Conventional steganography approaches embed a secret message into a carrier for concealed communication but are prone to attack by recent advanced steganalysis tools. In this paper, we propose Image DisEntanglement Autoencoder for Steganography (IDEAS) as a novel steganography without embedding (SWE) technique. Instead of directly embedding the secret message into a carrier image, our approach hides it by transforming it into a synthesised image, and is thus fundamentally immune to typical steganalysis attacks. By disentangling an image into two representations for structure and texture, we exploit the stability of structure representation to improve secret message extraction while increasing synthesis diversity via randomising texture representations to enhance steganography security. In addition, we design an adaptive mapping mechanism to further enhance the diversity of synthesised images when ensuring different required extraction levels. Experimental results convincingly demonstrate IDEAS to achieve superior performance in terms of enhanced security, reliable secret message extraction and flexible adaptation for different extraction levels, compared to state-of-the-art SWE methods.
Xiyao Liu 0001, Ziping Ma 0002, Junxing Ma, Jian Zhang 0048, Gerald Schaefer, Hui Fang 0003
CVPR5
2022 RWS-L-SHADE: An Effective L-SHADE Algorithm Incorporation Roulette Wheel Selection Strategy for Numerical Optimisation
Seyed Jalaleddin Mousavirad, Mahshid Helali Moghadam, Mehrdad Saadatmand, Ripon K. Chakrabortty, Gerald Schaefer, Diego Oliva 0001
EvoApplications5
2022 Automatic Foot Ulcer Segmentation Using an Ensemble of Convolutional Neural Networks
abstract
Foot ulcer is a common complication of diabetes mellitus and, associated with substantial morbidity and mortality, remains a major risk factor for lower leg amputations. Extracting accurate morphological features from foot wounds is crucial for appropriate treatment. Although visual inspection by a medical professional is the common approach for diagnosis, this is subjective and error-prone, and computer-aided approaches thus provide an interesting alternative. Deep learning-based methods, and in particular convolutional neural networks (CNNs), have shown excellent performance for various tasks in medical image analysis including medical image segmentation.In this paper, we propose an ensemble approach based on two encoder-decoder-based CNN models, namely LinkNet and U-Net, to perform foot ulcer segmentation. To deal with a limited number of available training samples, we use pre-trained weights (EfficientNetB1 for the LinkNet model and EfficientNetB2 for the U-Net model) and perform further pre-training using the Medetec dataset while also applying a number of morphological-based and colour-based augmentation techniques. To boost the segmentation performance, we incorporate five-fold cross-validation, test time augmentation and result fusion.Applied on the publicly available chronic wound dataset and the MICCAI 2021 Foot Ulcer Segmentation (FUSeg) Challenge, our method achieves state-of-the-art performance with databased Dice scores of 92.07% and 88.80%, respectively, and is the top ranked method in the FUSeg challenge leaderboard. The Dockerised guidelines, inference codes and saved trained models are publicly available at https://github.com/masih4/Foot_Ulcer_Segmentation.
Amirreza Mahbod, Gerald Schaefer, Rupert Ecker, Isabella Ellinger
ICPR2
2022 An Improved Novel View Synthesis Approach Based on Feature Fusion and Channel Attention
abstract
Single image novel view synthesis allows the generation of target images with different views from a single input image. Pixel generation methods are one of the main approaches for novel view synthesis, with previous methods typically using the input images to infer the target image in the new view. However, only features from input images in the source view might not be sufficient to generate a good target image, especially when only a single input image is available. In this paper, we present a deep learning-based novel view synthesis approach that fuses features from an input and a warped image to collaboratively generate pixels in the new view. The warped image here is an intermediate output generated by projecting pixels of the input image onto the target view via an estimated depth. Since the estimated depth and the generated warped image are not perfect, errors will be introduced when generating target pixels. To alleviate these and to ensure better channel information between the features from input and warped image, channel attention blocks are employed. Experimental results on standard benchmark datasets show that our method produces excellent view synthesis results and outperforms other state-of-the-art methods.
Gerald Schaefer, Qinggang Meng
SMC2
2022 A neural refinement network for single image view synthesis
Haibin Cai, Gerald Schaefer, Qinggang Meng
Neurocomputing3
2022 Multiple-feature-based zero-watermarking for robust and discriminative copyright protection of DIBR 3D videos
abstract
Zero-watermarking is a key technique for achieving lossless and flexible copyright protection of depth image-based rendering (DIBR) videos. Existing approaches extract features of both 2D frames and depth maps via a single mechanism to protect them simultaneously. However, it is difficult for these schemes to fully satisfy the copyright protection requirements of the two components, including the remarkable discriminative capability of 3D videos and robustness against various attacks. Hence, in this paper, we propose a novel multiple-feature-based zero-watermarking scheme to protect the copyright of DIBR 3D videos. To the best of our knowledge, this is the first scheme that integrates multiple features to improve both the discriminative capability and robustness against various attacks. Specifically, dual-tree complex wavelet transform and discrete cosine transform features enhance the robustness against DIBR conversion and noise addition, respectively, while ring-partition statistical residual features ensure robustness against geometric attacks and provide sufficient discriminative capacity. In addition, we use a logistic-logistic chaotic system to encrypt these multiple features for enhanced security and design an attention-based fusion approach to offer an optimal copyright protection solution. Extensive experimental results demonstrate that our proposed scheme has stronger robustness and discriminative capacity compared to state-of-the-art zero-watermarking methods.
Xiyao Liu 0001, Yayun Zhang, Yuying Sun, Gerald Schaefer, Hui Fang 0003
Inf. Sci.7
2022 Hiding multiple images into a single image via joint compressive autoencoders
abstract
Interest in image hiding has been continually growing. Recently, deep learning-based image hiding approaches improve the hidden capacity significantly. However, the major challenges of the existing methods are that they are difficult to balance between the errors of the modified cover image and those of the recovered secret image. To solve this problem, in this paper, we develop an image hiding algorithm based on a joint compressive autoencoder framework. Further, we propose a novel strategy to enlarge the hidden capacity, i.e., hiding multi-images in one container image. Specifically, our approach provides an extremely high image hidden capacity coupled with small reconstruction errors of the secret image. More importantly, we tackle the trade-off problem of earlier approaches by mapping the image representations in the latent spaces of the joint compressive autoencoder models, leading to both high visual quality of the container image and low reconstruction error the secret image. In an extensive set of experiments, we confirm our proposed approach to outperform several state-of-the-art image hiding methods, yielding high imperceptibility and steganalysis resistance of the container images with high recovery quality of the secret images, while improving the image hidden capacity significantly (four times higher than full-image hiding capacity).
Xiyao Liu 0001, Ziping Ma 0002, Fangfang Li 0004, Gerald Schaefer, Hui Fang 0003
Pattern Recognit.6
2021 Differential Evolution-based Neural Network Training Incorporating a Centroid-based Strategy and Dynamic Opposition-based Learning
abstract
Training multi-layer neural networks (MLNNs), a challenging task, involves finding appropriate weights and biases. MLNN training is important since the performance of MLNNs is mainly dependent on these network parameters. However, conventional algorithms such as gradient-based methods, while extensively used for MLNN training, suffer from drawbacks such as a tendency to getting stuck in local optima. Population-based metaheuristic algorithms can be used to overcome these problems. In this paper, we propose a novel MLNN training algorithm, CenDE-DOBL, that is based on differential evolution (DE), a centroid-based strategy (Cen-S), and dynamic opposition-based learning (DOBL). The Cen-S approach employs the centroid of the best individuals as a member of population, while other members are updated using standard crossover and mutation operators. This improves exploitation since the new member is obtained based on the best individuals, while the employed DOBL strategy, which uses the opposite of an individual, leads to enhanced exploration. Our extensive experiments compare CenDE-DOBL to 26 conventional and population-based algorithms and confirm it to provide excellent MLNN training performance.
Seyed Jalaleddin Mousavirad, Diego Oliva 0001, Salvador Hinojosa, Gerald Schaefer
CEC4
2021 RDE-OP: A Region-Based Differential Evolution Algorithm Incorporation Opposition-Based Learning for Optimising the Learning Process of Multi-layer Neural Networks
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Iakov Korovin, Diego Oliva 0001
EvoApplications2
2021 An Enhanced Differential Evolution Algorithm Using a Novel Clustering-based Mutation Operator
abstract
Differential evolution (DE) is an effective population-based metaheuristic algorithm for solving complex optimisation problems. However, the performance of DE is sensitive to the mutation operator. In this paper, we propose a novel DE algorithm, Clu-DE, that improves the efficacy of DE using a novel clustering-based mutation operator. First, we find, using a clustering algorithm, a winner cluster in search space and select the best candidate solution in this cluster as the base vector in the mutation operator. Then, an updating scheme is introduced to include new candidate solutions in the current population. Experimental results on CEC-2017 benchmark functions with dimensionalities of 30, 50 and 100 confirm that Clu-DE yields improved performance compared to DE.
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Iakov Korovin, Mahshid Helali Moghadam, Mehrdad Saadatmand, Mahdi Pedram
SMC2
2020 Many-level Image Thresholding using a Center-Based Differential Evolution Algorithm
abstract
Image thresholding is a crucial image processing task. Most of the time, it plays a pivotal role in an image processing chain, therefore, any error in image thresholding can propagate to other steps such as edge detection, area/volume estimation, or object recognition. Multi-level image thresholding is a popular method for image segmentation, dividing an image into homogeneous regions. Conventional algorithms are timeconsuming due to utilising an exhaustive search, especially when the number of threshold levels increases. On the other hand, population-based metaheuristic algorithms have been successfully applied to this problem. In this paper, we propose a center-based differential evolution (DE) algorithm for high-dimensional multilevel image thresholding (many-level image thresholding). While DE has been shown to yield satisfactory performance for various real-world optimisation problems, in our algorithm, DE is further boosted with a center-based sampling strategy. We evaluate our algorithm on a set of benchmark images on high-dimensional search spaces and with regards to an entropy-based objective function and peak signal-to-noise ratio (PSNR). The obtained results demonstrate that the proposed algorithm can improve upon the performance of other metaheuristic image thresholding techniques.
Seyed Jalaleddin Mousavirad, Shahryar Rahnamayan, Gerald Schaefer
CEC3
2020 On Improvements of the Human Mental Search Algorithm for Global Optimisation
abstract
Population-based metaheuristic algorithms are problem-independent approaches to solve global optimisation problems. The human mental search (HMS) algorithm is a powerful population-based metaheuristic algorithm that has been shown to yield competitive performance for a variety of optimisation problems. HMS comprises three main operators, mental search, grouping, and movement. Mental search explores the neighbourhood of candidate solutions based on a Levy flight distribution to allow for simultaneous exploration and exploitation. Grouping is used to cluster the current population in order to find a promising area in search space, while during movement, candidate solutions move towards the identified promising area. In this paper, we propose an improved HMS algorithm-HMS-IS-OSK - that introduces an adaptive selection of the number of mental processes to improve the exploitation ability of HMS, and a one-step k-means algorithm for grouping to decrease the computational complexity. To evaluate the proposed algorithm, we perform a set of experiments on the CEC 2017 bench-mark functions with dimensionalities of 30, 50, and 100. The obtained results show that HMS-IS-OSK outperforms standard HMS as well as other population-based metaheuristic algorithms including covariance matrix adaptation evolution strategy (CMAES), particle swarm optimisation (PSO), artificial bee colony algorithm (ABC), whale optimisation algorithm (WOA), grey wolf optimiser (GWO), and moth-flame optimisation (MFO).
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Leila Esmaeili, Iakov Korovin
CEC2
2020 Neural Network Training Using a Biogeography-Based Learning Strategy
Seyed Jalaleddin Mousavirad, Seyed Mohammad Jafar Jalali, Sajad Ahmadian, Abbas Khosravi, Gerald Schaefer, Saeid Nahavandi
ICONIP (5)5
2020 Joint compressive autoencoders for full-image-to-image hiding
abstract
Image hiding has received significant attention due to the need of enhanced multimedia services such as multimedia security and meta-information embedding for multimedia augmentation. Recently, deep learning-based methods have been introduced that are capable of significantly increasing the hidden capacity and supporting full-size image hiding. However, these methods suffer from the necessity to balance the errors of the modified cover image and the recovered hidden image. In this paper, we propose a novel joint compressive autoencoder (J-CAE) framework to design an image hiding algorithm that achieves full-size image hidden capacity with small reconstruction errors of the hidden image. More importantly, our approach addresses the trade-off problem of previous deep learning-based methods by mapping the image representations in the latent spaces of the joint CAE models. Thus, both visual quality of the container image and recovery quality of the hidden image can be simultaneously improved. Extensive experimental results demonstrate that our proposed method outperforms several state-of-the-art deep learning-based image hiding techniques in terms of imperceptibility and recovery quality of the hidden images while maintaining full-size image hidden capacity.
Xiyao Liu 0001, Ziping Ma 0002, Xingbei Guo, Jialu Hou, Lei Wang 0017, Jian Zhang 0048, Gerald Schaefer, Hui Fang 0003
ICPR7
2020 Investigating and Exploiting Image Resolution for Transfer Learning-based Skin Lesion Classification
abstract
Skin cancer is among the most common cancer types. Dermoscopic image analysis improves the diagnostic accuracy for detection of malignant melanoma and other pigmented skin lesions when compared to unaided visual inspection. Hence, computer-based methods to support medical experts in the diagnostic procedure are of great interest. Fine-tuning pre-trained convolutional neural networks (CNNs) has been shown to work well for skin lesion classification. Pre-trained CNNs are typically trained with natural images of a fixed image size significantly smaller than captured skin lesion images and consequently dermoscopic images are downsampled for fine-tuning. However, useful medical information may be lost during this transformation. In this paper, we explore the effect of input image size on skin lesion classification performance of fine-tuned CNNs. For this, we resize dermoscopic images to different resolutions, ranging from 64 × 64 to 768 × 768 pixels and investigate the resulting classification performance of three well-established CNNs, namely DenseNet-121, ResNet-18, and ResNet-50. Our results show that using very small images (of size 64 × 64 pixels) degrades the classification performance, while images of size 128 × 128 pixels and above support good performance with larger image sizes leading to slightly improved classification. We further propose a novel fusion approach based on a three-level ensemble strategy that exploits multiple fine-tuned networks trained with dermoscopic images at various sizes. When applied on the ISIC 2017 skin lesion classification challenge, our fusion approach yields an area under the receiver operating characteristic curve of 89.2% and 96.6% for melanoma classification and seborrheic keratosis classification, respectively, outperforming state-of-the-art algorithms.
Amirreza Mahbod, Gerald Schaefer, Chunliang Wang, Rupert Ecker, Georg Dorffner, Isabella Ellinger
ICPR2
2020 An Effective Approach for Neural Network Training Based on Comprehensive Learning
abstract
Multi-layer feed-forward neural networks have been used to tackle many complex practical applications. Their performance is closely related to the success of training algorithms which adapt the weights in the network. Although conventional algorithms such as back-propagation are widely used, they suffer from drawbacks such as a tendency to get trapped in local optima. Stochastic optimisation algorithms, and in particular population-based metaheuristics, represent a useful alternative in this context. In this paper, we have proposed an effective hybrid algorithm, CLPSO-LM, which is based on particle swarm optimisation (PSO), a population-based metaheuristic algorithm, the Levenberg- Marquardt (LM) algorithm as a local search algorithm, and a comprehensive learning (CL) strategy. The CL strategy in our algorithm is responsible for improving the exploration ability of the algorithm and preventing premature convergence using neighbour candidate solutions in PSO. The best position found by comprehensive learning PSO is then used as the initial network weights for the LM algorithm. An extensive set of experiments on different classification benchmark datasets and comparison to various conventional and population-based algorithms shows CLPSO-LM to yield very competitive performance.
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Iakov Korovin
ICPR2
2020 Colour Image Denoising using Curvelets and Scale Dependent Shrinkage
abstract
With the widespread use of image processing and computer vision applications, effective denoising methods are highly sought after, prompting the development of a variety of algorithms under different assumptions on noise and signal properties. However, most of these techniques are developed to deal with grayscale images, and are typically extended to colour images by processing each RGB channel separately. In this paper, we extend the curvelet power shrinkage algorithm, introduced previously for grayscale images, to colour image denoising, by applying the proposed method in the luminance/opponent-colour YCbCr colour space to take into consideration image inter-channel dependencies. The performance of the proposed algorithm on colour images corrupted by additive white Gaussian noise is evaluated in terms of both objective and subjective measures, and the obtained results show our method to be competitive to other methods including curvelet domain hard thresholding and MSt-SVD.
Oussama Kadri, Zine-Eddine Baarir, Gerald Schaefer, Iakov Korovin
SMC3
2020 Camouflage Generative Adversarial Network: Coverless Full-image-to-image Hiding
abstract
Image hiding, one of the most important data hiding techniques, is widely used to enhance cybersecurity when transmitting multimedia data. In recent years, deep learning-based image hiding algorithms have been designed to improve the embedding capacity whilst maintaining sufficient imperceptibility to malicious eavesdroppers. These methods can hide a full-size secret image into a cover image, thus allowing full-image-to-image hiding. However, these methods suffer from a trade-off challenge to balance the possibility of detection from the container image against the recovery quality of secret image. In this paper, we propose Camouflage Generative Adversarial Network (Cam-GAN), a novel two-stage coverless full-image-to-image hiding method named, to tackle this problem. Our method offers a hiding solution through image synthesis to avoid using a modified cover image as the image hiding container and thus enhancing both image hiding imperceptibility and recovery quality of secret images. Our experimental results demonstrate that Cam-GAN outperforms state-of-the-art full-image-to-image hiding algorithms on both aspects.
Xiyao Liu 0001, Ziping Ma 0002, Xingbei Guo, Jialu Hou, Gerald Schaefer, Lei Wang 0017, Victoria Wang, Hui Fang 0003
SMC5
2020 Colour Quantisation using Human Mental Search and Local Refinement
abstract
Colour quantisation is a common image processing technique to reduce the number of distinct colours in an image which are then represented by a colour palette. Selection of appropriate entries in this palette is challenging since the quality of the quantised image is directly dictated by the palette colours. In this paper, we propose a novel colour quantisation algorithm based on the human mental search (HMS) algorithm and subsequent refinement of the colour palette using k-means. HMS is a recent population-based metaheuristic algorithm that has been shown to yield good performance on a variety of optimisation problems. In the first stage, we use HMS to find a high-quality initial colour palette. In the second stage, this palette is refined using k-means to converge towards a local optimum and thus to further improve the quality of the quantised image. We evaluate our algorithm on a set of benchmark images and compare it to several conventional and soft computing-based colour quantisation algorithms to demonstrate excellent image quality, outperforming the other methods.
Seyed Jalaleddin Mousavirad, Gerald Schaefer, M. Emre Celebi 0001, Hui Fang 0003, Xiyao Liu 0001
SMC2
2020 A Multi-Organ Nucleus Segmentation Challenge
abstract
Generalized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics.
Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi
IEEE Trans. Medical Imaging30
2019 A Benchmark of Population-Based Metaheuristic Algorithms for High-Dimensional Multi-Level Image Thresholding
abstract
Multi-level image thresholding is a popular approach for image segmentation where the image is divided into several non-overlapping regions based on the image histogram. Conventional algorithms for multi-level image thresholding are time-consuming. This is in particular so when the number of thresholds increases due to the curse of dimensionality where the search space expands exponentially as the number of parameters (thresholds) increases. One approach to address this problem is to employ population-based metaheuristic algorithms. Since various such optimisation algorithms have been presented in the literature, in this paper, we benchmark the performance of 13 population-based algorithms in the high-dimensional search spaces of the multi-level image thresholding problem. The algorithms we assess include the whale optimisation algorithm (WOA), grey wolf optimiser (GWO), cuckoo optimisation algorithm (COA), biogeography-based optimisation (BBO), teaching-learning-based optimisation (TLBO), gravitational search algorithm (GSA), imperialist competitive algorithm (ICA), cuckoo search (CS), firefly algorithm (FA), bat algorithm (BA), differential evolution (DE), particle swarm optimisation (PSO), and genetic algorithm (GA). We evaluate these on different images with regards to objective function value as well as peak signal-to-noise ratio (PSNR) and also employ a non-parametric statistical test, the Wilcoxon signed rank test, to compare the algorithms and to draw conclusions about their performance for multi-level image thresholding.
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Hossein Ebrahimpour-Komleh
CEC2
2019 Skin Lesion Classification Using Hybrid Deep Neural Networks
abstract
Skin cancer is one of the major types of cancers with an increasing incidence over the past decades. Accurately diagnosing skin lesions to discriminate between benign and malignant skin lesions is crucial to ensure appropriate patient treatment. While there are many computerised methods for skin lesion classification, convolutional neural networks (CNNs) have been shown to be superior over classical methods. In this work, we propose a fully automatic computerised method for skin lesion classification which employs optimised deep features from a number of well-established CNNs and from different abstraction levels. We use three pre-trained deep models, namely AlexNet, VGG16 and ResNet-18, as deep feature generators. The extracted features then are used to train support vector machine classifiers. In a final stage, the classifier outputs are fused to obtain a classification. Evaluated on the 150 validation images from the ISIC 2017 classification challenge, the proposed method is shown to achieve very good classification performance, yielding an area under receiver operating characteristic curve of 83.83% for melanoma classification and of 97.55% for seborrheic keratosis classification.
Amirreza Mahbod, Gerald Schaefer, Chunliang Wang, Rupert Ecker, Isabella Ellinger
ICASSP2
2019 An Effective Hybrid Approach for Optimising the Learning Process of Multi-layer Neural Networks
Seyed Jalaleddin Mousavirad, Azam Asilian Bidgoli, Hossein Ebrahimpour-Komleh, Gerald Schaefer, Iakov Korovin
ISNN (1)4
2019 A Global-Best Guided Human Mental Search Algorithm with Random Clustering Strategy
abstract
Human mental search (HMS) is a recent population-based metaheuristic inspired by the exploration manner in the bid space of online auctions. It has three main operators: (1) mental search which explores the vicinity of each candidate solution based on Levy flight, (2) grouping which is performed using a clustering algorithm to find a promising area, and (3) moving towards the promising area. HMS has shown competitive performance in solving various optimisation problems.In this paper, an improved HMS algorithm, Global-Best Human Mental Search with Random Clustering Strategy (GHMS-RCS) is proposed as a variant of HMS for global optimisation. GHMS-RCS benefits from the information of global best solutions to improve the exploitation of the HMS algorithm. Also, to reduce the time complexity and enhance exploration and exploitation, a new strategy named random clustering is introduced to improve the grouping operator in HMS. Experimental results show that GHMS-RCS outperforms standard HMS as well as other population-based algorithms including particle swarm optimisation (PSO), shuffled frog-leaping algorithm (SFLA), and biogeography-based optimisation (BBO).
Seyed Jalaleddin Mousavirad, Gerald Schaefer, Iakov Korovin
SMC2
2018 Color Quantization Using Coreset Sampling
abstract
Color quantization is an important operation with many applications in computer graphics and image processing and analysis. Clustering algorithms have been extensively applied to this problem. However, despite its popularity as a general purpose clustering algorithm, k-means has not received much attention in the colour quantization literature because of its high computational requirements and sensitivity to initialization. In this paper, we propose a novel color quantization method based on the k-means algorithm. The proposed method utilizes adaptive initialization, deterministic sub-sampling and efficient coreset construction to attain high speed and high quality quantization. Experiments on a set of benchmark images demonstrate the proposed method to be significantly faster than k-means while delivering nearly identical results.
German Valenzuela, M. Emre Celebi 0001, Gerald Schaefer
SMC3
2018 Distributed Task Rescheduling With Time Constraints for the Optimization of Total Task Allocations in a Multirobot System
abstract
This paper considers the problem of maximizing the number of task allocations in a distributed multirobot system under strict time constraints, where other optimization objectives need also be considered. It builds upon existing distributed task allocation algorithms, extending them with a novel method for maximizing the number of task assignments. The fundamental idea is that a task assignment to a robot has a high cost if its reassignment to another robot creates a feasible time slot for unallocated tasks. Multiple reassignments among networked robots may be required to create a feasible time slot and an upper limit to this number of reassignments can be adjusted according to performance requirements. A simulated rescue scenario with task deadlines and fuel limits is used to demonstrate the performance of the proposed method compared with existing methods, the consensus-based bundle algorithm and the performance impact (PI) algorithm. Starting from existing (PI-generated) solutions, results show up to a 20% increase in task allocations using the proposed method.
Joanna Turner, Qinggang Meng, Gerald Schaefer, Amanda Whitbrook, Andrea Soltoggio
IEEE Trans. Cybern.3
2016 Simple and effective pre-processing for automated melanoma discrimination based on cytological findings
abstract
In this paper, we propose a simple and effective preprocessing method for melanoma classification by considering cytological properties of melanomas, in particular the alignment of the major axis of the tumor in the same direction. We evaluate our method with a set of 1,760 dermoscopic images (329 of melanomas and 1,431 of nevi) and a simple convolutional neural network (CNN) classifier with five-fold cross validation. The proposed tumor alignment method improves the classification performance by 5.8% in terms of the area under the ROC curve (AUC). In addition, it proves to be 2.1% better in term of AUC when compared with the same configured CNN trained using images that are nine times larger. Our results also show that considering the intrinsic features of the classification target is important even when the classifier has a capability to obtain effective features automatically through its learning process.
Takuya Yoshida, M. Emre Celebi 0001, Gerald Schaefer, Hitoshi Iyatomi
IEEE BigData3
2016 Effective classification of HEp-2 cells using fuzzy texture descriptors
abstract
Indirect immunofluorescence (IIF) imaging is an important technique for detecting antinuclear antibodies in HEp-2 cells and therefore employed in the diagnosis of autoimmune diseases and other important pathological conditions involving the immune system. HEp-2 cells are often categorised into six groups (homogeneous, fine speckled, coarse speckled, nucleolar, cytoplasmic, and centromere cells), which in turn give indications on different autoimmune diseases. While traditionally this classification is performed manually thus representing a subjective and laborous task, recently there is significant interest in computer vision based approaches to automatically categorise HEp-2 cell images and thus provide an objective and fast alternative. Various algorithms have been proposed for this purpose in which texture information often plays a dominant role. In this paper, we also employ texture descriptors for automated classification of HEp-2 cells but choose descriptors that allow to take into account the fuzzy nature of the images and the noise present in them. In particular, we employ a fuzzy version of the well known local binary pattern (LBP) paradigm, coupled with support vector machine (SVM) based classification. We benchmark our approach on the ICPR 2012 HEp-2 contest benchmark dataset and show it to provide excellent classification performance, outperforming not only conventional LBP features but also all algorithms that were entered in the competition as well as the performance of a human expert.
Gerald Schaefer, Niraj P. Doshi
FUZZ-IEEE1
2016 An Innovative Approach for Attribute Reduction Using Rough Sets and Flower Pollination Optimisation
abstract
Optimal search is a major challenge for wrapper-based attribute reduction. Rough sets have been used with much success, but current hill-climbing rough set approaches to attribute reduction are insufficient for finding optimal solutions. In this paper, we propose an innovative use of an intelligent optimisation method, namely the flower search algorithm (FSA), with rough sets for attribute reduction. FSA is a relatively recent computational intelligence algorithm, which is inspired by the pollination process of flowers. For many applications, the attribute space, besides being very large, is also rough with many different local minima which makes it difficult to converge towards an optimal solution. FSA can adaptively search the attribute space for optimal attribute combinations that maximise a given fitness function, with the fitness function used in our work being rough set-based classification. Experimental results on various benchmark datasets from the UCI repository confirm our technique to perform well in comparison with competing methods.
Waleed Yamany, Eid Emary, Aboul Ella Hassanien, Gerald Schaefer, Shao Ying Zhu
KES4
2016 Historic handwritten manuscript binarisation using whale optimisation
abstract
Preserving the content of historic handwritten manuscripts is important for a variety of reasons. On the other hand, digital libraries are rapidly expanding and thus facilitate to store this information directly in digital form. For digitising text documents, a crucial step is to binarise the captured images to separate the text from the background. In this paper, we propose an effective approach for binarisation of handwritten Arabic manuscripts which employs a whale optimisation algorithm, incorporating a fuzzy c-means objective function, to obtain optimal thresholds. Experimental results confirm the effectiveness of the proposed approach compared to earlier methods.
Aboul Ella Hassanien, Mohamed Abd Elfattah, Sherihan Aboulenin, Gerald Schaefer, Shao Ying Zhu, Iakov Korovin
SMC4
2016 Classifying HEp-2 cells in immunofluorescence images using multiple kernel learning
abstract
Indirect immunofluorescence (IIF) imaging is an important technique for detecting antinuclear antibodies in HEp-2 cells and therefore employed in the diagnosis of autoimmune diseases and other important pathological conditions involving the immune system. Here, HEp-2 cells are categorised into different groups, which allow to make implications about different autoimmune diseases. Traditionally, this categorisation is performed manually by an expert and is hence both subjective and time intensive. In this paper, we present an effective method for classification of HEp-2 cells in which we first extract local binary pattern (LBP) texture features in form of multi-dimensional LBP (MD-LBP) histograms and then employ a multiple kernel learning approach to classification that integrates a multitude of support vector kernels generated by sampling the feature space. We evaluate our algorithm on the ICPR 2012 HEp-2 contest benchmark dataset, and demonstrate that our employed texture features are indeed useful for the differentiation of HEp-2 cells and that our multiple kernel learning based classification approach outperforms single kernel classification schemes. Our algorithm is shown to provide super performance compared to all techniques that were entered in the competition and to rival results obtained by a human expert.
Gerald Schaefer, Niraj P. Doshi, Iakov Korovin, Shao Ying Zhu
SMC1
2016 Credibility investigation of newsworthy tweets using a visualising Petri net model
abstract
Investigating information credibility is an important problem in online social networks such as Twitter. Since misleading information can get easily propagated in Twitter, ranking tweets according to their credibility can help to detect rumors and identify misinformation. In this paper, we propose a Petri net model to visualise tweet credibility in Twitter. We consider the uniform resource locator (URL) as an effective feature in evaluating tweet credibility since it is used to identify the source of tweets, especially for newsworthy tweets. We perform an experimental evaluation on about 1000 tweets, and show that the proposed model is effective for assigning tweets to two classes: credible and incredible tweets, which each class being further divided into two sub-classes (“credible” and “seem credible” and “doubtful” and “incredible” tweets, respectively) based on appropriate features.
Mohamed Torky, Ramadan Babers, Ragia A. Ibrahim, Aboul Ella Hassanien, Gerald Schaefer, Iakov Korovin, Shao Ying Zhu
SMC5
2015 Strategies for addressing class imbalance in ensemble classification of thermography breast cancer features
abstract
Thermography provides an interesting alternative to mammography for diagnosing breast cancer as it is a noncontact, non-invasive and passive technique that is able to detect small tumors and thus can lead to earlier diagnosis. Computer-aided diagnostic approaches based on thermography are typically split into a feature extraction stage that derives useful information from the thermogram images, and a classification stage that distinguishes between malignant and benign cases. The latter is challenging since, as is the case for many medical decision making problems, there are (many) more benign cases available for classifier training compared to malignant cases, leading to an imbalanced classification problem. In this paper, we first perform image analysis to identify features describing bilateral differences in regions of interest in the thermogram. These features then form the input for a pattern classification stage for which we present several strategies to address the existing class imbalance in the context of ensemble classifiers. In particular, we discuss an ensemble constructed of cost-sensitive decision tree classifiers, an ensemble whose base classifiers are trained on balanced subspaces, and an ensemble that is based on the combination of one-class classifiers. All three strategies are evaluated on a challenging dataset of about 150 thermograms and it is shown that they provide very good classification performance and furthermore perform favourably compared to other state-of-the-art classifier ensembles for imbalanced data.
Gerald Schaefer, Tomoharu Nakashima
CEC1
2015 Increasing allocated tasks with a time minimization algorithm for a search and rescue scenario
abstract
Rescue missions require both speed to meet strict time constraints and maximum use of resources. This study presents a Task Swap Allocation (TSA) algorithm that increases vehicle allocation with respect to the state-of-the-art consensus-based bundle algorithm and one of its extensions, while meeting time constraints. The novel idea is to enable an online reconfiguration of task allocation among distributed and networked vehicles. The proposed strategy reallocates tasks among vehicles to create feasible spaces for unallocated tasks, thereby optimizing the total number of allocated tasks. The algorithm is shown to be efficient with respect to previous methods because changes are made to a task list only once a suitable space in a schedule has been identified. Furthermore, the proposed TSA can be employed as an extension for other distributed task allocation algorithms with similar constraints to improve performance by escaping local optima and by reacting to dynamic environments.
Joanna Turner, Qinggang Meng, Gerald Schaefer
ICRA3
2015 An Evaluation of Image Enhancement Techniques for Nailfold Capillary Skeletonisation
abstract
Nailfold capillaroscopy (NC) is a routine technique used to assess the characteristics and morphology of nailfold capillaries. Observation of micro-blood vessels in the nailfold is important for diagnosing diseases that lead to morphological changes of capillaries such as scleroderma, Raynaud's phenomenon and other connective tissue diseases. In order to support a computer-aided diagnosis approach to analysing NC images, several approaches have been proposed in the literature aiming to extract capillaries. In general, such capillary skeletonisation algorithms involve an image pre-processing step, followed by binarisation and finally extraction and definition of the capillary skeletons. Since image denoising and enhancement in the pre-processing step can have a major impact on the subsequent analysis, in this paper, we evaluate the performance of five enhancement techniques for the purpose for nailfold capillary skeletonisation. In particular, we investigate the α-trimmed filter, bilateral filter, bilateral enhancer, anisotropic diffusion filter and non-local means and integrate them with three capillary extraction algorithms from the literature. We report visual and quantitative performance on a set of diverse NC images. The obtained results indicate that a relatively simple α-trimmed filter, combined with a skeletonisation algorithm incorporating a difference-of-Gaussian approach to address non-uniform lighting and an iterative rule-based skeletonisation procedure, leads to the best results when comparing the obtained skeletonisations to a manually obtained ground truth.
Niraj P. Doshi, Gerald Schaefer, Shao Ying Zhu
KES2
2015 Human Action Recognition based on Spectral Domain Features
abstract
In this paper, we propose a novel approach towards human action recognition using spectral domain feature extraction. Action representations can be considered as image templates, which can be useful for understanding various actions or gestures as well as for recognition and analysis. An action recognition scheme is developed based on extracting spectral features from the frames of a video sequence using the two-dimensional discrete Fourier transform (2D-DFT). The proposed spectral feature selection algorithm offers the advantage of very low feature dimensionality and thus lower computational cost. We show that using frequency domain features enhances the distinguishability of different actions, resulting in high within-class compactness and between-class separability of the extracted features, while certain undesirable phenomena, such as camera movement and change in camera distance, are less severe in the frequency domain. Principal component analysis is performed to further reduce the dimensionality of the feature space. Experimental results on a benchmark action recognition database confirm that our proposed method offers not only computational savings but also a high degree of accuracy.
Hafiz Imtiaz, Upal Mahbub, Gerald Schaefer, Shao Ying Zhu, Md. Atiqur Rahman Ahad
KES3
2015 CT Liver Segmentation Using Artificial Bee Colony Optimisation
abstract
The automated segmentation of the liver area is an essential phase in liver diagnosis from medical images. In this paper, we propose an artificial bee colony (ABC) optimisation algorithm that is used as a clustering technique to segment the liver in CT images. In our algorithm, ABC calculates the centroids of clusters in the image together with the region corresponding to each cluster. Using mathematical morphological operations, we then remove small and thin regions, which may represents flesh regions around the liver area, sharp edges of organs or small lesions inside the liver. The extracted regions are integrated to give an initial estimate of the liver area. In a final step, this is further enhanced using a region growing approach. In our experiments, we employed a set of 38 images, taken in pre-contrast phase, and the similarity index calculated to judge the performance of our proposed approach. This experimental evaluation confirmed our approach to afford a very good segmentation accuracy of 93.73% on the test dataset.
Abdalla Mostafa, Ahmed Fouad Ali, Mohamed Abd Elfattah, Aboul Ella Hassanien, Hesham A. Hefny, Shao Ying Zhu, Gerald Schaefer
KES7
2015 An Evaluation of LBP Texture Descriptors for the Classification of HEp-2 Cells
abstract
Indirect immunofluorescence imaging is a fundamental technique for detecting antinuclear antibodies in HEp-2 cells and consequently important for the diagnosis of autoimmune diseases and other important pathological conditions involving the immune system. HEp-2 cells can be categorised into six groups: homogeneous, fine speckled, coarse speckled, nucleolar, cytoplasmic, and Centro mere cells, which give indications on different autoimmune diseases. In the literature, various algorithms have been proposed for automatic classification of HEp-2 cells based typically on shape features, texture features and classification algorithms. Local binary pattern (LBP) features are simple yet powerful texture descriptors, which encode the neighbours of a pixels into a binary pattern. While over the years a variety of LBP algorithms have been introduced, only a few descriptors are utilised in the context of HEp-2 cell classification. In this paper, we benchmarked eight rotation invariant LBP variants and a total of 16 descriptors on the ICPR 2012 HEp-2 contest benchmark dataset. We found rotation invariant multi-dimensional LBP features to lead to the best classification performance.
Niraj P. Doshi, Gerald Schaefer, Shao Ying Zhu
SMC2
2015 A hybrid cost-sensitive ensemble for imbalanced breast thermogram classification
Bartosz Krawczyk, Gerald Schaefer, Michal Wozniak 0001
Artif. Intell. Medicine2
2015 A Hybrid Color Quantization Algorithm Incorporating a Human Visual Perception Model
abstract
Color quantization is a common image processing technique where full color images are to be displayed using a limited palette of colors. The choice of a good palette is crucial as it directly determines the quality of the resulting image. Standard quantization approaches aim to minimize the mean squared error (MSE) between the original and the quantized image, which does not correspond well to how humans perceive the image differences. In this article, we introduce a color quantization algorithm that hybridizes an optimization scheme based with an image quality metric that mimics the human visual system. Rather than minimizing the MSE, its objective is to maximize the image fidelity as evaluated by S‐CIELAB, an image quality metric that has been shown to work well for various image processing tasks. In particular, we employ a variant of simulated annealing with the objective function describing the S‐CIELAB image quality of the quantized image compared with its original. Experimental results based on a set of standard images demonstrate the superiority of our approach in terms of achieved image quality.
Gerald Schaefer, Lars Nolle
Comput. Intell.1
2015 Interactive browsing of image collections on mobile devices
Gerald Schaefer, Matthew Tallyn, Daniel Felton, William Plant, David Edmundson
Multim. Tools Appl.1
2014 Cost-sensitive texture classification
abstract
Texture recognition plays an important role in many computer vision tasks including segmentation, scene understanding and interpretation, medical imaging and object recognition. In some situations, the correct identification of particular textures is more important compared to others, for example recognition of enemy uniforms for automatic defense systems, or isolation of textures related to tumors in medical images. Such cost-sensitive texture classification is the focus of this paper, which we address by reformulating the classification problem as a cost minimisation problem. We do this by constructing a cost-sensitive classifier ensemble that is tuned using a genetic algorithm. Based on experimental results obtained on several Outex datasets with cost definitions, we show our approach to work well in comparison with canonical classification methods and the ensemble approach to lead to better performance compared to single predictors.
Gerald Schaefer, Bartosz Krawczyk, Niraj P. Doshi, Tomoharu Nakashima
IEEE Congress on Evolutionary Computation1
2014 Fusion of multi-spectral and panchromatic satellite images using principal component analysis and fuzzy logic
abstract
In this paper, we propose a fuzzy-based multi-spectral (MS) and panchromatic (PAN) image fusion approach which provides a tradeoff solution between spectral and spatial fidelity and is able to preserve more detail in terms of spectral and spatial information. First, we perform principal component analysis on the multi-spectral images and utilise the first principal component to extract matched low and high frequency coefficients. We then apply fuzzy-based image fusion rules to fuse the first principal component with the PAN image, followed by fusing the approximation coefficients. The proposed approach is tested on several satellite images and shown to provide a feasible and effective approach.
Reham Gharbia, Ali Hassan El Baz, Aboul Ella Hassanien, Gerald Schaefer, Tomoharu Nakashima, Ahmad Taher Azar
FUZZ-IEEE4
2014 Retinal blood vessel segmentation using bee colony optimisation and pattern search
abstract
Accurate segmentation of retinal blood vessels is an important task in computer aided diagnosis of retinopathy. In this paper, we propose an automated retinal blood vessel segmentation approach based on artificial bee colony optimisation in conjunction with fuzzy c-means clustering. Artificial bee colony optimisation is applied as a global search method to find cluster centers of the fuzzy c-means objective function. Vessels with small diameters appear distorted and hence cannot be correctly segmented at the first segmentation level due to confusion with nearby pixels. We employ a pattern search approach to optimisation in order to localise small vessels with a different fitness function. The proposed algorithm is tested on the publicly available DRIVE and STARE retinal image databases and confirmed to deliver performance that is comparable with state-of-the-art techniques in terms of accuracy, sensitivity and specificity.
Eid Emary, Hossam M. Zawbaa, Aboul Ella Hassanien, Gerald Schaefer, Ahmad Taher Azar
IJCNN4
2014 Retinal vessel segmentation based on possibilistic fuzzy c-means clustering optimised with cuckoo search
abstract
Automated analysis of retinal vessels is essential for the diagnosis of a wide range of eye diseases and plays an important role in automatic retinal disease screening systems. In this paper, we present an approach to automatic vessel segmentation in retinal images that utilises possibilistic fuzzy c-means (PFCM) clustering to overcome the problems of the conventional fuzzy c-means objective function. In order to obtain optimised clustering results using PFCM, a cuckoo search method is used. The cuckoo search algorithm, which is based on the brood parasitic behaviour of some cuckoo species in combination with the Levy flight behaviour of some birds and fruit flies, is applied to drive the optimisation of the fuzzy clustering. The performance of our algorithm is analysed on two benchmark databases, the DRIVE and STARE datasets, and encouraging segmentation performance is observed.
Eid Emary, Hossam M. Zawbaa, Aboul Ella Hassanien, Gerald Schaefer, Ahmad Taher Azar
IJCNN4
2014 Vicinal support vector classifier using supervised kernel-based clustering
Xulei Yang, Aize Cao, Qing Song 0001, Gerald Schaefer, Yi Su 0001
Artif. Intell. Medicine4
2014 Exploiting diversity for optimizing margin distribution in ensemble learning
Qinghua Hu, Leijun Li, Xiangqian Wu 0002, Gerald Schaefer, Daren Yu
Knowl. Based Syst.4
2013 Improved LBP texture classification using ensemble learning
abstract
Texture analysis and classification play an important role in many multimedia and computer vision applications. Local binary patterns (LBP) form a simple yet powerful texture descriptor characterising local neighbourhood properties, and consequently LBP variants are widely employed. In this paper, we demonstrate that through appropriate construction of a multiple classifier system, improved texture classification based on LBP features is possible. In particular, we employ a classifier ensemble where each classifier (a support vector machine) is trained in conjunction with a different feature selection method. The ensemble is then pruned based on a diversity measure, and the remaining models are combined using a neural fuser. Experimental results, obtained on Outex benchmark datasets and employing four LBP variants, confirm that our proposed approach leads to statistically significantly improved texture classification.
Gerald Schaefer, Bartosz Krawczyk, Niraj P. Doshi
ICME1
2013 Similarity-Based Browsing of Image Search Results
abstract
In this demo paper, we present an image browsing system that is suitable for online visualisation and browsing of search results from Google Images. Our approach is based on the Huffman tables available in the JPEG headers of Google Images thumbnails. Since these are adapted to the images, we employ them directly as image features. We then generate a visualisation of the search results by projection onto a 2-dimensional visualisation space based on principal component analysis derived from the Huffman entries. Images are dynamically placed into a grid structure and organised in a tree-like hierarchy for visual browsing. Since we utilise information only from the JPEG header, the requirements in terms of bandwidth are low, while no explicit feature calculation needs to be performed, thus allowing for interactive browsing of online image search results.
David Edmundson, Gerald Schaefer, M. Emre Celebi 0001
ISM2
2013 Sorting JPEG images at a glance
abstract
While content-based image retrieval (CBIR) has been an active research area for more than two decades, the computational overhead associated with image feature extraction is often high, making existing methods unsuitable for on-line retrieval where image features need to be extracted during the retrieval process. In this paper, we present an image retrieval algorithm for JPEG images that works in an extremely fast fashion, and is based solely on information contained in the file headers. In particular, we demonstrate that optimising the Huffman tables of JPEG files not only leads to improved compression but also allows retrieval based on the (image adapted) Huffman tables. Exploiting this leads to a retrieval method that is about 40 times faster than existing compressed domain algorithms and at least 150 times faster than common pixel-domain methods. While retrieval performance on its own does not quite match that of current techniques, the method is shown to work well as an image filter to discard a large part of a database in an efficient way. Combined with a more accurate compressed-domain retrieval algorithm it is found that retrieval time can be shortened by about 80% without sacrificing retrieval accuracy.
David Edmundson, Gerald Schaefer
MMSys2
2013 An Improved Binarisation Algorithm for Nailfold Capillary Skeleton Extraction
abstract
Nailfold capillaroscopy (NC) is a non-invasive imaging technique employed to assess the condition of blood capillaries in the nail fold, and is routinely used for the detection of scleroderma spectral disorders, Raynaud's phenomenon and other connective tissue diseases. While NC image evaluation is typically performed through manual inspection by an expert, computer-aided approaches of capillary inspection can reduce the time required for diagnosis. The aim of NC image analysis is usually to extract the skeletons of the capillaries present in the image which form the basis of further analysis for diagnosis. In this paper, we propose an improved binarisation technique for NC image analysis that addresses the challenge of non-uniform background by employing a Difference-of-Gaussian approach before thresholding, coupled with a post-processing stage to remove smaller image artefacts. Based on two previously published NC skeletonisation algorithms, we demonstrate that our technique leads to significantly better skeleton extraction on different types of NC images.
Niraj P. Doshi, Gerald Schaefer, Shao Ying Zhu
SMC2
2013 Exploring mapping-based visualisations of large remote image databases
abstract
Image database visualisations, in particular mapping-based visualisations, provide an interesting approach to accessing image repositories as they are able to overcome some of the drawbacks associated with retrieval based approaches. However, making a mapping-based approach work efficiently on large remote image databases, has yet to be explored. In this paper, we present Web-Based Images Browser (WBIB), a novel system that efficiently employs image pyramids to reduce bandwidth requirements so that users can interactively explore large remote image databases.
William Plant, Gerald Schaefer
VINCI2
2013 Visualisation of large image collections
abstract
Visual information, in particular in form of images, is becoming increasingly important, and consequently efficient and effective tools for managing these rapidly growing collections are highly sought after. Interactive image database browsing systems provide an interesting alternative to retrieval-based approaches as they let the user explore an image dataset in an intuitive fashion. Based on content-based concepts, large image collections are visualised so that visually similar images are located close in the visualisation space. Once an image collection is displayed, the user is then given the opportunity to interactively explore it through various browsing operations. In my talk, I will highlight the three main approaches to visualising image databases, namely mapping-based, clustering-based and graph-based visualisations and will then present some of the systems that we have developed in our lab for this purpose, such as the Hue Sphere Image Browser, the Honeycomb Image Browser and their recent ports to large multi-touch screen environments and mobile devices.
Gerald Schaefer
VINCI1
2013 Mean shift based gradient vector flow for image segmentation
Huiyu Zhou 0001, Xuelong Li 0001, Gerald Schaefer, M. Emre Celebi 0001, Paul Miller 0003
Comput. Vis. Image Underst.3
2012 Recompressing images to improve image retrieval performance
abstract
Virtually all images are stored in compressed form, most in (lossy) JPEG format. Compressing images however has been shown to cause a small but not negligible drop in performance for content-based image retrieval (CBIR) algorithms. In this paper, we show that it is possible to reverse this performance drop. We achieve this by what might at a first glance seem counter-intuitive, namely by compressing the images even more. In detail, what we perform is recompressing images (or rather re-quantising the DCT coefficients) to their lowest common image quality setting. We demonstrate, on a benchmark image retrieval database and using standard CBIR algorithms, that this results in improved image retrieval performance rivalling that of running the algorithms on uncompressed data.
David Edmundson, Gerald Schaefer
ICASSP2
2012 Robust texture retrieval of compressed images
abstract
Almost all images are stored in compressed form, most commonly in (lossy) JPEG format. In this paper, we show that compression leads to a drop in performance of texture retrieval algorithms, and propose a method that reverses this performance drop. We achieve this by what might at first glance seem counter-intuitive, namely by compressing the images even more. In particular, we recompress images (or rather re-quantising their DCT coefficients) to their lowest common image quality setting. We demonstrate, on a large benchmark texture retrieval database and using standard texture algorithms, that this results in improved image retrieval performance close to that obtained on uncompressed images.
David Edmundson, Gerald Schaefer, M. Emre Celebi 0001
ICIP2
2012 A comprehensive benchmark of local binary pattern algorithms for texture retrieval
Niraj P. Doshi, Gerald Schaefer
ICPR2
2012 Fast JPEG image retrieval using optimised Huffman tables
David Edmundson, Gerald Schaefer
ICPR2
2012 Effective multiple classifier systems for breast thermogram analysis
Bartosz Krawczyk, Gerald Schaefer
ICPR2
2012 Multi-dimensional local binary pattern descriptors for improved texture analysis
Gerald Schaefer, Niraj P. Doshi
ICPR1
2012 Accurate genomic signal recovery using compressed sensing
Bakhtiyar Uddin, M. Emre Celebi 0001, Hassan A. Kingravi, Gerald Schaefer
ICPR4
2012 JIRL - A C++ Library for JPEG Compressed Domain Image Retrieval
abstract
In this paper we present JIRL, an open source C++ software suite that allows to perform content-based image retrieval in the JPEG compressed domain. We provide implementations of nine retrieval algorithms representing the current state-of-the-art. For each algorithm, methods for compressed domain feature extraction as well as feature comparison are provided in an object-oriented framework. In addition, our software suite includes functionality for benchmarking retrieval algorithms in terms of retrieval performance and retrieval time. An example full image retrieval application is also provided to demonstrate how the library can be used. JIRL is made available to fellow researchers under the LGPL.
David Edmundson, Gerald Schaefer
ISM2
2012 Efficient Filtering of JPEG Images
abstract
With image databases growing rapidly, efficient methods for content-based image retrieval (CBIR) are highly sought after. In this paper, we present a very fast method for filtering JPEG compressed images to discard irrelevant pictures. We show that compressing images using individually optimised quantisation tables not only maintains high image quality and therefore allows for improved compression rates, but that the quantisation tables themselves provide a useful image descriptor for CBIR. Visual similarity between images can thus be expressed as similarity between their quantisation tables. As these are stored in the JPEG header, feature extraction and similarity computation can be performed extremely fast, and we consequently employ our method as an initial filtering step for a subsequent CBIR algorithm. We show, on a benchmark dataset of more than 30,000 images, that we can filter 80% or more of the images without a drop in retrieval performance while reducing the online retrieval time by a factor of at about 5.
David Edmundson, Gerald Schaefer
ISM2
2012 Exploiting JPEG Compression for Image Retrieval
abstract
Content-based image retrieval (CBIR) has been an active research area for many years, yet much of the research ignores the fact that most images are stored in compressed form which affects retrieval both in terms of processing speed and retrieval accruacy. In this paper, we address various aspects of JPEG compressed images in the context of image retrieval. We first analyse the effect of JPEG quantisation on image retrieval and present a robust method to address the resulting performance drop. We then compare various retrieval methods that work in the JPEG compressed domain and finally propose two new methods that are based solely on information available in the JPEG header. One of these is using optimised Huffman tables for retrieval, while the other is based on tuned quantisation tables. Both techniques are shown to give retrieval performance comparable to existing methods while being magnitudes faster.
David Edmundson, Gerald Schaefer
ISM2
2012 Mobile image browsing on a 3D globe: demo paper
abstract
With users increasingly using their mobile devices such as smartphones as digital photo albums, effective methods for managing these collections are becoming increasingly important. Standard solutions provide only limited facilities for organising, browsing and searching image collections on mobile devices, making it challenging and time-consuming to locate images of interest.
Klaus Schöffmann, Marco A. Hudelist, Manfred del Fabro, Gerald Schaefer
ICMR4
2012 A user study on image browsing on touchscreens
abstract
Default image browsing interfaces on touch-based mobile devices provide limited support for image search tasks. To facilitate fast and convenient searches we propose an alternative interface that takes advantage of 3D graphics and arranges images on a rotatable globe according to color similarity. In a user study we compare the new design to the iPad's image browser. Results collected from 24 participants show that for color-sorted image collections the globe can reduce search time by 23% without causing more errors and that it is perceived as being fun to use and preferred over the standard browsing interface by 70% of the participants.
David Ahlström, Marco A. Hudelist, Klaus Schöffmann, Gerald Schaefer
ACM Multimedia4
2012 Interactive exploration of large remote image databases
abstract
Mapping-based visualisations of image databases are well suited to users wanting to survey the overall content of a collection. Given the large amount of image data contained within such visualisations, however, this approach has yet to be applied to large image databases stored remotely. In this technical demonstration, we showcase our Web-Based Images Browser (WBIB). Our novel system makes use of image pyramids so that users can interactively explore mapping-based visualisations of large remote image databases.
William Plant, Gerald Schaefer
ACM Multimedia2
2012 Interacting with image collections: visualisation and browsing of image repositories
abstract
In this tutorial paper, we look at a variety of techniques and methods for effective and intuitive image database visualisation and browsing. While interaction with traditional image retrieval systems can lead to a confusing and frustrating user experience, image browsing systems attempt to provide the user with an intuitive interface to manage potentially large image databases.
Gerald Schaefer
ACM Multimedia1
2012 An evaluation of image enhancement techniques for capillary imaging
abstract
Nailfold capillaroscopy (NC) is a non-invasive imaging technique employed to assess the condition of blood capillaries in the nailfold, and is known to be effective particularly for early detection of scleroderma spectrum disorders and evaluation of Raynaud's phenomenon. Manual inspection of NC images can be aided by a computerised system avoiding the inherent ambiguity present in human judgment and improving the diagnosis speed. For such an automated analysis, image enhancement is typically the first step. The performance of an employed image enhancement algorithm is crucial due to its influence on subsequent algorithms. In this paper, we aim to provide a comparative evaluation of different image enhancement techniques for nailfold capillaroscopy (NC) images. In particular, we evaluate the performance of ten image enhancement/noise removal techniques for NC images as a pre-cursor to edge detection aimed at identifying capillaries. Results on a variety of NC images show that bilateral filters and enhancers, non local means and anisotropic diffusion provide the best image quality for this task.
Niraj P. Doshi, Gerald Schaefer, Arcangelo Merla
SMC2
2012 Efficient and effective online image retrieval
abstract
With visual information becoming increasingly important, efficient and effective methods for querying and retrieving this kind of information are highly sought after. In this paper, we focus on image information and querying from image collections in an online retrieval fashion. In online retrieval, image features for performing retrieval are not pre-calculated but need to be extracted during the retrieval stage. Consequently, and in particular for large image datasets, the time required for feature extraction becomes crucial so as to not exceed interactive retrieval speeds. Our aim in this work is to match the retrieval accuracy of a high performing yet rather slow image retrieval method, but perform the retrieval in only a fraction of the time. We achieve this by a combination of two carefully crafted filtering stages both of which are based on the way data is stored in JPEG compressed images. The first of these performs extremely fast image retrieval using solely information contained in the JPEG headers. The second stage employs a compressed domain retrieval method that utilises features calculated from JPEG coefficient data. The first filter discards a large part of irrelevant images in a very fast fashion. The remaining images are filtered by the second technique in order to arrive at a relatively small subset of the complete database in a timely fashion. Finally, on this subset the high performing algorithm of choice, the MPEG-7 colour structure descriptor in this paper, is applied to produce a final ranking of the images to be returned to the user. Our experimental results demonstrate that on a large dataset of over 25,000 images our approach achieves retrieval scores nearly identical to those of the high performing technique while reducing the overall retrieval time by a factor of 15.
David Edmundson, Gerald Schaefer
SMC2
2012 A comparative analysis of local binary pattern texture classification
abstract
Texture recognition is an important aspect of many computer vision applications. Local binary pattern (LBP) based texture algorithms have gained significant popularity in recent years and have been shown to be useful for a variety of tasks. While over the years a variety of LBP algorithms have been introduced in the literature, what is missing is a comprehensive evaluation of their performance. In this paper, we fill this gap and benchmark 37 texture descriptors based on 15 LBP variants for texture classification against common standard datasets of textures including those captured at different rotation angles and under different illumination conditions. Overall, LBP variance (LBPV) is found to give the best texture classification performance.
Niraj P. Doshi, Gerald Schaefer
VCIP2
2012 Proposal for VCIP 2012 tutorial on "image database visualisation and browsing"
abstract
This tutorial discuss the following: Content-based image retrieval - progress and challenges; Retrieval vs. visualisation and browsing; Image database visualisation through dimensionality reduction; Image database visualisation on graphs and networks; Time-based image database visualisation and hybrid visualisation approaches; Horizontal image database browsing; Vertical image database browsing; Immersive image database browsing; Evaluating image database browsing systems;Image database browsing-challenges and future directions.
Gerald Schaefer
VCIP1
2012 Intuitive mobile image browsing on a hexagonal lattice
abstract
Following miniaturisation of cameras and their integration into mobile devices such as smartphones combined with the intensive use of the latter, it is likely that in the near future the majority of digital images will be captured using such devices rather than using dedicated cameras. Since many users decide to keep their photos on their mobile devices, effective methods for managing these image collections are required. Common image browsers prove to be only of limited use, especially for large image sets [1].
Gerald Schaefer, Matthew Tallyn, Daniel Felton, David Edmundson, William Plant
VCIP1
2012 Flickr Retriever - Fast Retrieval of Flickr Photos
abstract
With image repositories growing at an exponential rate, efficient methods for querying such databases are highly sought after. In this demo paper, we present Flickr Retriever, a software program that performs fast retrieval of images from the photo sharing site Flickr using content-based image retrieval concepts. Our methods allows retrieval of JPEG compressed images using only information stored in the header of the files. In particular, we make use of the Huffman tables in the JPEG headers and the fact that during photo upload Flickr optimises these Huffman tables in an image adaptive way. Flickr Retriever can be downloaded from our website.
David Edmundson, Gerald Schaefer
Web Intelligence2
2012 Rough C-means and Fuzzy Rough C-means for Colour Quantisation
abstract
Colour quantisation algorithms are essential for displaying true colour images using a limited palette of distinct colours. The choice of a good colour palette is crucial as it directly determines the quality of the resulting image. Colour quantisati
Gerald Schaefer, Qinghua Hu, Huiyu Zhou 0001, James F. Peters, Aboul Ella Hassanien
Fundam. Informaticae1
2010 Robust codebook-based video background subtraction
abstract
Dynamic backgrounds and sudden illumination changes are two of the major problems associated with background subtraction techniques. In this paper, we present a novel approach to background subtraction that addresses both of these challenges. In particular, we present an improved codebook background modelling and subtraction technique. We utilise image segmentation on the background image and model the background with a codebook for each pixel along with a pseudo background layer. We perceive background motion as an occlusion of one background layer by a nearby background layer. In other words, sliding of one background layer over a neighbouring layer causes background motion and will hence result in false segmentation. We present our approach of codeword spreading across layer boundaries to handle background motion. Furthermore, we present a two-step update of the background codebook to handle both sudden and gradual illumination changes.
Amit Pal, Gerald Schaefer, M. Emre Celebi 0001
ICASSP2
2010 Robust border detection in dermoscopy images using threshold fusion
abstract
Dermoscopy is one of the major imaging modalities used in the diagnosis of melanoma and other pigmented skin lesions. Due to the difficulty and subjectivity of human interpretation, automated analysis of dermoscopy images has become an important research area. Border detection is often the first step in this analysis. In many cases, the lesion can be roughly separated from the background skin using a thresholding method applied to the blue channel. However, no single thresholding method appears to be robust enough to successfully handle a wide variety of dermoscopic images. In this paper, we present an automated method for detecting lesion borders in dermoscopy images using a fusion of several thresholding methods. Experiments on a difficult set of 90 images demonstrate that the proposed method achieves both fast and accurate results when compared to six state-of-the-art methods.
M. Emre Celebi 0001, Sae Hwang, Hitoshi Iyatomi, Gerald Schaefer
ICIP4
2010 A new family of order-statistics based switching vector filters
abstract
In this paper, we present a family of order-statistics based vector filters for the removal of impulsive noise from color images. These filters preserve the edges and fine image details by switching between the identity (no filtering) operation and a robust order-statistics based filter operation based on the univariate median operator. Experiments on a diverse set of images and comparisons with state-of-the-art filters show that the proposed filters combine simplicity, flexibility, good filtering quality, and low computational requirements.
M. Emre Celebi 0001, Gerald Schaefer, Huiyu Zhou 0001
ICIP2
2010 Automated color normalization for dermoscopy images
abstract
Accurate color information in dermoscopy images is very important for melanoma diagnosis since inappropriate white balance or brightness in the images adversely affects the diagnostic performance. In this paper, we present an automated color normalization method for dermoscopy images of skin lesions. We develop color normalization filters based on a total of 319 images which normalize color of images using the HSV color system. We determined that the color characteristics of the peripheral part of the tumors significantly influence the color normalization and confirmed that the developed normalization filter achieved satisfactory normalization performance as evaluated by a cross-validation test.
Hitoshi Iyatomi, M. Emre Celebi 0001, Gerald Schaefer, Masaru Tanaka
ICIP3
2010 Image retrieval on the Honeycomb Image Browser
abstract
Efficient and effective approaches of dealing with the vast amount of visual information available nowadays are highly sought after. This is particularly the case for image collections, both personal and commercial. Due to the magnitude of these ever expanding image repositories, annotation of all images images is infeasible, and search in such an image collection therefore becomes inherently difficult. Although content-based image retrieval techniques have shown much potential, such approaches also suffer from various problems making it difficult to adopt them in practice. In this paper, we follow a different approach, namely that of browsing image databases for image retrieval. In our Honeycomb Image Browser, large image databases are visualised on a hexagonal lattice with image thumbnails occupying hexagons. Arranged in a space filling manner, visually similar images are located close together enabling large image datasets to be navigated in a hierarchical manner. Various browsing tools are incorporated to allow for interactive exploration of the database. Experimental results confirm that our approach affords efficient image retrieval.
William Plant, Gerald Schaefer
ICIP2
2010 An uncompressed benchmark image dataset for colour imaging
abstract
Virtually all vision and imaging applications and algorithms require validation and evaluation, yet benchmarking and evaluation have proven to be difficult and challenging. This is partly due to the amount of work involved in performing an appropriate and convincing test, but is also often hindered by the fact that there are rather few test datasets that are publicly available. In this paper we present UCID10K, an image database comprising 10,000 colour images. Importantly, all images in the dataset are captured and preserved in uncompressed form allowing its use for benchmarking a variety of imaging applications including image compression, colour quantisation, steganography and image forensics. In addition, a ground truth for the evaluation of content-based image retrieval algorithms is provided enabling the UCID dataset to be used for assessing the performance of image querying techniques but also the testing of compressed-domain retrieval techniques. The UCID10K database is publicly available to fellow researchers.
Gerald Schaefer
ICIP1
2010 Visual appearance based document image classification
abstract
In the paper, we present a new method for classifying documents with rigid geometry. Our approach is based on the fast and robust Viola-Jones object detection algorithm. The advantages of our proposed method are high speed, the possibility of automatic model construction using a training set, and processing of raw source images without any pre-processing steps such as draft recognition, layout analysis or binarisation. Furthermore, our algorithm allows not only to classify documents, but also to detect the placement and orientation of documents within an image.
Sergey A. Usilin, Dmitry P. Nikolaev, Vasiliy V. Postnikov, Gerald Schaefer
ICIP4
2010 Robust estimation of the fundamental matrix
abstract
Most approaches to estimate the fundamental matrix assume a Gaussian distribution in the errors in view of mathematical tractability. However, this assumption is violated if the distribution computed is not normal. In this paper we propose a robust approach of estimating the fundamental matrix which does not rely on the Gaussian assumption. The proposed technique, weighted least squares (WLS), is the application of linear mixed-effects models considering the correlation between different data sub-samples. It provides an unbiased estimation of the fundamental matrix which is not affected by outlier samples. Experimental results on synthetic and real images confirm the accuracy of our method and its superiority to standard estimation methods.
Huiyu Zhou 0001, Gerald Schaefer
ICIP2
2010 A next generation browsing environment for large image repositories
Gerald Schaefer
Multim. Tools Appl.1
2010 Segmentation of optic disc in retinal images using an improved gradient vector flow algorithm
Huiyu Zhou 0001, Gerald Schaefer, Tangwei Liu, Faquan Lin
Multim. Tools Appl.2
2010 Data mining of gene expression data by fuzzy and hybrid fuzzy methods
abstract
Microarray studies and gene expression analysis have received tremendous attention over the last few years and provide many promising avenues toward the understanding of fundamental questions in biology and medicine. Data mining of these vasts amount of data is crucial in gaining this understanding. In this paper, we present a fuzzy rule-based classification system that allows for effective analysis of gene expression data. The applied classifier consists of a set of fuzzy if-then rules that enable accurate nonlinear classification of input patterns. We further present a hybrid fuzzy classification scheme in which a small number of fuzzy if-then rules are selected through means of a genetic algorithm, leading to a compact classifier for gene expression analysis. Extensive experimental results on various well-known gene expression datasets confirm the efficacy of our approaches.
Gerald Schaefer, Tomoharu Nakashima
IEEE Trans. Inf. Technol. Biomed.1
2009 Cellular neural network based algorithms in the early detection of hand osteoarthritis
abstract
Cellular neural network (CNN) algorithms have been successfully used in a plethora of image processing applications including the medical imaging domain. Analogic CNN algorithms use CNN templates combined with logic operations to perform operations such as blurring and thresholding for image processing. In this paper we apply CNN based techniques incorporating image enhancement, region segmentation and line detection for detecting the manifestations of osteoarthritis, a metabolic disease afflicting the elderly population caused by wear and tear of cartilage surrounding weight bearing bone joints like the human hand. The two main indicators of osteoarthritis that we examine are the cystic regions, and osteophytes or bony spurs in the vicinity of the joints, produced by the rubbing together of bones due to joint space narrowing.
Sreeparna Banerjee, Gerald Schaefer, Ioannis K. Vlachos
FUZZ-IEEE2
2009 Application of cost-sensitive fuzzy classifiers to image understanding problems
abstract
Image understanding applications often involve a pattern classification stage. In this paper we show how a fuzzy rule-based classifier, extended to incorporate a cost function, can be successfully used in various imaging applications. The antecedent part of fuzzy if-then rules are specified by partitioning each attributes into fuzzy sets while the consequent class and the degree of certainty are determined from compatibility training patterns. Extension to include a cost term is shown to be straightforward and experimental results on several image processing tasks demonstrate the efficacy of our method.
Gerald Schaefer, Takashima Nakashima
FUZZ-IEEE1
2009 Contrast enhancement in dermoscopy images by maximizing a histogram bimodality measure
abstract
Dermoscopy is one of the major imaging modalities used in the diagnosis of melanoma and other pigmented skin lesions. Due to the difficulty and subjectivity of human interpretation, automated analysis of dermoscopy images has become an important research area. Border detection is typically the first step in this analysis yet is often limited by the quality of the images to be analyzed. In this paper, we present an effective method to enhance the contrast in dermoscopy images. Given an input RGB image, we determine the optimal weights to convert it to grayscale by maximizing a histogram bimodality measure. Experiments on a large set of images demonstrate that this adaptive optimization scheme increases the contrast between the lesion and the background skin, and leads to a more accurate separation of the two regions using Otsu's thresholding method.
M. Emre Celebi 0001, Hitoshi Iyatomi, Gerald Schaefer
ICIP3
2009 Skin lesion extraction in dermoscopic images based on colour enhancement and iterative segmentation
abstract
Accurate extraction of lesion borders is a crucial step in analysing dermoscopic skin lesion images. In this paper we present an effective approach to extracting lesion areas by combining an iterative segmentation algorithm with a preprocessing step that enhances colour information and image contrast. Following the pre-processing, analysis of the image background is conducted by iterative measurements based on median and standard deviation of non-lesion pixels, which in turn facilitates automatic and recurring noise reduction and enhancement. The algorithm does not depend on the use of rigid threshold values as an optimal thresholding algorithm is used to determine the optimal threshold iteratively. Extensive experimental evaluation is carried out on a dataset of 90 dermoscopy images with known ground truths obtained from three expert dermatologists. The results show that our approach is capable of providing good segmentation performance and that the colour enhancement step is indeed crucial as demonstrated by comparison with results obtained from the original RGB images.
Gerald Schaefer, Maher I. Rajab, M. Emre Celebi 0001, Hitoshi Iyatomi
ICIP1
2009 Bayesian image segmentation with mean shift
abstract
Image segmentation plays a key role in many image content analysis applications, and a lot of effort has aimed at improving the performance of established segmentation algorithms. In this paper, we present a mean shift-based combined Dirichlet process mixture (MDP)/Markov Random Field (MRF) image segmentation algorithm. Our method incorporates a mean shift process to iteratively reduce the difference between the mean of cluster centres and image pixels within the standard MDP/MRF procedure. Experimental results show that the proposed segmentation technique outperforms the classical MDP/MRF algorithm.
Huiyu Zhou 0001, Gerald Schaefer, M. Emre Celebi 0001, Minrui Fei
ICIP2
2009 3-D structure recovery from 2-D observations
abstract
In this paper we present a novel method for simultaneously determining three dimensional motion and structure of a non-rigid object from its uncalibrated two dimensional data with Gaussian or non-Gaussian distributions. A non-rigid motion can be treated as a combination of a rigid component and a non-rigid deformation. To reduce the high dimensionality of the deformable structure or shape, we estimate the probability distribution function of the structure through random sampling, integrating an established probabilistic model. The fitting between the observations and the estimated 3-D structure is evaluated using the pooled variance estimator. Applications of the proposed method to both synthetic and real image sequences show promising results.
Huiyu Zhou 0001, Gerald Schaefer, Tangwei Liu, Faquan Lin
ICIP2
2009 Visualising image databases
abstract
In this paper we explore different ways in which large collections of images can be visualised. We discuss the three principle visualisation techniques employed for this purpose, namely dimensionality reduced mappings, clustering-based visualisations and graph-based representations. Mapping-based techniques try to present the relationships between images described by high-dimensional features in a low-dimensional visualisation space. Clustered visualisations group similar images based on content, metadata or time stamp information, while in graph-based approaches links between images are exploited to arrive at an intuitive display of the dataset. We highlight advantages and disadvantages of the various approaches and emphasise the need for a benchmark which allows objective evaluation of these systems.
William Plant, Gerald Schaefer
MMSP2
2009 Localization of Lesions in Dermoscopy Images Using Ensembles of Thresholding Methods
M. Emre Celebi 0001, Hitoshi Iyatomi, Gerald Schaefer, William V. Stoecker
PSIVT3
2009 Thermography based breast cancer analysis using statistical features and fuzzy classification
Gerald Schaefer, Michal Zavisek, Tomoharu Nakashima
Pattern Recognit.1
2009 Rough Sets and Near Sets in Medical Imaging: A Review
abstract
This paper presents a review of the current literature on rough-set- and near-set-based approaches to solving various problems in medical imaging such as medical image segmentation, object extraction, and image classification. Rough set frameworks hybridized with other computational intelligence technologies that include neural networks, particle swarm optimization, support vector machines, and fuzzy sets are also presented. In addition, a brief introduction to near sets and near images with an application to MRI images is given. Near sets offer a generalization of traditional rough set theory and a promising approach to solving the medical image correspondence problem as well as an approach to classifying perceptual objects by means of features in solving medical imaging problems. Other generalizations of rough sets such as neighborhood systems, shadowed sets, and tolerance spaces are also briefly considered in solving a variety of medical imaging problems. Challenges to be addressed and future directions of research are identified and an extensive bibliography is also included.
Aboul Ella Hassanien, Ajith Abraham, James F. Peters, Gerald Schaefer, Christopher J. Henry
IEEE Trans. Inf. Technol. Biomed.4
2008 An overview of Genetic Algorithms in simulation soccer
abstract
This paper discusses the use of genetic algorithms and genetic programming within the simulation soccer domain. Genetic algorithms (GAs) are based on the Darwinian theory of evolution and provide techniques to execute an effective search on a large range of potential solutions to a specific problem. Genetic Programming (GP) uses GA concepts to evolve a computer program. We show how GAs and GP have been applied to the challenging real-time and noisy domain of RoboCup simulation soccer. Among others, genetic approaches can be used to find appropriate actions for a soccer agent during a game, to improve different aspects of team strategy as well as to strengthen the ability of a player or a team in training exercises.
William Plant, Gerald Schaefer, Tomoharu Nakashima
IEEE Congress on Evolutionary Computation2
2008 Comparison of simulated annealing and SASS for parameter estimation of biochemical networks
abstract
Estimating the parameters of biochemical networks from time-courses is becoming increasingly important. There have been some attempts in the past to carry out this task in an automatic way. In this research, SASS, a novel heuristic optimisation algorithm that has only one control parameter, has been used to solve this problem. While the obtained estimations are similar to those using other recent techniques, the method presented here offers a better resistance to local minima and a decrease of a 20% in average in computational cost, without the need of finding a suitable set of control parameters.
Josep Sayol, Lars Nolle, Gerald Schaefer, Tomoharu Nakashima
IEEE Congress on Evolutionary Computation3
2008 An overview of rough-hybrid approaches in image processing
abstract
Rough set theory offers a novel approach to manage uncertainty that has been used for the discovery of data dependencies, importance of features, patterns in sample data, feature space dimensionality reduction, and the classification of objects. Consequently, rough sets have been successfully employed for various image processing tasks including image segmentation, enhancement and classification. Nevertheless, while rough sets on their own provide a powerful technique, it is often the combination with other computational intelligence techniques that results in a truly effective approach. In this paper we show how rough sets have been combined with various other methodologies such as neural networks, wavelets, mathematical morphology, fuzzy sets, genetic algorithms, Bayesian approaches, swarm optimization, and support vector machines in the image processing domain.
Aboul Ella Hassanien, Ajith Abraham, James F. Peters, Gerald Schaefer
FUZZ-IEEE4
2008 Intensity-based image registration using multiple distributed agents
Roger J. Tait, Gerald Schaefer, Adrian A. Hopgood
Knowl. Based Syst.2
2007 Cost-Sensitive Fuzzy Classification for Medical Diagnosis
abstract
Medical diagnosis essentially represents a pattern classification problem: based on a certain input an expert arrives at a diagnosis which often takes on a binary form, i.e. the patient suffering from a certain disease or not. A lot of research has focussed on computer assisted diagnosis where objective measurements are passed to a classifier algorithm which then proposes diagnostic output based on a previous learning process. However, these classifiers put equal emphasis on a learning patterns irrespective of the class they belong to. In this paper we apply a fuzzy rule-based classification system to medical diagnosis. Importantly, we extend the classifier to incorporate a concept of cost which can be used to emphasize those cases that signify illness as it is usually more costly to incorrectly diagnose such a patient as being healthy. Experimental results on various medical datasets confirm the usefulness and efficacy of our approach
Gerald Schaefer, Tomoharu Nakashima, Yasuyuki Yokota, Hisao Ishibuchi
CIBCB1
2007 Introducing Class-Based Classification Priority in Fuzzy Rule-Based Classification Systems
abstract
In this paper we propose a fuzzy rule-generation method for pattern classification problems with classification priority. The assumption in this paper is that a classification priority is given a priori in relation to other classes. Our fuzzy rule-based classification system consists of a set of fuzzy if-then rules that are automatically generated from a set of given training patterns. The proposed method decides the consequent class of fuzzy if-then rules based on the number of covered training patterns for each class. In computational experiments we first show the effect of introducing classification priority for a synthetic two-dimensional problem. Then we show the effectiveness of the proposed method for several real-world pattern classification problems.
Tomoharu Nakashima, Yasuyuki Yokota, Gerald Schaefer, Hisao Ishibuchi
FUZZ-IEEE3
2007 Fuzzy Classification of Gene Expression Data
abstract
Microarray expression studies measure, through a hybridisation process, the levels of genes expressed in biological samples. Knowledge gained from these studies is deemed increasingly important due to its potential of contributing to the understanding of fundamental questions in biology and clinical medicine. One important aspect of microarray expression analysis is the classification of the recorded samples which poses many challenges due to the vast number of recorded expression levels compared to the relatively small numbers of analysed samples. In this paper we show how fuzzy rule-based classification can be applied successfully to analyse gene expression data. The generated classifier consists of an ensemble of fuzzy if-then rules which together provide a reliable and accurate classification of the underlying data. Experimental results on several standard microarray datasets confirm the efficacy of the approach.
Gerald Schaefer, Tomoharu Nakashima, Yasuyuki Yokota, Hisao Ishibuchi
FUZZ-IEEE1
2007 Breast Cancer Classification Using Statistical Features and Fuzzy Classification of Thermograms
abstract
Advances in camera technologies and reduced equipment costs have lead to an increased interest in the application of thermography in the medical fields. Thermography is of particular interest for detection of breast cancer as it has been shown that it is capable of detecting the cancer earlier and is also allows diagnosis of fatty breast tissue. In this paper we perform breast cancer detection based on thermography, using a series of statistical features extracted from the thermograms coupled with a fuzzy rule-based classification system for diagnosis. The features stem from a comparison of left and right breast areas and quantify the bilateral differences encountered. Following this asymmetry analysis the features are fed to a fuzzy classification system. This classifier is used to extract fuzzy if-then rules based on a training set of known cases. Experimental results on a set of nearly 150 cases show the proposed system to work well accurately classifying about 80% of cases, a performance that is comparable to other imaging modalities such as mammography.
Gerald Schaefer, Tomoharu Nakashima, Michal Zavisek, Yasuyuki Yokota, Ales Drastich, Hisao Ishibuchi
FUZZ-IEEE1
2007 An Integrative Semantic Framework for Image Annotation and Retrieval
abstract
Most public image retrieval engines utilise free-text search mechanisms, which often return inaccurate matches as they in principle rely on statistical analysis of query keyword recurrence in the image annotation or surrounding text. In this paper we present a semantically-enabled image annotation and retrieval engine that relies on methodically structured ontologies for image annotation, thus allowing for more intelligent reasoning about the image content and subsequently obtaining a more accurate set of results and a richer set of alternatives matchmaking the original query. Our semantic retrieval technology is designed to satisfy the requirements of the commercial image collections market in terms of both accuracy and efficiency of the retrieval process. We also present our efforts in further improving the recall of our retrieval technology by deploying an efficient query expansion technique.
Taha Osman, Dhavalkumar Thakker, Gerald Schaefer, Phil Lakin
Web Intelligence3
2007 A weighted fuzzy classifier and its application to image processing tasks
Tomoharu Nakashima, Gerald Schaefer, Yasuyuki Yokota, Hisao Ishibuchi
Fuzzy Sets Syst.2
2006 A Cost-based Fuzzy Rule-based System for Pattern Classification Problems
abstract
This paper proposes a cost-based fuzzy classification system for pattern classification problems with misclassification costs. The task is to minimize the total misclassification cost incurred by a fuzzy classification system consisting of a number of fuzzy if-then rules where the number of generated fuzzy if-then rules depends on the specification of fuzzy partitions for each axis. In the proposed fuzzy classification system the consequent class of a fuzzy if-then rule is determined so that the misclassification cost is minimal over the covered training patterns by the antecedent part of the rule. On the other hand, conventional fuzzy classification systems are compatibility-based. That is, the consequent class of a fuzzy if-then rule is determined from the compatibility of training patterns covered by the antecedent part of the rule. The grade of certainty of the fuzzy if-then rules in both classification systems is calculated by using the compatibility of training patterns from each class. In a series of computational experiments, we compare the performance of the proposed cost-based fuzzy classification systems with that of the conventional compatibility-based systems. The performance of both classifiers is measured for three real-world pattern classification problems.
Tomoharu Nakashima, Yasuyuki Yokota, Gerald Schaefer, Hisao Ishibuchi
FUZZ-IEEE3
2006 Quality Metric Based Colour Palette Optimisation
abstract
Colour quantisation is a common image processing technique where full colour images are to be displayed using a limited palette. The choice of a good palette is therefore crucial as it directly determines the quality of the resulting image. Standard quantisation approaches typically try to minimise the (squared) error between the original and the quantised image which does not correspond well to how humans perceive the images. In this paper we introduce a new colour quantisation algorithm that is designed not to minimise these errors but to maximise the image quality as evaluated by S-CIELAB, an image quality metric that has been shown to work well for various image processing tasks. Experimental results based on a set of standard images demonstrate the superiority in terms of achieved image quality of our novel method.
Gerald Schaefer, Lars Nolle
ICIP1
2006 A Self-Adaptive Hybrid Genetic Algorithm for Color Clustering
abstract
Color palettes are inherent to color quantized images and represent the range of possible colors in such images. When converting full true color images to palletized counterparts, the color palette should be chosen so as to minimize the resulting distortion compared to the original. In this paper, we show that in contrast to previous approaches on color quantization, which rely on either heuristics or clustering techniques, a generic optimization algorithm such as a self-adaptive hybrid genetic algorithm can be employed to generate a palette of high quality. Experiments on a set of standard test images using a novel self-adaptive hybrid genetic algorithm show that this approach is capable of outperforming several conventional color quantization algorithms and provide superior image quality.
Tarek A. El-Mihoub, Lars Nolle, Gerald Schaefer, Tomoharu Nakashima, Adrian A. Hopgood
SMC3
2006 Examining the Effect of Cost Assignment on the Performance of Cost-Based Classification Systems
abstract
This paper examines the performance of cost-based fuzzy classification systems for pattern classification problems with misclassification costs. The task here is to minimize the total sum of misclassification costs by fuzzy classification systems. In the cost-based fuzzy classification systems, the consequent class of a fuzzy if-then rule is determined so that the misclassification cost is minimized over the covered training patterns by the antecedent part of the fuzzy if-then rule. On the other hand, the conventional fuzzy classification systems are compatibility-based. That is, the consequent class of a fuzzy if-then rule is determined from the compatibility of training patterns covered by the antecedent part of the fuzzy if-then rule. The grade of certainty of the fuzzy if-then rules in both fuzzy classification systems is calculated by using the compatibility of training patterns from each class. In computational experiments, we compare the performance of the cost-based fuzzy classification systems with that of the conventional compatibility-based fuzzy classification systems.
Tomoharu Nakashima, Yasuyuki Yokota, Gerald Schaefer, Hisao Ishibuchi
SMC3
2005 A Combined Physical and Statistical Approach to Colour Constancy
abstract
Computational colour constancy tries to recover the colour of the scene illuminant of an image. Colour constancy algorithms can, in general, be divided into two groups: statistics-based approaches that exploit statistical knowledge of common lights and surfaces, and physics-based algorithms which are based on an understanding of how physical processes such as highlights manifest themselves in images. A combined physical and statistical colour constancy algorithm that integrates the advantages of the statistics-based colour by correlation method with those of a physics-based technique based on the dichromatic reflectance model is introduced. In contrast to other approaches not only a single illuminant estimate is provided but a set of likelihoods for a given illumination set. Experimental results on the benchmark Simon Fraser image database show the combined method to clearly outperform purely statistical and purely physical algorithms.
Gerald Schaefer, Steven D. Hordley, Graham D. Finlayson
CVPR (1)1
2005 Illuminant and device invariant colour using histogram equalisation
Graham D. Finlayson, Steven D. Hordley, Gerald Schaefer
Pattern Recognit.3
2004 JPEG2000 vs. JPEG from an image retrieval point of view
abstract
It is well known that JPEG2000 has a variety of advantages over its predecessor JPEG. JPEG2000 not only allows images to be coded with clearly better visual image quality, it also addresses a series of other issues. Unification of lossless and lossy compression modes, robustness to bit-errors to allow image transmission over noisy channels and provision of regions of interest (ROI) are only some of those that have been incorporated into the new standard. In this paper we look at the JPEG2000 versus JPEG debate from an image retrieval standpoint. While it seems evident that image compression will have a negative effect on the performance of retrieval algorithms our aim is to provide quantitative results of how severe this performance drop would be for JPEG2000 compression in comparison to standard JPEG encoding. Our results show that while high compression causes problems for retrieval of JPEG images, the retrieval performance of JPEG2000 images is almost independent of compression ratio.
Gerald Schaefer
ICIP1
2004 CVPIC image retrieval based on block colour co-occurance matrix and pattern histogram
abstract
Compressed domain image processing techniques are becoming increasingly important. Compressed domain retrieval allows the calculation of image features and hence content-based image retrieval (CBIR) to be performed directly on the compressed data without the need for decoding it beforehand. The colour visual pattern image coding (CVPIC) technique represents a compression algorithm where the compressed form is directly meaningful. Based on CVPIC, we introduce a compressed domain retrieval algorithm that makes immediate use of the fact that colour and pattern information is readily available in the CVPIC domain. Colour features are exploited by building a block co-occurance matrix of colour indices while shape information is represented through pattern histograms. Combining these two types of descriptors results in an efficient and effective image retrieval method that even outperforms popular pixel-based algorithms such as colour histograms, colour coherence vectors and colour correlograms.
Gerald Schaefer, Simon Lieutaud, Guoping Qiu
ICIP1
2002 Compressed domain image retrieval by comparing vector quantization codebooks
Gerald Schaefer
VCIP1
2001 Hue that is invariant to brightness and gamma
abstract
Hue provides a useful and intuitive cue that is used in a variety of computer vision applications. Hue is an attractive feature as it captures intrinsic information about the colour of objects or surfaces in a scene. Moreover, hue is invariant to confounding factors such as illumination brightness. However hue is not stable to all of the types of confounding factors that one might reasonably encounter. Specifically, the RGBs captured in images are sometimes raised to the power gamma. This is done for two reasons. First, to make the images suitable for display (since monitors have an intrinsic non-linearity). Second, applying a gamma is the simplest way to change the contrast in images. It has also been observed that digital cameras often apply a scene dependent gamma type function (which is unknown to the user). In this paper we show that a simple photometric ratio in log RGB space cancels both brightness and gamma. Furthermore, some simple manipulation reveals that the brightness/gamma invariant can usefully be interpreted as a hue in a log opponent colour space. We carried out indexing experiments to evaluate the usefulness of the derived hue correlate. In situations where gamma is held fixed, the new hue supports recognition equal to conventional definitions. In situations where gamma varies the new correlate supports better indexing. The new hue is also found to predict some psychophysical data quite accurately.
Graham D. Finlayson, Gerald Schaefer
BMVC2
2001 Convex and Non-convex Illuminant Constraints for Dichromatic Colour Constancy
abstract
The dichromatic reflectance model introduced by S. Shafer (1985) predicts that the colour signals of most materials fall on a plane spanned by a vector due to the material and a vector that represents the scene illuminant. Since the illuminant is in the span of all dichromatic planes, colour constancy can be achieved by finding the intersection of two or more planes. Unfortunately, this approach has proven to be hard to get to work in practice. First, segmentation needs to be carried out and second, the actual intersection computation is quite unstable: small changes in the orientation of a dichromatic plane can significantly alter the location of the intersection point. We propose to ameliorate the instability problem by regularising the intersection. Specifically, we introduce a constraint on the colour of the illuminant. We show how the intersection problem in the context of convex and non-convex illuminant constraints, based on the distribution of common light sources, can be solved. This algorithm coupled with the simplest of segmentations results in good estimation results for a large set of real images. Estimation performance is significantly better than for the unconstrained algorithm.
Graham D. Finlayson, Gerald Schaefer
CVPR (1)2
2001 JPEG Compressed Domain Image Retrieval by Colour and Texture
Gerald Schaefer
Data Compression Conference1
2001 Solving for Colour Constancy using a Constrained Dichromatic Reflection Model
Graham D. Finlayson, Gerald Schaefer
Int. J. Comput. Vis.2
2000 Constrained Dichromatic Colour Constancy
Graham D. Finlayson, Gerald Schaefer
ECCV (1)2