Eyad Elyan

dblp:49/1289 · DBLP profile ↗
← Back
63ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-8342-9026ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 5 first-author · 24 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Chirality-Aware Grammar-Guided Surgical Action Anticipation from Video
abstract
Anticipating surgical actions requires more than recognising motion patterns, it also demands adherence to procedural logic and the resolution of subtle ambiguities, such as distinguishing mirrored grasp–retract interactions. However, existing Transformer-based models often fall short in this domain, often producing structurally invalid step sequences and misclassifying chirally opposite actions that appear visually similar. To address these limitations, we introduce a neuro-symbolic framework centred on a Probabilistic Temporal Grammar (PTG). The grammar was constructed from a unified corpus of surgical data (Ground-Truth), by capturing procedural structure, temporal priors, and chirality-aware terminals for opposite actions (e.g., push_needle <-> pull_suture) directly into its rules. To enforce causal consistency, the PTG incorporates a Goal-conditioned Multivariate Markov Chain (GcMMC) that models evolving object-action dependencies. Our framework employs a two-stage process: a V-JEPA-powered Transformer generates raw forecasts of future actions and durations, which are then refined by a constrained parsing algorithm guided by the PTG. Candidate futures are jointly scored for structural validity, temporal plausibility, and causal grounding. By explicitly encoding surgical logic into a unified neuro-symbolic system. Experiments across three publicly available surgical datasets show that our approach outperforms state-of-the-art anticipation models. Importantly, by generating interpretable and procedurally consistent forecasts of upcoming actions, PTG establishes the predictive foundation required for proactive robotic assistance and safe human–robot collaboration in the operating room.
Md Rezowan Hossain Ferdous Shuvo, Mahammad Shareef Mekala, Eyad Elyan
HRI3
2026 Spot the Difference: Bilateral Contrastive Representation Learning for Nodule Classification
Sophie Crawford Haynes, Mahammad Shareef Mekala, Eyad Elyan
ICPR (13)3
2025 A Feature Transformation Technique for Improving Ensemble Learning Systems
Truong Thanh Nguyen, Hieu Vu, Eyad Elyan, Thanh Son Vu, Tien Thanh Nguyen
ACIIDS (2)4
2025 An Evolutionary Neural Architecture Search-Based Approach for Time Series Forecasting
abstract
Time series forecasting (TSF) is one of the most prevalent research topics in artificial intelligence and has garnered significant attention in the research community. In recent years, significant breakthroughs have been made in TSF research, shifting from traditional statistical models to deep learning (DL), and attention-based methods. While attention-based methods excel at capturing global dependencies, they often face challenges in effectively modeling local patterns. Additionally, the design of these networks typically demands substantial human expertise, experimental work, and manual configuration. To address these issues, we propose an Evolutionary Neural Architecture Search for Time Series Forecasting, entitled ENAS-TSF to automate TSF architecture design. Concretely, we propose a novel local/ global context module encoding strategy to define a search space. Each local context module includes various convolutions with different kernel sizes to capture temporal dependencies, aiming to enhance the local features and pattern recognition. Global encoding meanwhile contains attention mechanisms and feedforward layers for global context modeling. We propose an evolutionary neural architecture search approach to identify the optimal ENAS-TSF architectures, achieving the ideal balance of local/ global context modeling. Extensive experiments on common benchmark datasets show that ENAS-TSF achieves competitive performance compared to state-of-the-art methods, demonstrating the proposed framework’s effectiveness.
Tien Thanh Nguyen, Eyad Elyan
CEC3
2025 Offshore Asset Inspection Redefined: Expert- Validated Deep Learning for Critical Defects
abstract
The offshore energy industry faces challenges in maintaining ageing infrastructure, with over half of North Sea platforms past their 25-year design life. This creates a need for scalable inspection methods beyond traditional manual reviews. We present a collaborative Human-AI framework that assists in detecting structural defects whilst preserving expert oversight under challenging marine conditions. Our two-stage system first employs a lightweight classifier to filter video frames by risk level. High-risk frames are then analysed by a modified Pyramid Attention Network that performs precise defect localisation. Experts validate the results at both stages, ensuring the system's continuous improvement. To better identify rare but critical flaws, we design Enhanced Tversky, a composite loss function to mitigate severe class imbalance by explicitly prioritising rare yet safety-critical defects like cracks. Evaluation on 5,525 classified frames and 1,013 segmented images demonstrates F2 score of 87.59% and a mean IoU of 78.73%, with crack detection reaching a crucial 89.62% F2 score. The framework reduces expert review time by 3.88 x whilst maintaining safety standards, offering a practical approach to scaling offshore inspection capabilities through Human-AI collaboration.
Urmila Gurung, Mahammad Shareef Mekala, Carlos Francisco Moreno-García, Nikolaus Zolnhofer, Barry Marshall, Eyad Elyan
DSAA6
2025 KD-LSRED : Knowledge Distillation for Lightweight Symbol Recognition in Engineering Diagrams
Ikenna Ekeke, Carlos Francisco Moreno-García, Eyad Elyan
ICDAR (5)3
2025 Mitigating Class Imbalance in Multiclass Educational Data: A Hybrid One-vs-One and Density-Based Resampling Approach
John-Bosco Diekuu, Mahammad Shareef Mekala, John P. Isaacs, Eyad Elyan
IDEAL (1)4
2025 On the Impact of Greyscale ImageNet Pre-training for Chest X-Ray Model Transferability
Sophie Crawford Haynes, Pamela Johnston, Eyad Elyan
IDEAL (1)3
2025 MSBATN: Multi-Stage Boundary-Aware Transformer Network for action segmentation in untrimmed surgical videos
Rezowan Shuvo, Mahammad Shareef Mekala, Eyad Elyan
Comput. Vis. Image Underst.3
2025 Blind sonar image quality assessment via machine learning: Leveraging micro- and macro-scale texture and contour features in the wavelet domain
abstract
In subsea environments, sound navigation and ranging (SONAR) images are widely used for exploring and monitoring infrastructures due to their robustness and insensitivity to low-light conditions. However, their quality can degrade during acquisition and transmission, where standard SONAR image processing techniques can hardly produce high-quality outcomes. An effective image quality assessment (IQA) method can assess their usefulness and aid to develop refinement techniques by identifying the degradation issues, ensuring the reliability of SONAR data. Existing methods often fail to account for degradations from noise, distortion, and resolution changes simultaneously. To address this challenge, we propose a new blind quality assessment method that measures the overall quality of SONAR images by quantifying both the perceptual and utility qualities using the micro- and macro-scale texture and contour features derived from the wavelet domain. By combining the local binary pattern (LBP) micro-scale texture features with the proposed histograms of Schmid Gabor-like edge maps as macro-scale features, a support vector regression model is learned to map from these features to subjective quality scores. Extensive experiments have demonstrated the superiority of our method over existing SONAR IQA techniques on distorted and reconstructed super-resolution side-scan, acoustic lens, and forward-looking SONAR images. Specifically, our method achieves Pearson’s and Spearman’s correlation metrics of 0.8616 and 0.8541, respectively, for distorted SONAR images, demonstrating improvements of 4.69% and 4.8%. For reconstructed super-resolution SONAR images, our method attains correlation metrics of 0.9415 and 0.9408, reflecting improvements of 0.8% and 1.6% over the second-best method, respectively. To facilitate ease of access, a comprehensive list of key abbreviations and their full names is provided in Table A.9 in the Appendix section. The source code of the proposed method will be shared at https://github.com/hfarhaditolie/BSIQA .
Hamidreza Farhadi Tolie, Jinchang Ren, Rongjun Chen 0001, Huimin Zhao 0001, Eyad Elyan
Eng. Appl. Artif. Intell.5
2025 Towards fully automated processing and analysis of construction diagrams: AI-powered symbol detection
abstract
Abstract Construction drawings are frequently stored in undigitised formats and consequently, their analysis requires substantial manual effort. This is true for many crucial tasks, including material takeoff where the purpose is to obtain a list of the equipment and respective amounts required for a project. Engineering drawing digitisation has recently attracted increased attention, however construction drawings have received considerably less interest compared to other types. To address these issues, this paper presents a novel framework for the automatic processing of construction drawings. Extensive experiments were performed using two state-of-the-art deep learning models for object detection in challenging high-resolution drawings sourced from industry. The results show a significant reduction in the time required for drawing analysis. Promising performance was achieved for symbol detection across various classes, with a mean average precision of 79% for the YOLO-based method and 83% for the Faster R-CNN-based method. This framework enables the digital transformation of construction drawings, improving tasks such as material takeoff and many others.
Laura Jamieson, Carlos Francisco Moreno-García, Eyad Elyan
Int. J. Document Anal. Recognit.3
2025 Artificial Intelligence-Based Conversational Agents Used for Sustainable Fashion: Systematic Literature Review
abstract
In the past five years, the textile industry has undergone significant transformations in response to evolving fashion trends and increased consumer garment turnover.To address the environmental impacts of fast fashion, the industry is embracing artificial intelligence (AI) and immersive technologies, particularly leveraging conversational agents as personalized guides for sustainable fashion practices.In this research article, we conduct a systematic literature review to categorize techniques, platforms, and applications of conversational agents in promoting sustainability within the fashion industry.Additionally, the review aims to scrutinize the solutions offered, identify gaps in the existing literature, and provide insights into the effectiveness and limitations of these conversational agents.Utilizing a predefined search strategy on IEEE Xplore, Google Scholar, SCOPUS, and Web of Science, 15 relevant articles were selected through a step-by-step procedure based on the guidelines of the PRISMA framework.The findings reveal a notable global interest in AIpowered conversational agents, with Italy emerging as a significant center for research in this domain.The studies predominantly focus on consumer perceptions and intentions regarding the adoption of AI technologies, indicating a broader curiosity about how individuals incorporate such innovations into their daily lives.Moreover, a substantial proportion of the studies employ diverse methods, reflecting a comprehensive approach to understanding the functionality and performance of conversational agents in various contexts.While acknowledging the historical precedence of text-based agents, the review highlights a research gap related to embodied agents.The conclusion emphasizes the need for continued exploration, particularly in understanding the broader impact of these technologies on creating sustainable and environmentally friendly business models in the e-retail sector.
Diana S. Hernandez Manzo, Eyad Elyan, John P. Isaacs
Int. J. Hum. Comput. Interact.3
2025 A Multimodel-Based Screening Framework for C-19 Using Deep Learning-Inspired Data Fusion
abstract
In recent times, there has been a notable rise in the utilization of Internet of Medical Things (IoMT) frameworks particularly those based on edge computing, to enhance remote monitoring in healthcare applications. Most existing models in this field have been developed temperature screening methods using RCNN, face temperature encoder (FTE), and a combination of data from wearable sensors for predicting respiratory rate (RR) and monitoring blood pressure. These methods aim to facilitate remote screening and monitoring of Severe Acute Respiratory Syndrome Coronavirus (SARS-CoV) and COVID-19. However, these models require inadequate computing resources and are not suitable for lightweight environments. We propose a multimodal screening framework that leverages deep learning-inspired data fusion models to enhance screening results. A Variation Encoder (VEN) design proposes to measure skin temperature using Regions of Interest (RoI) identified by YoLo. Subsequently, the multi-data fusion model integrates electronic records features with data from wearable human sensors. To optimize computational efficiency, a data reduction mechanism is added to eliminate unnecessary features. Furthermore, we employ a contingent probability method to estimate distinct feature weights for each cluster, deepening our understanding of variations in thermal and sensory data to assess the prediction of abnormal COVID-19 instances. Simulation results using our lab dataset demonstrate a precision of 95.2%, surpassing state-of-the-art models due to the thoughtful design of the multimodal data-based feature fusion model, weight prediction factor, and feature selection model.
Achyut Shankar, Rizwan Patan, Mahammad Shareef Mekala, Eyad Elyan, Amir Hossein Gandomi, Carsten Maple, Joel J. P. C. Rodrigues
IEEE J. Biomed. Health Informatics4
2024 A Multiclass Imbalanced Dataset Classification of Symbols from Piping and Instrumentation Diagrams
Laura Jamieson, Carlos Francisco Moreno-García, Eyad Elyan
ICDAR (1)3
2024 A Novel Ensemble Aggregation Method Based on Deep Learning Representation
Truong Thanh Nguyen, Eyad Elyan, Tien Thanh Nguyen, Martin Longmuir
ICPR (24)2
2024 FAT: Fusion-Attention Transformer for Remaining Useful Life Prediction
Eyad Elyan, Will Vorley, Joe Goodlad, Tien Thanh Nguyen
ICPR (10)2
2024 DICAM: Deep Inception and Channel-wise Attention Modules for underwater image enhancement
Hamidreza Farhadi Tolie, Jinchang Ren, Eyad Elyan
Neurocomputing3
2024 Which classifiers are connected to others? An optimal connection framework for multi-layer ensemble systems
abstract
• A multi-layer ensemble connects each classifier to multiple ones in prior layer. • Connections signify the use of previous-layer-classifiers’ outputs as new inputs. • A binary encoding scheme is proposed to encode the topology of proposed ensemble. • The optimal topology of proposed ensemble is found by using Differential Evolution. • Our ensemble performs better than benchmark algorithms on experimental datasets. Ensemble learning is a powerful machine learning strategy that combines multiple models e.g. classifiers to improve predictions beyond what any single model can achieve. Until recently, traditional ensemble methods typically use only one layer of models which limits the exploration of different aspects in the classifiers’ predictions. On the other hand, the rise of deep learning has introduced multi-layer architectures that can learn complex functions by transforming data into multiple levels of representation. This characteristic of deep learning suggests that multi-layer ensembles may potentially provide better performance compared to single-layer ensembles. However, a problem which might arise is that in the subsequent layers, not all the inputs to a classifier are desirable, leading to lower performance. In this paper, we introduce a novel multi-layer ensemble of classifiers named COME in which each classifier at a specific layer is connected to multiple classifiers in the previous layer. These connections signify the use of the previous-layer-classifiers’ outputs as inputs for training the current layer's classifier. Each classifier can be connected to different classifiers in the previous layer, which allows inputs in each layer to be optimally selected. We propose a binary encoding scheme to encode the topology of the proposed multi-layer ensemble with defined connections between layers. Differential Evolution, a popular evolutionary computation method, is used as the optimisation algorithm to search for the optimal set of connections. Experimental results on 30 datasets from the UCI Machine Learning Repository and OpenML demonstrate that our proposed ensemble outperforms many state-of-the-art ensemble learning algorithms.
Tien Thanh Nguyen, Alan Wee-Chung Liew, Eyad Elyan, John A. W. McCall
Knowl. Based Syst.4
2024 Generalisation challenges in deep learning models for medical imagery: insights from external validation of COVID-19 classifiers
abstract
Abstract The generalisability of deep neural network classifiers is emerging as one of the most important challenges of our time. The recent COVID-19 pandemic led to a surge of deep learning publications that proposed novel models for the detection of COVID-19 from chest x-rays (CXRs). However, despite the many outstanding metrics reported, such models have failed to achieve widespread adoption into clinical settings. The significant risk of real-world generalisation failure has repeatedly been cited as one of the most critical concerns, and is a concern that extends into general medical image modelling. In this study, we propose a new dataset protocol and, using this, perform a thorough cross-dataset evaluation of deep neural networks when trained on a small COVID-19 dataset, comparable to those used extensively in recent literature. This allows us to quantify the degree to which these models can generalise when trained on challenging, limited medical datasets. We also introduce a novel occlusion evaluation to quantify model reliance on shortcut features. Our results indicate that models initialised with ImageNet weights then fine-tuned on small COVID-19 datasets, a standard approach in the literature, facilitate the learning of shortcut features, resulting in unreliable, poorly generalising models. In contrast, pre-training on related CXR imagery can stabilise cross-dataset performance. The CXR pre-trained models demonstrated a significantly smaller generalisation drop and reduced feature dependence outwith the lung region, as indicated by our occlusion test. This paper demonstrates the challenging problem of model generalisation, and the need for further research on developing techniques that will produce reliable, generalisable models when learning with limited datasets.
Sophie Crawford Haynes, Pamela Johnston, Eyad Elyan
Multim. Tools Appl.3
2023 Object-aware Multi-criteria Decision-Making Approach using the Heuristic data-driven Theory for Intelligent Transportation Systems
abstract
Sharing up-to-date information about the surrounding measured by On-Board Units (OBUs) and Roadside Units (RSUs) is crucial in accomplishing traffic efficiency and pedestrians safety towards Intelligent Transportation Systems (ITS). Transferring measured data demands $\geq$10Gbit/s transfer rate and $\geq$1GHz bandwidth though the data is lost due to unusual data transfer size and impaired line of sight (LOS) propagation. Most existing models concentrated on resource optimization instead of measured data optimization. Subsequently, RSU-LiDARs have become increasingly popular in addressing object detection, mapping and resource optimization issues of Edge-based Software-Defined Vehicular Orchestration (ESDVO). In this regard, we design a two-step data-driven optimization approach called Object-aware Multi-criteria Decision-Making (OMDM) approach. First, the surroundings-measured data by RSUs and OBUs is processed by cropping object-enabled frames using YoLo and FRCNN at RSU. The cropped data likely share over the environment based on the RSU Computation-Communication method. Second, selecting the potential vehicle/device is treated as an NP-hard problem that shares information over the network for effective path trajectory and stores the cosine data at the fog server for end-user accessibility. In addition, we use a nonlinear programming multi-tenancy heuristic method to improve resource utilization rates based on device preference predictions (Like detection accuracy and bounding box tracking) which elaborately concentrate in future work. The simulation results agree with the targeted effectiveness of our approach, i.e., mAP($\geq$71%) with processing delay ($\leq3.5\times 10^{6}$bits/slot), and transfer delay ($\leq$3Sms). Our simulation results indicate that our approach is highly effective.
Mahammad Shareef Mekala, Eyad Elyan, Gautam Srivastava 0001
DSAA2
2023 Digital Transformation for Offshore Assets: A Deep Learning Framework for Weld Classification in Remote Visual Inspections
Luis Toral, Eyad Elyan, Carlos Francisco Moreno-García, Jan Stander
EANN2
2023 Unmasking the Imposters: Task-specific feature learning for face presentation attack detection
abstract
Presentation attacks pose a threat to the reliability of face recognition systems. A photograph, a video, or a mask representing an authorised user can be used to circumvent the face recognition system. Recent research has demonstrated high accuracy in intra-dataset evaluations using existing face presentation attack detection models. Nonetheless, these models did not achieve similar performance when evaluated across datasets due to limited generalisation. Consequently, this article presents task-specific feature learning using deep pre-trained models. Model performance was evaluated using three public datasets: the SiW dataset was used for intra-dataset evaluation, while CASIA and Replay Attack were used for cross-dataset evaluation. Custom task-specific feature learning, compared to deep and hybrid models, demonstrated improved cross-dataset performance and exhibited more generalisability. The results suggest future direction for further research toward improving the model's generalisation using custom task-specific feature learning.
Faseela Abdullakutty, Eyad Elyan, Pamela Johnston
IJCNN2
2022 Cross Domain Evaluation of Text Detection Models
Adamu Ali-Gombe, Eyad Elyan, Carlos Francisco Moreno-García, Chrisina Jayne
ICANN (3)2
2021 Weighted Ensemble of Deep Learning Models based on Comprehensive Learning Particle Swarm Optimization for Medical Image Segmentation
abstract
In recent years, deep learning has rapidly become a method of choice for segmentation of medical images. Deep neural architectures such as UNet and FPN have achieved high performances on many medical datasets. However, medical image analysis algorithms are required to be reliable, robust, and accurate for clinical applications which can be difficult to achieve for some single deep learning methods. In this study, we introduce an ensemble of classifiers for semantic segmentation of medical images. The ensemble of classifiers here is a set of various deep learning-based classifiers, aiming to achieve better performance than using a single classifier. We propose a weighted ensemble method in which the weighted sum of segmentation outputs by classifiers is used to choose the final segmentation decision. We use a swarm intelligence algorithm namely Comprehensive Learning Particle Swarm Optimization to optimize the combining weights. Dice coefficient, a popular performance metric for image segmentation, is used as the fitness criteria. Experiments conducted on some medical datasets of the CAMUS competition on cardiographic image segmentation show that our method achieves better results than both the constituent segmentation models and the reported model of the CAMUS competition.
Tien Thanh Nguyen, Carlos Francisco Moreno-García, Eyad Elyan, John A. W. McCall
CEC4
2021 Face Spoof Detection: An Experimental Framework
Faseela Abdullakutty, Eyad Elyan, Pamela Johnston
EANN2
2021 Face Detection with YOLO on Edge
Adamu Ali-Gombe, Eyad Elyan, Carlos Francisco Moreno-García, Johan Zwiegelaar
EANN2
2021 Class-Decomposition and Augmentation for Imbalanced Data Sentiment Analysis
abstract
Significant progress has been made in the area of text classification and natural language processing. However, like many other datasets from across different domains, text-based datasets may suffer from class-imbalance. This problem leads to model's bias toward the majority class instances. In this paper, we present a new approach to handle class-imbalance in text data by means of unsupervised learning algorithms. We present class-decomposition using two different unsupervised methods, namely k-means and Density-Based Spatial Clustering of Applications with Noise, applied to two different sentiment analysis data sets. The experimental results show that utilizing clustering to find within-class similarities can lead to significant improvement in learning algorithm's performances as well as reducing the dominance of the majority class instances without causing information loss.
Carlos Francisco Moreno-García, Chrisina Jayne, Eyad Elyan
IJCNN3
2021 On the class overlap problem in imbalanced data classification
Pattaramon Vuttipittayamongkol, Eyad Elyan, Andrei Petrovski 0001
Knowl. Based Syst.2
2021 CDSMOTE: class decomposition and synthetic minority class oversampling technique for imbalanced-data classification
abstract
Abstract Class-imbalanced datasets are common across several domains such as health, banking, security, and others. The dominance of majority class instances (negative class) often results in biased learning models, and therefore, classifying such datasets requires employing some methods to compact the problem. In this paper, we propose a new hybrid approach aiming at reducing the dominance of the majority class instances using class decomposition and increasing the minority class instances using an oversampling method. Unlike other undersampling methods, which suffer data loss, our method preserves the majority class instances, yet significantly reduces its dominance, resulting in a more balanced dataset and hence improving the results. A large-scale experiment using 60 public datasets was carried out to validate the proposed methods. The results across three standard evaluation metrics show the comparable and superior results with other common and state-of-the-art techniques.
Eyad Elyan, Carlos Francisco Moreno-García, Chrisina Jayne
Neural Comput. Appl.1
2020 Towards a Reliable Face Recognition System
Adamu Ali-Gombe, Eyad Elyan, Johan Zwiegelaar
EANN2
2020 Symbols in Engineering Drawings (SiED): An Imbalanced Dataset Benchmarked by Convolutional Neural Networks
Eyad Elyan, Carlos Francisco Moreno-García, Pamela Johnston
EANN1
2020 Predicting Permeability Based on Core Analysis
Harry Kontopoulos, Hatem Ahriz, Eyad Elyan, Richard Arnold
EANN3
2020 Deep Learning for Text Detection and Recognition in Complex Engineering Diagrams
abstract
Engineering drawings such as Piping and Instrumentation Diagrams contain a vast amount of text data which is essential to identify shapes, pipeline activities, tags, amongst others. These diagrams are often stored in undigitised format, such as paper copy, meaning the information contained within the diagrams is not readily accessible to inspect and use for further data analytics. In this paper, we make use of the benefits of recent deep learning advances by selecting models for both text detection and text recognition, and apply them to the digitisation of text from within real world complex engineering diagrams. Results show that 90% of text strings were detected including vertical text strings, however certain non text diagram elements were detected as text. Text strings were obtained by the text recognition method for 86% of detected text instances. The findings show that whilst the chosen Deep Learning methods were able to detect and recognise text which occurred in simple scenarios, more complex representations of text including those text strings located in close proximity to other drawing elements were highlighted as a remaining challenge.
Laura Jamieson, Carlos Francisco Moreno-García, Eyad Elyan
IJCNN3
2020 Improved Overlap-based Undersampling for Imbalanced Dataset Classification with Application to Epilepsy and Parkinson's Disease
abstract
Classification of imbalanced datasets has attracted substantial research interest over the past decades. Imbalanced datasets are common in several domains such as health, finance, security and others. A wide range of solutions to handle imbalanced datasets focus mainly on the class distribution problem and aim at providing more balanced datasets by means of resampling. However, existing literature shows that class overlap has a higher negative impact on the learning process than class distribution. In this paper, we propose overlap-based undersampling methods for maximizing the visibility of the minority class instances in the overlapping region. This is achieved by the use of soft clustering and the elimination threshold that is adaptable to the overlap degree to identify and eliminate negative instances in the overlapping region. For more accurate clustering and detection of overlapped negative instances, the presence of the minority class at the borderline areas is emphasized by means of oversampling. Extensive experiments using simulated and real-world datasets covering a wide range of imbalance and overlap scenarios including extreme cases were carried out. Results show significant improvement in sensitivity and competitive performance with well-established and state-of-the-art methods.
Pattaramon Vuttipittayamongkol, Eyad Elyan
Int. J. Neural Syst.2
2020 Response to Discussion on "Improved Overlap-Based Undersampling for Imbalanced Dataset Classification with Application to Epilepsy and Parkinson's Disease, "
abstract
In the paper Improved Overlap-Based Undersampling for Imbalanced Dataset Classification with Application to Epilepsy and Parkinson's Disease, the authors introduced two new methods that address the class overlap problem in imbalanced datasets. The methods involve identification and removal of potentially overlapped majority class instances. Extensive evaluations were carried out using 136 datasets and compared against several state-of-the-art methods. Results showed competitive performance with those methods, and statistical tests proved significant improvement in classification results. The discussion on the paper related to the behavioral analysis of class overlap and method validation was raised by Fernández. In this article, the response to the discussion is delivered. Detailed clarification and supporting evidence to answer all the points raised are provided.
Pattaramon Vuttipittayamongkol, Eyad Elyan
Int. J. Neural Syst.2
2020 Neighbourhood-based undersampling approach for handling imbalanced and overlapped data
Pattaramon Vuttipittayamongkol, Eyad Elyan
Inf. Sci.2
2020 Video tampering localisation using features learned from authentic content
abstract
Video tampering detection remains an open problem in the field of digital media forensics. As video manipulation techniques advance, it becomes easier for tamperers to create convincing forgeries that can fool human eyes. Deep learning methods have already shown great promise in discovering effective features from data, particularly in the image domain; however, they are exceptionally data hungry. Labelled datasets of varied, state-of-the-art, tampered video which are large enough to facilitate machine learning do not exist and, moreover, may never exist while the field of digital video manipulation is advancing at such an unprecedented pace. Therefore, it is vital to develop techniques which can be trained on authentic or synthesised video but used to localise the patterns of manipulation within tampered videos. In this paper, we developed a framework for tampering detection which derives features from authentic content and utilises them to localise key frames and tampered regions in three publicly available tampered video datasets. We used convolutional neural networks to estimate quantisation parameter, deblock setting and intra/inter mode of pixel patches from an H.264/AVC sequence. Extensive evaluation suggests that these features can be used to aid localisation of tampered regions within video.
Pamela Johnston, Eyad Elyan, Chrisina Jayne
Neural Comput. Appl.2
2020 Deep learning for symbols detection and classification in engineering drawings
Eyad Elyan, Laura Jamieson, Adamu Ali-Gombe
Neural Networks1
2019 Multiple Fake Classes GAN for Data Augmentation in Face Image Dataset
abstract
Class-imbalanced datasets often contain one or more class that are under-represented in a dataset. In such a situation, learning algorithms are often biased toward the majority class instances. Therefore, some modification to the learning algorithm or the data itself is required before attempting a classification task. Data augmentation is one common approach used to improve the presence of the minority class instances and rebalance the dataset. However, simple augmentation techniques such as applying some affine transformation to the data, may not be sufficient in extreme cases, and often do not capture the variance present in the dataset. In this paper, we propose a new approach to generate more samples from minority class instances based on Generative Adversarial Neural Networks (GAN). We introduce a new Multiple Fake Class Generative Adversarial Networks (MFC-GAN) and generate additional samples to rebalance the dataset. We show that by introducing multiple fake class and oversampling, the model can generate the required minority samples. We evaluate our model on face generation task from attributes using a reduced number of samples in the minority class. Results obtained showed that MFC-GAN produces plausible minority samples that improve the classification performance compared with state-of-the-art AC-GAN generated samples.
Adamu Ali-Gombe, Eyad Elyan, Chrisina Jayne
IJCNN2
2019 MFC-GAN: Class-imbalanced dataset classification using Multiple Fake Class Generative Adversarial Network
Adamu Ali-Gombe, Eyad Elyan
Neurocomputing2
2019 New trends on digitisation of complex engineering drawings
abstract
Engineering drawings are commonly used across different industries such as oil and gas, mechanical engineering and others. Digitising these drawings is becoming increasingly important. This is mainly due to the legacy of drawings and documents that may provide rich source of information for industries. Analysing these drawings often requires applying a set of digital image processing methods to detect and classify symbols and other components. Despite the recent significant advances in image processing, and in particular in deep neural networks, automatic analysis and processing of these engineering drawings is still far from being complete. This paper presents a general framework for complex engineering drawing digitisation. A thorough and critical review of relevant literature, methods and algorithms in machine learning and machine vision is presented. Real-life industrial scenario on how to contextualise the digitised information from specific type of these drawings, namely piping and instrumentation diagrams, is discussed in details. A discussion of how new trends on machine vision such as deep learning could be applied to this domain is presented with conclusions and suggestions for future research directions.
Carlos Francisco Moreno-García, Eyad Elyan, Chrisina Jayne
Neural Comput. Appl.2
2018 Deep Imitation Learning with Memory for Robocup Soccer Simulation
Ahmed Hussein 0001, Eyad Elyan, Chrisina Jayne
EANN2
2018 Toward Video Tampering Exposure: Inferring Compression Parameters from Pixels
Pamela Johnston, Eyad Elyan, Chrisina Jayne
EANN2
2018 Overlap-Based Undersampling for Improving Imbalanced Data Classification
Pattaramon Vuttipittayamongkol, Eyad Elyan, Andrei Petrovski 0001, Chrisina Jayne
IDEAL (1)2
2018 Few-shot Classifier GAN
abstract
Fine-grained image classification with a few-shot classifier is a highly challenging open problem at the core of a numerous data labeling applications. In this paper, we present Few-shot Classifier Generative Adversarial Network as an approach for few-shot classification. We address the problem of few-shot classification by designing a GAN model in which the discriminator and the generator compete to output labeled data in any case. In contrast to previous methods, our techniques generate then classify images into multiple fake or real classes. A key innovation of our adversarial approach is to allow fine- grained classification using multiple fake classes with semi- supervised deep learning. A major strength of our techniques lies in its label-agnostic characteristic, in the sense that the system handles both labeled and unlabeled data during training. We validate quantitatively our few-shot classifier on the MNIST and SVHN datasets by varying the ratio of labeled data over unlabeled data in the training set. Our quantitative analysis demonstrates that our techniques produce better classification performance when using multiple fake classes and larger amount of unlabelled data.
Adamu Ali-Gombe, Eyad Elyan, Yann Savoye, Chrisina Jayne
IJCNN2
2018 Symbols Classification in Engineering Drawings
abstract
Technical drawings are commonly used across different industries such as Oil and Gas, construction, mechanical and other types of engineering. In recent years, the digitization of these drawings is becoming increasingly important. In this paper, we present a semi-automatic and heuristic-based approach to detect and localise symbols within these drawings. This includes generating a labeled dataset from real world engineering drawings and investigating the classification performance of three different state-of the art supervised machine learning algorithms. In order to improve the classification accuracy the dataset was pre-processed using unsupervised learning algorithms to identify hidden patterns within classes. Testing and evaluating the proposed methods on a dataset of symbols representing one standard of drawings, namely Process and Instrumentation (P&ID) showed very competitive results.
Eyad Elyan, Carlos Francisco Moreno-García, Chrisina Jayne
IJCNN1
2018 Spatial Effects of Video Compression on Classification in Convolutional Neural Networks
abstract
A collection of computer vision applications reuse pre-learned features to analyse video frame-by-frame. Those features are classically learned by Convolutional Neural Networks (CNN) trained on high quality images. However, available video content is almost always subject to compression which is nearly never considered during the analysis process. In this paper, we present an empirical study to measure how the visual discrepancy of compressed data limit the learning performance of the CNN model. The learning performance is evaluated using a benchmark of synthetic datasets compressed at various levels using H.264/AVC. We measure the image quality quantitatively using classical evaluation metrics such as Peak Signal to Noise Ratio and Structural SIMilarity. A cross-evaluation is performed to measure the robustness of the CNN model in processing for a wide range of quality-varying visual data. Our experimental results have shown that the performance of the CNN depends on the compression rate. The results show that, in general, higher compression results in lower performance. However performance on lower quality test data can be improved by using lower quality data for CNN training. Finally, our work demonstrates that conditioning the CNN with the compression properties could potentially lead to better learning.
Pamela Johnston, Eyad Elyan, Chrisina Jayne
IJCNN2
2018 Deep imitation learning for 3D navigation tasks
abstract
Deep learning techniques have shown success in learning from raw high-dimensional data in various applications. While deep reinforcement learning is recently gaining popularity as a method to train intelligent agents, utilizing deep learning in imitation learning has been scarcely explored. Imitation learning can be an efficient method to teach intelligent agents by providing a set of demonstrations to learn from. However, generalizing to situations that are not represented in the demonstrations can be challenging, especially in 3D environments. In this paper, we propose a deep imitation learning method to learn navigation tasks from demonstrations in a 3D environment. The supervised policy is refined using active learning in order to generalize to unseen situations. This approach is compared to two popular deep reinforcement learning techniques: deep-Q-networks and Asynchronous actor-critic (A3C). The proposed method as well as the reinforcement learning methods employ deep convolutional neural networks and learn directly from raw visual input. Methods for combining learning from demonstrations and experience are also investigated. This combination aims to join the generalization ability of learning by experience with the efficiency of learning by imitation. The proposed methods are evaluated on 4 navigation tasks in a 3D simulated environment. Navigation tasks are a typical problem that is relevant to many real applications. They pose the challenge of requiring demonstrations of long trajectories to reach the target and only providing delayed rewards (usually terminal) to the agent. The experiments show that the proposed method can successfully learn navigation tasks from raw visual input while learning from experience methods fail to learn an effective policy. Moreover, it is shown that active learning can significantly improve the performance of the initially learned policy using a small number of active samples.
Ahmed Hussein 0001, Eyad Elyan, Mohamed Medhat Gaber, Chrisina Jayne
Neural Comput. Appl.2
2017 Fish Classification in Context of Noisy Images
Adamu Ali-Gombe, Eyad Elyan, Chrisina Jayne
EANN2
2017 Heuristics-Based Detection to Improve Text/Graphics Segmentation in Complex Engineering Drawings
Carlos Francisco Moreno-García, Eyad Elyan, Chrisina Jayne
EANN2
2017 Deep reward shaping from demonstrations
abstract
Deep reinforcement learning is rapidly gaining attention due to recent successes in a variety of problems. The combination of deep learning and reinforcement learning allows for a generic learning process that does not consider specific knowledge of the task. However, learning from scratch becomes more difficult when tasks involve long trajectories with delayed rewards. The chances of finding the rewards using trial and error become much smaller compared to tasks where the agent continuously interacts with the environment. This is the case in many real life applications which poses a limitation to current methods. In this paper we propose a novel method for combining learning from demonstrations and experience to expedite and improve deep reinforcement learning. Demonstrations from a teacher are used to shape a potential reward function by training a deep supervised convolutional neural network. The shaped function is added to the reward function used in deep-Q-learning (DQN) to perform off-policy training through trial and error. The proposed method is demonstrated on navigation tasks that are learned from raw pixels without utilizing any knowledge of the problem. Navigation tasks represent a typical AI problem that is relevant to many real applications and where only delayed rewards (usually terminal) are available to the agent. The results show that using the proposed shaped rewards significantly improves the performance of the agent over standard DQN. This improvement is more pronounced the sparser the rewards are.
Ahmed Hussein 0001, Eyad Elyan, Mohamed Medhat Gaber, Chrisina Jayne
IJCNN2
2017 A genetic algorithm approach to optimising random forests applied to class engineered data
Eyad Elyan, Mohamed Medhat Gaber
Inf. Sci.1
2016 An Outlier Ranking Tree Selection Approach to Extreme Pruning of Random Forests
Khaled Fawagreh, Mohamed Medhat Gaber, Eyad Elyan
EANN3
2016 Deep Active Learning for Autonomous Navigation
Ahmed Hussein 0001, Mohamed Medhat Gaber, Eyad Elyan
EANN3
2016 A fine-grained Random Forests using class decomposition: an application to medical diagnosis
Eyad Elyan, Mohamed Medhat Gaber
Neural Comput. Appl.1
2014 Diversified Random Forests Using Random Subspaces
Khaled Fawagreh, Mohamed Medhat Gaber, Eyad Elyan
IDEAL3
2014 Rebuilding Visual Vocabulary via Spatial-temporal Context Similarity for Video Retrieval
Lei Wang 0198, Eyad Elyan, Dawei Song 0001
MMM (1)2
2014 The Cognitive Benefit of Dynamic Representations on Procedural Skill Acquisition: A Computational Modeling Approach
abstract
Cognitive computational modeling is a viable methodology for further investigation of the hitherto inconclusive findings on the cognitive benefits of dynamic versus static visualization components of instructions. This is more so as contemporary cognitive architectures such as the Adaptive Control of Thought–Rational (ACT–R) 6.0 are increasingly applied to traditional cognitive psychology research problems. The application of this methodology is, however, restricted by the limited capability of existing architectures for implementing detailed atomic motor actions such as those involved in complex skill acquisition and performance. This article presents a 2-component computational modeling methodology for investigating the cognitive processes involved in the acquisition and performance of skilled motor tasks. The approach specifies a novel combination of a sequence-of-point technique with a movement control mechanism to implement variously acquired cognitive mental task representations and their intertwined role in postlearning performance as evident in the atomic control of motor actions. This paradigm is validated for 2 experiments using incrementally developed cognitive models developed in ACT–R 6.0. The model's quantitative outputs correlate significantly with equivalent empirical human data. This has implications for multimedia instructional design, especially where rapid, transferrable skill acquisition is desired on initial exposure.
Olurotimi Richard Akinlofa, Patrik O'Brian Holt, Eyad Elyan
Int. J. Hum. Comput. Interact.3
2013 Effect of Interface Dynamism on Learning Procedural Motor Skills
abstract
The effectiveness of dynamic versus static visualizations in computer-based training (CBT) systems has generated a lot of research effort with divergent findings. The work reported in this paper examines a novel paradigm that learning a procedural motor skill may be enhanced by instructional visualizations that optimizes the construction of mental task models. We investigated the interaction of different interface visualizations of a CBT system with the cognitive characteristics of trainees by comparing three conditions of interface dynamism in a mechanical motor skills learning task. Ninety-one participants across three treatment groups performed a disassembly motor task. After controlling for effects of spatial visualization abilities, participants who used training interfaces with dynamic information content completed the post-learning motor task faster and more accurately than those who used interfaces with a static visual content. These findings suggest that instructional interfaces having motor coordinating information, which is intrinsic to the execution of procedural motor tasks, are more suitable for CBT of novice trainees. It may also imply the possibility of a common approach to the design and implementation of CBT systems, which is independent of learner's cognitive abilities.
Olurotimi Richard Akinlofa, Patrik O'Brian Holt, Eyad Elyan
Interact. Comput.3
2012 Improving bag-of-visual-words model with spatial-temporal correlation for video retrieval
abstract
Most of the state-of-art approaches to Query-by-Example (QBE) video retrieval are based on the Bag-of-visual-Words (BovW) representation of visual content. It, however, ignores the spatial-temporal information, which is important for similarity measurement between videos. Direct incorporation of such information into the video data representation for a large scale data set is computationally expensive in terms of storage and similarity measurement. It is also static regardless of the change of discriminative power of visual words for different queries. To tackle these limitations, in this paper, we propose to discover Spatial-Temporal Correlations (STC) imposed by the query example to improve the BovW model for video retrieval. The STC, in terms of spatial proximity and relative motion coherence between different visual words, is crucial to identify the discriminative power of the visual words. We develop a novel technique to emphasize the most discriminative visual words for similarity measurement, and incorporate this STC-based approach into the standard inverted index architecture. Our approach is evaluated on the TRECVID2002 and CC\_WEB\_VIDEO datasets for two typical QBE video retrieval tasks respectively. The experimental results demonstrate that it substantially improves the BovW model as well as a state of the art method that also utilizes spatial-temporal information for QBE video retrieval.
Lei Wang 0198, Dawei Song 0001, Eyad Elyan
CIKM3
2012 Enhanced interactivity and engagement: Learning by doing to simplify mathematical concepts in computer graphics and animation
abstract
Interactive 3D Animation is a module that is taught at the School of Computing for fourth year students. Amongst the learning outcomes of this module is the understanding and use of B-Splines, Bezier Curves, Free Form Deformation (FFD), hierarchal modeling, and inverse and forward kinematics. Such concepts appear to be challenging to students and demanding to the instructors. In this paper we present an intuitive, practical and non-mathematical approach to teach students these concepts. During the lecture, and before introducing the formal definitions of these topics, the students are asked to carry out simple and short activities using an appropriate software tool. These activities are closely related to the topic to be discussed and followed by questions and discussion. Finally, formal definitions are introduced to give a mathematical meaning to what the students have already done. Overall, results indicate that engaging students in practical activities and reflecting on these activities help them better understand some of the challenging mathematical theories.
Eyad Elyan
EDUCON1
2011 Video Retrieval Based on Words-of-Interest Selection
Lei Wang 0198, Dawei Song 0001, Eyad Elyan
ECIR3
2011 Words-of-interest selection based on temporal motion coherence for video retrieval
abstract
The "Bag of Visual Words" (BoW) framework has been widely used in query-by-example video retrieval to model the visual content by a set of quantized local feature descriptors. In this paper, we propose a novel technique to enhance BoW by the selection of Word-of-Interest (WoI) that utilizes the quantified temporal motion coherence of the visual words between the adjacent frames in the query example. Experiments carried out using TRECVID datasets show that our technique improves the retrieval performance of the classical BoW-based approach.
Lei Wang 0198, Dawei Song 0001, Eyad Elyan
SIGIR3