Frans Coenen

dblp:c/FransCoenen · DBLP profile ↗
← Back
147ranked-venue papers
25as first author
21since 2021 · last 2026
0000-0003-1026-6649ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 116 · 20 first-author · 14 since 2021Databases, data management, data science and information retrieval · 68 · 18 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 3 since 2021Security and privacy · 6Human-computer interaction and ubiquitous computing · 5Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Physics-Guided 3D Convolutional Learning for Accurate Springback Error Prediction in Single Point Incremental Forming
Frans Coenen, Mariluz Penalva Oscoz, Ander Martin Rebe, Yang Hai, Anh Nguyen 0003
ICPR (1)2
2025 Attention-SSM Network for Predicting Springback Error in Single Point Incremental Forming
abstract
Predicting springback when employing a Single Point Incremental Forming (SPIF) process is a complex industrial manufacturing task. Current methodologies predominantly depend on Recurrent Neural Networks (RNNs), such as LSTM and GRU, for the processing of sequential 3D point cloud data representations. However, these systems exhibit constrained scalability, inefficiency stemming from sequential processing, and a deficiency in interpretability, rendering them less appropriate for practical manufacturing contexts. We introduce the use of Attention-State-Space Modeling(SSM) Network in this study. This innovative approach utilizes Transformer-based self-attention mechanisms in conjunction with SSM to improve sprinback prediction efficiency, scalability, and explainability. Our model encodes the 3D point sequence through self-attention layers and processes the encoded sequence with the Mamba module, effectively capturing contextual dependencies and essential spatial information essential for forecasting springback failures. Extensive evaluation indicates that our proposed approach attains state-of-the-art performance across diverse grid sizes, providing insights into essential sequence characteristics through visualizations. Our contributions are threefold: (1) the introduction of the Attention-SSM Network for SPIF springback prediction, (2) state-of-the-art performance tested on benchmark datasets, and (3) a comprehensive feature analysis emphasizing model explainability. Our code and models are available at https://github.com/DarrenChen0923/Mamba-back.
Mariluz Penalva Oscoz, Ander Martin Rebe, Yang Hai, Frans Coenen, Anh Nguyen 0003
IJCNN5
2023 Triple-kernel gated attention-based multiple instance learning with contrastive learning for medical image analysis
Huafeng Hu, Ruijie Ye, Jeyan Thiyagalingam, Frans Coenen, Jionglong Su
Appl. Intell.4
2023 From deterministic to stochastic: an interpretable stochastic model-free reinforcement learning framework for portfolio optimization
Zitao Song, Pin Qian, Sifan Song, Frans Coenen, Zhengyong Jiang, Jionglong Su
Appl. Intell.5
2023 Deep ensemble learning for high-dimensional subsurface fluid flow modeling
abstract
The accuracy of Deep Learning (DL) algorithms can be improved by combining several deep learners into an ensemble. This avoids the continuous endeavor required to adjust the architecture of individual networks or the nature of the propagation. This study investigates prediction improvements possible using Deep Ensemble Learning (DEL) to determine four distinct multiscale basis functions in the mixed Generalized Multiscale Finite Element Method (GMsFEM), involving the permeability field as the only input. 376,250 samples were initially generated, filtered down to 367,811 after data pre-processing. A standard Convolutional Neural Network (CNN) named SkiplessCNN and three skip connection-based CNNs named FirstSkipCNN, MidSkipCNN, and DualSkipCNN were developed for the base learners. For each basis function, these four CNNs were combined into an ensemble model using linear regression and ridge regression, separately, as part of the stacking technique. A comparison of the coefficient of determination (R2) and Mean Squared Error (MSE) confirms the effectiveness of all three skip connections in enhancing the performance of the standard CNN, with DualSkip being the most effective among them. Additionally, as evaluated on the testing subset, the combined models meaningfully outperform the individual models for all basis functions. The case that applies linear regression delivers R2 ranging from 0.8456 to 0.9191 and MSE ranging from 0.0092 to 0.0369. The ridge regression case achieves marginally better predictions with R2 ranging from 0.8539 to 0.922, and MSE ranging from 0.009 to 0.0349 because its solution involves more evenly distributed weights.
Abouzar Choubineh, Jie Chen 0029, David A. Wood 0003, Frans Coenen, Fei Ma 0002
Eng. Appl. Artif. Intell.4
2023 Retrieval-based language model adaptation for handwritten Chinese text recognition
Shuying Hu, Qiufeng Wang 0001, Kaizhu Huang, Frans Coenen
Int. J. Document Anal. Recognit.5
2022 Pathology Data Prioritisation: A Study Using Multi-variate Time Series
Girvan Burnside, Frans Coenen
DaWaK3
2022 Electrocardiogram Two-Dimensional Motifs: A Study Directed at Cardio Vascular Disease Classification
Hanadi Aldosari, Frans Coenen, Gregory Yoke Hong Lip, Yalin Zheng
IC3K2
2022 Zero-Shot Text Classification via Knowledge Graph Embedding for Social Media Data
abstract
The idea of “citizen sensing” and “human as sensors” is crucial for social Internet of Things, an integral part of cyber–physical–social systems (CPSSs). Social media data, which can be easily collected from the social world, has become a valuable resource for research in many different disciplines, e.g., crisis/disaster assessment, social event detection, or the recent COVID-19 analysis. Useful information, or knowledge derived from social data, could better serve the public if it could be processed and analyzed in more efficient and reliable ways. Advances in deep neural networks have significantly improved the performance of many social media analysis tasks. However, deep learning models typically require a large amount of labeled data for model training, while most CPSS data is not labeled, making it impractical to build effective learning models using traditional approaches. In addition, the current state-of-the-art, pretrained natural language processing (NLP) models do not make use of existing knowledge graphs, thus often leading to unsatisfactory performance in real-world applications. To address the issues, we propose a new zero-shot learning method which makes effective use of existing knowledge graphs for the classification of very large amounts of social text data. Experiments were performed on a large, real-world tweet data set related to COVID-19, the evaluation results show that the proposed method significantly outperforms six baseline models implemented with state-of-the-art deep learning models for NLP.
Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Frans Coenen
IEEE Internet Things J.4
2022 A Novel 3D Unsupervised Domain Adaptation Framework for Cross-Modality Medical Image Segmentation
abstract
We consider the problem of volumetric (3D) unsupervised domain adaptation (UDA) in cross-modality medical image segmentation, aiming to perform segmentation on the unannotated target domain (e.g. MRI) with the help of labeled source domain (e.g. CT). Previous UDA methods in medical image analysis usually suffer from two challenges: 1) they focus on processing and analyzing data at 2D level only, thus missing semantic information from the depth level; 2) one-to-one mapping is adopted during the style-transfer process, leading to insufficient alignment in the target domain. Different from the existing methods, in our work, we conduct a first of its kind investigation on multi-style image translation for complete image alignment to alleviate the domain shift problem, and also introduce 3D segmentation in domain adaptation tasks to maintain semantic consistency at the depth level. In particular, we develop an unsupervised domain adaptation framework incorporating a novel quartet self-attention module to efficiently enhance relationships between widely separated features in spatial regions on a higher dimension, leading to a substantial improvement in segmentation accuracy in the unlabeled target domain. In two challenging cross-modality tasks, specifically brain structures and multi-organ abdominal segmentation, our model is shown to outperform current state-of-the-art methods by a significant margin, demonstrating its potential as a benchmark resource for the biomedical and health informatics research community.
Zixian Su, Kaizhu Huang, Xi Yang 0008, Jie Sun 0024, Amir Hussain 0001, Frans Coenen
IEEE J. Biomed. Health Informatics7
2021 Motif-based Classification using Enhanced Sub-Sequence-Based Dynamic Time Warping
Mohammed Alshehri, Frans Coenen, Keith Dures
DATA2
2021 Motif Based Feature Vectors: Towards a Homogeneous Data Representation for Cardiovascular Diseases Classification
Hanadi Aldosari, Frans Coenen, Gregory Yoke Hong Lip, Yalin Zheng
DaWaK2
2021 Document Ranking for Curated Document Databases Using BERT and Knowledge Graph Embeddings: Introducing GRAB-Rank
Iqra Muhammad, Danushka Bollegala, Frans Coenen, Carrol Gamble, Anna Kearney, Paula R. Williamson
DaWaK3
2021 Pathology Data Prioritisation: A Study of Using Multi-variate Time Series Without a Ground Truth
Girvan Burnside, Frans Coenen
IC3K3
2021 Capturing Expert Knowledge for Building Enterprise SME Knowledge Graphs
abstract
Whilst Knowledge Graphs (KGs) are increasingly used in business scenarios, the construction of enterprise ontologies and the population of KGs from existing relational data remains a significant challenge. In this paper we report our experience in supporting CSols (an SME operating in the analytical laboratory domain) in transitioning their data from legacy databases to a bespoke KG. We modelled the KG using a streamlined approach based on state of the art ontology engineering methodologies, that addresses the challenges faced by SMEs when transitioning to new technologies: lack of resources to devote to the transition, paucity of comprehensive data governance policies, and resistance within the organisation to accepting new practices and knowledge. Our approach uses a combination of UML diagrams and a controlled language glossary to support stakeholders in reaching consensus during the knowledge capture phase, thus reducing the intervention of the ontology engineer only to cases where no agreement can be found. We present a case study illustrating the generation of the KG from a UML specification of part of the analytical domain and from legacy relational data, and we discuss the benefits and challenges of the approach.
Martin Mansfield, Valentina Tamma, Phil Goddard, Frans Coenen
K-CAP4
2021 Weakly supervised learning of RNA modifications from low-resolution epitranscriptome data
abstract
MOTIVATION: Increasing evidence suggests that post-transcriptional ribonucleic acid (RNA) modifications regulate essential biomolecular functions and are related to the pathogenesis of various diseases. Precise identification of RNA modification sites is essential for understanding the regulatory mechanisms of RNAs. To date, many computational approaches for predicting RNA modifications have been developed, most of which were based on strong supervision enabled by base-resolution epitranscriptome data. However, high-resolution data may not be available. RESULTS: We propose WeakRM, the first weakly supervised learning framework for predicting RNA modifications from low-resolution epitranscriptome datasets, such as those generated from acRIP-seq and hMeRIP-seq. Evaluations on three independent datasets (corresponding to three different RNA modification types and their respective sequencing technologies) demonstrated the effectiveness of our approach in predicting RNA modifications from low-resolution data. WeakRM outperformed state-of-the-art multi-instance learning methods for genomic sequences, such as WSCNN, which was originally designed for transcription factor binding site prediction. Additionally, our approach captured motifs that are consistent with existing knowledge, and visualization of the predicted modification-containing regions unveiled the potentials of detecting RNA modifications with improved resolution. AVAILABILITY IMPLEMENTATION: The source code for the WeakRM algorithm, along with the datasets used, are freely accessible at: https://github.com/daiyun02211/WeakRM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daiyun Huang, Jingjue Wei, Jionglong Su, Frans Coenen, Jia Meng 0001
Bioinform.5
2021 MetaTX: deciphering the distribution of mRNA-related features in the presence of isoform ambiguity, with applications in epitranscriptome analysis
abstract
MOTIVATION: The distribution of biological features strongly indicates their functional relevance. Compared to DNA-related features, deciphering the distribution of mRNA-related features is non-trivial due to the existence of isoform ambiguity and compositional diversity of mRNAs. RESULTS: We propose here a rigorous statistical framework, MetaTX, for deciphering the distribution of mRNA-related features. Through a standardized mRNA model, MetaTX firstly unifies various mRNA transcripts of diverse compositions, and then corrects the isoform ambiguity by incorporating the overall distribution pattern of the features through an EM algorithm. MetaTX was tested on both simulated and real data. Results suggested that MetaTX substantially outperformed existing direct methods on simulated datasets, and that a more informative distribution pattern was produced for all the three datasets tested, which contain N6-Methyladenosine sites generated by different technologies. MetaTX should make a useful tool for studying the distribution and functions of mRNA-related biological features, especially for mRNA modifications such as N6-Methyladenosine. AVAILABILITY AND IMPLEMENTATION: The MetaTX R package is freely available at GitHub: https://github.com/yue-wang-biomath/MetaTX.1.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yue Wang 0060, Kunqi Chen, Frans Coenen, Jionglong Su, Jia Meng 0001
Bioinform.4
2021 Secure third-party data clustering using SecureCL, Φ-data and multi-user order preserving encryption
abstract
Abstract Secure collaborative data clustering using SecureCL is presented. SecureCL is founded on the concept of Φ‐data implemented using Super Secure Chain Distance Matrices and encrypted using Multi‐User Order Preserving Encryption. The advantage offered, unlike comparable systems, is that SecureCL does not require any user participation once the Φ‐data proxy has been encrypted; it does not require recourse to Secure Multi‐Party Computation protocols or ‘secret sharing’ mechanisms. The utility of SecureCL is illustrated using Nearest Neighbour Clustering and Density‐Based Spatial Clustering of Applications with Noise, although it can be applied to any data clustering algorithm that involves distance comparison. The reported experiments demonstrate that SecureCL can produce securely cluster configurations comparable to those produced using standard, non‐encrypted, approaches without entailing any significant computational overhead, thus indicating its suitability in the context of Data Mining as a Service.
Nawal Almutairi, Frans Coenen, Keith Dures
Expert Syst. J. Knowl. Eng.2
2021 Multi-modal generative adversarial networks for traffic event detection in smart cities
Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Suparna De, Frans Coenen
Expert Syst. Appl.5
2021 A survey on presentation attack detection for automatic speaker verification systems: State-of-the-art, taxonomy, issues and future direction
abstract
Abstract The emergence of biometric technology provides enhanced security compared to the traditional identification and authentication techniques that were less efficient and secure. Despite the advantages brought by biometric technology, the existing biometric systems such as Automatic Speaker Verification (ASV) systems are weak against presentation attacks. A presentation attack is a spoofing attack launched to subvert an ASV system to gain access to the system. Though numerous Presentation Attack Detection (PAD) systems were reported in the literature, a systematic survey that describes the current state of research and application is unavailable. This paper presents a systematic analysis of the state-of-the-art voice PAD systems to promote further advancement in this area. The objectives of this paper are two folds: (i) to understand the nature of recent work on PAD systems, and (ii) to identify areas that require additional research. From the survey, a taxonomy of voice PAD and the trend analysis of recent work on PAD systems were built and presented, whereby the recent and relevant articles including articles from Interspeech and ICASSP Conferences, mostly indexed by Scopus, published between 2015 and 2021 were considered. A total of 172 articles were surveyed in this work. The findings of this survey present the limitation of recent works, which include spoof-type dependent PAD. Consequently, the future direction of work on voice PAD for interested researchers is established. The findings of this survey present the limitation of recent works, which include spoof-type dependent PAD. Consequently, the future direction of work on voice PAD for interested researchers is established.
Choon Beng Tan, Mohd. Hanafi Ahmad Hijazi, Norazlina Khamis, Puteri Nor Ellyza binti Nohuddin, Zuraini Zainol, Frans Coenen, Abdullah Gani
Multim. Tools Appl.6
2021 Automated Social Text Annotation With Joint Multilabel Attention Networks
abstract
Automated social text annotation is the task of suggesting a set of tags for shared documents on social media platforms. The automated annotation process can reduce users' cognitive overhead in tagging and improve tag management for better search, browsing, and recommendation of documents. It can be formulated as a multilabel classification problem. We propose a novel deep learning-based method for this problem and design an attention-based neural network with semantic-based regularization, which can mimic users' reading and annotation behavior to formulate better document representation, leveraging the semantic relations among labels. The network separately models the title and the content of each document and injects an explicit, title-guided attention mechanism into each sentence. To exploit the correlation among labels, we propose two semantic-based loss regularizers, i.e., similarity and subsumption, which enforce the output of the network to conform to label semantics. The model with the semantic-based loss regularizers is referred to as the joint multilabel attention network (JMAN). We conducted a comprehensive evaluation study and compared JMAN to the state-of-the-art baseline models, using four large, real-world social media data sets. In terms of F1, JMAN significantly outperformed bidirectional gated recurrent unit (Bi-GRU) relatively by around 12.8%-78.6% and the hierarchical attention network (HAN) by around 3.9%-23.8%. The JMAN model demonstrates advantages in convergence and training speed. Further improvement of performance was observed against latent Dirichlet allocation (LDA) and support vector machine (SVM). When applying the semantic-based loss regularizers, the performance of HAN and Bi-GRU in terms of F1was also boosted. It is also found that dynamic update of the label semantic matrices (JMANd) has the potential to further improve the performance of JMAN but at the cost of substantial memory and warrants further study.
Hang Dong 0002, Wei Wang 0042, Kaizhu Huang, Frans Coenen
IEEE Trans. Neural Networks Learn. Syst.4
2020 Graph Convolution over Multiple Dependency Sub-graphs for Relation Extraction
abstract
We propose in this paper a contextualised graph convolution network over multiple dependency sub-graphs for relation extraction.A novel method to construct multiple sub-graphs using words in shortest dependency path and words linked to entities in the dependency graph is proposed.Graph convolution operation is performed over the resulting multiple sub-graphs to obtain more informative features useful for relation extraction.Our experimental results show that the proposed method achieves superior performance over existing GCN-based models achieving stateof-the-art performance on cross-sentence n-ary relation extraction and SemEval 2010 Task 8 sentence-level relation extraction task.Our model also achieves a comparable performance to the SoTA on the TACRED dataset.
Angrosh Mandya, Danushka Bollegala, Frans Coenen
COLING3
2020 Sustainable Development Goal Relational Modelling: Introducing the SDG-CAP Methodology
Yassir Alharbi, Frans Coenen, Daniel Arribas-Bel
DaWaK2
2020 Do not let the history haunt you: Mitigating Compounding Errors in Conversational Question Answering
abstract
The Conversational Question Answering (CoQA) task involves answering a sequence of inter-related conversational questions about a contextual paragraph. Although existing approaches employ human-written ground-truth answers for answering conversational questions at test time, in a realistic scenario, the CoQA model will not have any access to ground-truth answers for the previous questions, compelling the model to rely upon its own previously predicted answers for answering the subsequent questions. In this paper, we find that compounding errors occur when using previously predicted answers at test time, significantly lowering the performance of CoQA systems. To solve this problem, we propose a sampling strategy that dynamically selects between target answers and model predictions during training, thereby closely simulating the situation at test time. Further, we analyse the severity of this phenomena as a function of the question type, conversation length and domain type.
Angrosh Mandya, James O'Neill, Danushka Bollegala, Frans Coenen
LREC4
2020 Multi-modal Adversarial Training for Crisis-related Data Classification on Social Media
abstract
Social media platforms such as Twitter are increasingly used to collect data of all kinds. During natural disasters, users may post text and image data on social media platforms to report information about infrastructure damage, injured people, cautions and warnings. Effective processing and analysing tweets in real time can help city organisations gain situational awareness of the affected citizens and take timely operations. With the advances in deep learning techniques, recent studies have significantly improved the performance in classifying crisis-related tweets. However, deep learning models are vulnerable to adversarial examples, which may be imperceptible to the human, but can lead to model's misclassification. To process multi-modal data as well as improve the robustness of deep learning models, we propose a multi-modal adversarial training method for crisis-related tweets classification in this paper. The evaluation results clearly demonstrate the advantages of the proposed model in improving the robustness of tweet classification.
Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Suparna De, Frans Coenen
SMARTCOMP5
2020 A Cryptographic Ensemble for secure third party data analysis: Collaborative data clustering without data owner participation
Nawal Almutairi, Frans Coenen, Keith Dures
Data Knowl. Eng.2
2020 Knowledge base enrichment by relation learning from social tagging data
Hang Dong 0002, Wei Wang 0042, Frans Coenen, Kaizhu Huang
Inf. Sci.3
2019 Ontology Learning from Twitter Data
abstract
Copyright © 2019 by SCITEPRESS – Science and Technology Publications, Lda. All rights reserved This paper presents and compares three mechanisms for learning an ontology describing a domain of discoursed as defined in a collection of tweets. The task in part involves the identification of entities and relations in the free text data, which can then be used to produce a set of RDF triples from which an ontology can be generated. The first mechanism is therefore founded on the Stanford CoreNLP Toolkit.; in particular the Named Entity Recognition and Relation Extraction mechanisms that come with this tool kit. The second is founded on the GATE General Architecture for Text Engineering which provides an alternative mechanism for relation extraction from text. Both require a substantial amount of training data. To reduce the training data requirement the third mechanism is founded on the concept of Regular Expressions extracted from a training data “seed set”. Although the third mechanism still requires training data the amount of training data is significantly reduced without adversely affecting the quality of the ontologies generated.
Saad Alajlan, Frans Coenen, Boris Konev, Angrosh Mandya
KEOD2
2019 From Semi-automated to Automated Methods of Ontology Learning from Twitter Data
Saad Alajlan, Frans Coenen, Angrosh Mandya
IC3K2
2019 Automated Bundle Pagination Using Machine Learning
abstract
Coherent division of legal document bundles, whether this is done in the context of court bundles, briefs or some other application, is a time consuming and challenging task. We propose an approach whereby this process can be automated. Two variations are considered. The first addresses the scenario where the topic labelling is pre-defined and adopts a supervised learning approach. The second addresses the scenario where the topic labelling, for whatever reason, is not specified in advance and adopts an unsupervised learning approach. This paper reports on an investigation of both mechanisms using accident claims bundles. The evaluation results indicate that the proposed approaches can be successfully applied to divide legal document bundles.
Alessandro Torrisi, Robert Bevan, Katie Atkinson, Danushka Bollegala, Frans Coenen
ICAIL5
2019 Combining Textual and Visual Information for Typed and Handwritten Text Separation in Legal Documents
Alessandro Torrisi, Robert Bevan, Katie Atkinson, Danushka Bollegala, Frans Coenen
JURIX5
2018 A Re-evaluation of Intrusion Detection Accuracy: Alternative Evaluation Strategy
abstract
This work tries to evaluate the existing approaches used to benchmark the performance of machine learning models applied to network-based intrusion detection systems (NIDS). First, we demonstrate that we can reach a very high accuracy with most of the traditional machine learning and deep learning models by using the existing performance evaluation strategy. It just requires the right hyperparameter tuning to outperform the existing reported accuracy results in deep learning models. We further question the value of the existing evaluation methods in which the same datasets are used for training and testing the models. We are proposing the use of an alternative strategy that aims to evaluate the practicality and the performance of the models and datasets as well. In this approach, different datasets with compatible sets of features are used for training and testing. When we evaluate the models that we created with the proposed strategy, we demonstrate that the performance is very bad. Thus, models have no practical usage, and it performs based on a pure randomness. This research is important for security-based machine learning applications to re-think about the datasets and the model's quality.
Said Al-Riyami, Frans Coenen, Alexei Lisitsa 0001
CCS2
2018 Traversal-aware Encryption Adjustment for Graph Databases
Nahla Aburawi, Frans Coenen, Alexei Lisitsa 0001
DATA2
2018 Third Party Data Clustering Over Encrypted Data Without Data Owner Participation: Introducing the Encrypted Distance Matrix
Nawal Almutairi, Frans Coenen, Keith Dures
DaWaK2
2018 Secure Outsourced kNN Data Classification over Encrypted Data Using Secure Chain Distance Matrices
Nawal Almutairi, Frans Coenen, Keith Dures
IC3K2
2018 Querying Encrypted Graph Databases
abstract
Copyright © 2018 by SCITEPRESS – Science and Technology Publications, Lda. All rights reserved. We present an approach to execution of queries on encrypted graph databases. The approach is inspired by CryptDB system for relational DBs (R. A. Popa et al). Before processing a graph query is translated into encrypted form which then executed on a server without decrypting any data; the encrypted results are sent back to a client where they are finally decrypted. In this way data privacy is protected at the server side. We present the design of the system and empirical data obtained by experimentation with a prototype, implemented for Neo4j graph DBMS and Cypher query language, utilizing Java API. We report the efficiency of query execution for various types of queries on encrypted and non-encrypted Neo4j graph databases.
Nahla Aburawi, Alexei Lisitsa 0001, Frans Coenen
ICISSP3
2018 Spectral Analysis of Keystroke Streams: Towards Effective Real-time Continuous User Authentication
abstract
Copyright © 2018 by SCITEPRESS – Science and Technology Publications, Lda. All rights reserved. Continuous authentication using keystroke dynamics is significant for applications where continuous monitoring of a user’s identity is desirable, for example in the context of the online assessments and examinations frequently encountered in eLearning environments. In this paper, a novel approach to realtime keystroke continuous authentication is proposed that is founded on a sinusoidal signal based approach that takes into consideration the sequencing of keystrokes. Three alternative time series representations are considered and compared: Keystroke Time Series (KTS), Discrete Fourier Transform (DFT) and Discrete Wavelet Transform (DWT). The proposed process is fully described and analysed using three keystroke dynamics datasets. The evaluation also includes a comparison with the established Feature Vector Representation (FVR) approach. The reported evaluation demonstrates that the proposed method, coupled with the DWT representation, outperforms other approaches to keystroke continuous authentication with a best overall accuracy of 98.24%; a clear indicator that the proposed keystroke continuous authentication using time series analysis has significant potential.
Abdullah Alshehri 0001, Frans Coenen, Danushka Bollegala
ICISSP2
2018 Efficient and Effective Case Reject-Accept Filtering: A Study Using Machine Learning
abstract
The decision whether to accept or reject a new case is a well established task undertaken in legal work. This task frequently necessitates domain knowledge and is consequently resource expensive. In this paper it is proposed that early rejection/acceptance of at least a proportion of new cases can be effectively achieved without requiring significant human intervention. The paper proposes, and evaluates, five different AI techniques whereby early case reject-accept can be achieved. The results suggest it is possible for at least a proportion of cases to be processed in this way.
Robert Bevan, Alessandro Torrisi, Katie Atkinson, Danushka Bollegala, Frans Coenen
JURIX5
2018 A Dataset for Inter-Sentence Relation Extraction using Distant Supervision
Angrosh Mandya, Danushka Bollegala, Frans Coenen, Katie Atkinson
LREC3
2018 Multi-dimensional Banded Pattern Mining
Fatimah Binta Abdullahi, Frans Coenen
PKAW2
2018 Segmenting Sound Waves to Support Phonocardiogram Analysis: The PCGseg Approach
Hajar Al-Hujailan, Frans Coenen, Jo Dukes-McEwan, Jeyan Thiyagalingam
PRICAI2
2018 Learning Relations from Social Tagging Data
Hang Dong 0002, Wei Wang 0042, Frans Coenen
PRICAI (1)3
2017 Classifier-Based Pattern Selection Approach for Relation Instance Extraction
Angrosh Mandya, Danushka Bollegala, Frans Coenen, Katie Atkinson
CICLing (1)3
2017 K-Means Clustering Using Homomorphic Encryption and an Updatable Distance Matrix: Secure Third Party Data Clustering with Limited Data Owner Interaction
Nawal Almutairi, Frans Coenen, Keith Dures
DaWaK2
2017 Behavioural Biometric Continuous User Authentication Using Multivariate Keystroke Streams in the Spectral Domain
Abdullah Alshehri 0001, Frans Coenen, Danushka Bollegala
IC3K2
2017 CLIEL: context-based information extraction from commercial law documents
abstract
The effectiveness of document Information Extraction (IE) is greatly affected by the structure and layout of the documents being considered. In the case of legal documents relating to commercial law, an additional challenge is the many different and varied formats, structures and layouts used. In this paper, we present work on a flexible and scalable IE environment, the CLIEL (Commercial Law Information Extraction based on Layout) environment, for application to commercial law documentation that allows layout rules to be derived and then utilised to support IE. The proposed CLIEL environment operates using NLP (Natural Language Processing) techniques, JAPE (Java Annotation Patterns Engine) rules and some GATE (General Architecture for Text Engineering) modules. The system is fully described and evaluated using a commercial law document corpus. The results demonstrate that considering the layout is beneficial for extracting data point instances from legal document collections.
Matias Garcia-Constantino, Katie Atkinson, Danushka Bollegala, Karl Chapman, Frans Coenen, Claire Roberts, Katy Robson
ICAIL5
2017 Attribute Permutation Steganography Detection using Attribute Position Changes Count
Iman Sedeeq, Frans Coenen, Alexei Lisitsa 0001
ICISSP2
2017 A Prediction Model Based Approach to Open Space Steganography Detection in HTML Webpages
Iman Sedeeq, Frans Coenen, Alexei Lisitsa 0001
IWDW2
2017 TSP: Learning Task-Specific Pivots for Unsupervised Domain Adaptation
Xia Cui 0001, Frans Coenen, Danushka Bollegala
ECML/PKDD (2)2
2017 FCNN: Fourier Convolutional Neural Networks
Harry Pratt, Bryan M. Williams 0001, Frans Coenen, Yalin Zheng
ECML/PKDD (1)3
2017 Mining the information architecture of the WWW using automated website boundary detection
abstract
The world wide web has two main forms of architecture, the first is that which is explicitly encoded into web pages, and the second is that which is implied by the web content, particularly pertaining to look and feel. The latter is exemplified by the concept of a website, a concept that is only loosely defined, although users intuitively understand it. The Website Boundary Detection (WBD) problem is concerned with the task of identifying the complete collection of web pages/resources that are contained within a single website. Whatever the case, the concept of a website is used with respect to a number of application domains including; website archiving, spam detection, and www analysis. In the context of such applications it is beneficial if a website can be automatically identified. This is usually done by identifying a website of interest in terms of its boundary, the so called WBD problem. In this paper seven WBD techniques are proposed and compared, four statistical techniques where the web data to be used is obtained apriori, and three dynamic techniques where the data to be used is obtained as the process progresses. All seven techniques are presented in detail and evaluated.
Ayesh Alshukri, Frans Coenen
Web Intell.2
2016 Keyboard Usage Authentication Using Time Series Analysis
Abdullah Alshehri 0001, Frans Coenen, Danushka Bollegala
DaWaK2
2016 Image Representation for Image Mining: A Study Focusing on Mining Satellite Images for Census Data Collection
Frans Coenen, Kwankamon Dittakan
IC3K1
2016 A Statistical Approach to the Detection of HTML Attribute Permutation Steganography
Iman Sedeeq, Frans Coenen, Alexei Lisitsa 0001
ICISSP2
2016 Extracting Movement Patterns from Video Data to Drive Multi-Agent Based Simulations
Muhammad Tufail, Frans Coenen, Tintin Mu
MABS2
2016 Early Detection of Osteoarthritis Using Local Binary Patterns: A Study Directed at Human Joint Imagery
Kwankamon Dittakan, Frans Coenen
PRICAI2
2016 3-D Volume of Interest Based Image Classification
Akadej Udomchaiporn, Frans Coenen, Marta García-Fiñana, Vanessa Sluming
PRICAI2
2016 Driving posture recognition by convolutional neural networks
abstract
Driver fatigue and inattention have long been recognised as the main contributing factors in traffic accidents. This study presents a novel system which applies convolutional neural network (CNN) to automatically learn and predict pre‐defined driving postures. The main idea is to monitor driver hand position with discriminative information extracted to predict safe/unsafe driving posture. In comparison to previous approaches, CNNs can automatically learn discriminative features directly from raw images. In the authors' works, a CNN model was first pre‐trained by an unsupervised feature learning method called sparse filtering, and subsequently fine‐tuned with classification. The approach was verified using the Southeast University driving posture dataset, which comprised of video clips covering four driving postures, including normal driving, responding to a cell phone call, eating, and smoking. Compared with other popular approaches with different image descriptors and classification methods, the authors' scheme achieves the best performance with an overall accuracy of 99.78%. To evaluate the effectiveness and generalisation performance in more realistic conditions, the method was further tested using other two specially designed datasets which takes into account of the poor illuminations and different road conditions, achieving an overall accuracy of 99.3 and 95.77%, respectively.
Chao Yan 0003, Frans Coenen
IET Comput. Vis.2
2016 Video-Based Classification of Driving Behavior Using a Hierarchical Classification System with Multiple Features
abstract
Driver fatigue and inattention have long been recognized as one of the main contributing factors in traffic accidents. Therefore, the development of intelligent driver assistance systems, which provides automatic monitoring of driver's vigilance, is an urgent and challenging task. This paper presents a novel system for video-based driving behavior recognition. The fundamental idea is to monitor driver's hand movements and to use these as predictors for safe/unsafe driving behavior. In comparison to previous work, the proposed method utilizes hierarchical classification and treats driving behavior in terms of a spatio-temporal reference framework as opposed to a static image. The approach was verified using the Southeast University Driving-Posture Dataset, a dataset comprised of video clips covering aspects of driving such as: normal driving, responding to a cell phone call, eating and smoking. After pre-processing for illumination variations and motion sequence segmentation, eight classes of behavior were identified. The overall prediction accuracy obtained using the proposed approach was [Formula: see text] when using a hierarchical classification approach. The proposed approach was able to clearly identify two dangerous driving behaviors, Responding to a cellphone call and Eating, with recognition rates of 92.39% and 92.29% respectively.
Chao Yan 0003, Frans Coenen, Yong Yue 0001, Xiaosong Yang
Int. J. Pattern Recognit. Artif. Intell.2
2016 Face Occlusion Detection Using Deep Convolutional Neural Networks
abstract
With the rise of crimes associated with Automated Teller Machines (ATMs), security reinforcement by surveillance techniques has been a hot topic on the security agenda. As a result, cameras are frequently installed with ATMs, so as to capture the facial images of users. The main objective is to support follow-up criminal investigations in the event of an incident. However, in the case of miss-use, the user’s face is often occluded. Therefore, face occlusion detection has become very important to prevent crimes connected with ATM usage. Traditional approaches to solving the problem typically comprise a succession of steps: localization, segmentation, feature extraction and recognition. This paper proposes an end-to-end facial occlusion detection framework, which is robust and effective by combining region proposal algorithm and Convolutional Neural Networks (CNN). The framework utilizes a coarse-to-fine strategy, which consists of two CNNs. The first CNN detects the head element within an upper body image while the second distinguishes which facial part is occluded from the head image. In comparison with previous approaches, the usage of CNN is optimal from a system point of view as the design is based on the end-to-end principle and the model operates directly on image pixels. For evaluation purposes, a face occlusion database consisting of over 50[Formula: see text]000 images, with annotated facial parts, was used. Experimental results revealed that the proposed framework is very effective. Using the bespoke face occlusion dataset, Aleix and Robert (AR) face dataset and the Labeled Face in the Wild (LFW) database, we achieved over 85.61%, 97.58% and 100% accuracies for head detection when the Intersection over Union-section (IoU) is larger than 0.5, and 94.55%, 98.58% and 95.41% accuracies for occlusion discrimination, respectively.
Yizhang Xia, Frans Coenen
Int. J. Pattern Recognit. Artif. Intell.3
2015 Finding banded patterns in big data using sampling
abstract
A mechanism for identifying bandings in large "zero-one" N-dimensional data sets, using a sampling technique, is presented. The challenge of identifying bandings in data is the large number of potential permutations that need to be considered. To circumvent this a banding score mechanism is proposed that avoids the need to consider large numbers of permutations. This has been incorporated into a proposed banded pattern mining algorithm, the Exact ND Banded Pattern Mining (END BPM) algorithm. Although this operates well on reasonably sized datasets, there is still a challenge with respect to large N-dimensional data sets that cannot be held in primary storage. To this end a sampling technique is also proposed. The approach is fully described and evaluated using the GB cattle movement database, a "real life" database that records all movements of cattle in GB.
Fatimah Binta Abdullahi, Frans Coenen, Russell Martin
IEEE BigData2
2015 Finding Banded Patterns in Data: The Banded Pattern Mining Algorithm
Fatimah Binta Abdullahi, Frans Coenen, Russell Martin
DaWaK2
2015 Data Stream Mining with Limited Validation Opportunity: Towards Instrument Failure Prediction
Katie Atkinson, Frans Coenen, Phil Goddard, Terry R. Payne, Luke Riley
DaWaK2
2015 Scalable distributed collaborative tracking and mapping with Micro Aerial Vehicles
abstract
This paper describes work on a distributed framework for collaborative multi-robot localisation and mapping with large teams of Micro Aerial Vehicles (MAVs). We demonstrate the benefits of running both image capture and frame-to-frame tracking on the same device while offloading the more computationally intensive aspects of map creation and optimization to an off-board computer. We show no impact on the accuracy of pose estimates of this distributed approach and indeed demonstrate a robustness to delay that improves localisation performance. The bandwidth requirements of our system are much lower than similar systems which enables us to accommodate larger teams of MAVs. In the results section we demonstrate the performance of our system in both simulated and real-world environments.
Boris Konev, Frans Coenen
IROS3
2015 Towards an Intuitionistic Fuzzy Agglomerative Hierarchical Clustering Algorithm for Music Recommendation in Folksonomy
abstract
Folksonomy, a system for social tagging or collaborative tagging, is popular in Semantic Web research. Folksonomy is applied to items, such as music pieces, which their personalized tags can be annotated by users. Recommendation systems can use these tags to produce meaningful information. Clustering methods, such as the Agglomerative Hierarchical Clustering (AHC) method, can be applied in the context of recommendation system. This paper proposes the Intuitionistic Fuzzy Agglomerative Hierarchical Clustering (IFAHC) algorithm for recommendation using social tagging. The Intuitionistic Fuzzy Set (IFS) concept is used to represent tag values which are vague and uncertain. IFAHC can cluster items represented by using IFS into different groups. The application of IFAHC to music recommendation is used to demonstrate the usability of the proposed method.
Chun Guan, Kevin Kam Fung Yuen, Frans Coenen
SMC3
2015 Predicting "springback" using 3D surface representation techniques: A case study in sheet metal forming
Subhieh El-Salhi, Frans Coenen, Clare Dixon, M. Sulaiman Khan
Expert Syst. Appl.2
2014 A Scalable Algorithm for Banded Pattern Mining in Multi-dimensional Zero-One Data
Fatimah Binta Abdullahi, Frans Coenen, Russell Martin
DaWaK2
2014 3-D MRI Brain Scan Classification Using A Point Series Based Representation
Akadej Udomchaiporn, Frans Coenen, Marta García-Fiñana, Vanessa Sluming
DaWaK2
2014 Content-Based Readability Assessment: A Study Using A Syllabic Alphabetic Language (Thai)
Nattapong Tongtep, Frans Coenen, Thanaruk Theeramunkong
PRICAI2
2014 Trend mining in social networks: from trend identification to visualization
abstract
Abstract A four‐stage social network trend mining framework, the Identification, Grouping, Clustering and Visualization framework, is described. The framework extracts trends from social network data and then applies a sequence of techniques (‘tools’) to this data to facilitate interpretation of the identified trends. Of particular note is the visualization of trend migrations (changes) that feature within time‐stamped network data. The framework is illustrated using a sequence of four social networks extracted from the Cattle Tracing System in operation in Great Britain, although it could equally well be applied to other forms of temporal data. The presented analysis of the Identification, Grouping, Clustering and Visualization framework indicates advantages, with respect to network trend mining, that can be gained, especially when the framework is applied to real data.
Puteri Nor Ellyza binti Nohuddin, Wataru Sunayama, Rob M. Christley, Frans Coenen, Christian Setzkorn
Expert Syst. J. Knowl. Eng.4
2013 Vertex Unique Labelled Subgraph Mining for Vertex Label Classification
Wen Yu 0003, Frans Coenen, Michele Zito 0001, Subhieh El-Salhi
ADMA (1)2
2013 Hierarchical Classification for Solving Multi-class Problems: A New Approach Using Naive Bayesian Classification
Esra'a Alshdaifat, Frans Coenen, Keith Dures
ADMA (1)2
2013 A Comparative Study of Three Image Representations for Population Estimation Mining Using Remote Sensing Imagery
Kwankamon Dittakan, Frans Coenen, Rob M. Christley, Maya Wardeh
ADMA (1)2
2013 Predicting Features in Complex 3D Surfaces Using a Point Series Representation: A Case Study in Sheet Metal Forming
Subhieh El-Salhi, Frans Coenen, Clare Dixon, M. Sulaiman Khan
ADMA (1)2
2013 Generating Domain-Specific Sentiment Lexicons for Opinion Mining
Zaher Salah, Frans Coenen, Davide Grossi
ADMA (1)2
2013 3-D MRI Brain Scan Feature Classification Using an Oct-Tree Representation
Akadej Udomchaiporn, Frans Coenen, Marta García-Fiñana, Vanessa Sluming
ADMA (1)2
2013 Classification of volumetric retinal images using overlapping decomposition and tree analysis
abstract
Methods for the classification of volumetric three-dimensional (3D) volumes (images) play an important role in the context of medical applications. In this paper, a dedicated tree based 3D representation is proposed that serves to directly capture 3D image features in such a way that classification techniques can be applied. More specifically an Overlapping Hierarchical Decomposition (OHD) technique is presented to generate a tree representation of a given 3D volume. The OHD method recursively decomposes a given 3D volume into sub-volumes forming a tree. Once the tree has been generated, a frequent sub-graph mining algorithm is applied to mine the tree representation so as to generate sub-graphs. These sub-graphs are then used to define a feature space from which feature vectors representing 3D images (one per 3D volume) can be extracted and fed into a classifier generator. To demonstrate the applicability of the proposed method a 3D Optical Coherence Tomography (OCT) retinal image screening application is considered directed at the identification of Age-related Macular Degeneration (AMD). The results show a promising performance with a best Area Under the receiver operating Curve (AUC) value of 98.7%.
Abdulrahman Albarrak, Frans Coenen, Yalin Zheng
CBMS2
2013 Population Estimation Mining Using Satellite Imagery
Kwankamon Dittakan, Frans Coenen, Rob M. Christley, Maya Wardeh
DaWaK2
2013 Minimal Vertex Unique Labelled Subgraph Mining
Wen Yu 0003, Frans Coenen, Michele Zito 0001, Subhieh El-Salhi
DaWaK2
2013 Extracting debate graphs from parliamentary transcripts: a study directed at UK house of commons debates
abstract
The paper proposes a framework---the Debate Graph Extraction (DGE) framework---for extracting debate graphs from transcripts of political debates. The idea is to represent the structure of a debate as a graph with speakers as nodes and "exchanges" as links. Links between nodes are established according to the semantic similarity between the speeches and indicate an alignment of content between them. Nodes are labelled according to the "attitude" (sentiment) of the speakers, positive or negative, using a lexicon based technique founded on SentiWordNet. The attitude of the speakers is then used to label the graph links as being either "supporting" or "opposing". If both speakers have the same attitude (both negative or both positive) the link is labelled as being supporting; otherwise the link is labelled as being opposing. The resulting graphs capture the abstract representation of a debate as two opposing fractions exchanging arguments on related content.
Zaher Salah, Frans Coenen, Davide Grossi
ICAIL2
2013 An Efficient Algorithm for Mining Erasable Itemsets Using the Difference of NC-Sets
abstract
This paper proposes an improved version of the MERIT algorithm, deMERIT+, for mining all "erasable item sets". We first establish an algorithm MERIT+, a revised version of MERIT, which is then used as the foundation for deMERIT+. The proposed algorithm uses: a weight index, a hash table and the "difference" of Node Code Sets (dNC-Sets) to improve the mining time. A theorem is derived first to show that dNC-Sets can be used for mining erasable item sets. The experimental results show that deMERIT+ is more effective than MERIT+ in terms of the runtime.
Tuong Le, Bay Vo, Frans Coenen
SMC3
2013 A Hybrid Approach for Mining Frequent Itemsets
abstract
Frequent item set mining is a fundamental element with respect to many data mining problems. Recently, the PrePost algorithm has been proposed, a new algorithm for mining frequent item sets based on the idea of N-lists. PrePost in most cases outperforms other current state-of-the-art algorithms. In this paper, we present an improved version of PrePost that uses a hash table to enhance the process of creating the N-lists associated with 1-itemsets and an improved N-list intersection algorithm. Furthermore, two new theorems are proposed for determining the "subsume index" of frequent 1-itemsets based on the N-list concept. The experimental results show that the performance of the proposed algorithm improves on that of PrePost.
Bay Vo, Frans Coenen, Tuong Le, Tzung-Pei Hong
SMC2
2013 A new method for mining Frequent Weighted Itemsets based on WIT-trees
Bay Vo, Frans Coenen, Bac Le
Expert Syst. Appl.2
2013 Breast cancer diagnosis from biopsy images with highly reliable random subspace classifier ensembles
Yungang Zhang, Frans Coenen, Wenjin Lu
Mach. Vis. Appl.3
2012 Highly reliable breast cancer diagnosis with cascaded ensemble classifiers
abstract
Accuracy and reliability are two important issues in computer assisted breast cancer diagnosis. In this paper, a new cascade Random Subspace ensembles scheme with reject options is proposed for automatic breast cancer diagnosis. The diagnosis system is built as a serial fusion of two different Random Subspace classifier ensembles with rejection options to enhance the classification reliability. The first ensemble consists of a set of Support Vector Machine (SVM) classifiers that converts the original K-class classification problem into a number of K 2-class problems. The second ensemble consists of a Multi-Layer Perceptron (MLP) ensemble, that focuses on the rejected samples from the first ensemble. For both of the ensembles, the reject option is implemented by relating the consensus degree from majority voting to a confidence measure, and abstaining to classify ambiguous samples if the consensus degree is lower than some threshold. Using a microscopic breast biopsy image dataset from Israel Institute of Technology and benchmark datasets from UCI, promising results are obtained using the proposed system.
Yungang Zhang, Frans Coenen, Wenjin Lu
IJCNN3
2012 Identification and Visualisation of Pattern Migrations in Big Network Data
Puteri Nor Ellyza binti Nohuddin, Frans Coenen, Rob M. Christley, Wataru Sunayama
PRICAI2
2012 A framework for Multi-Agent Based Clustering
Santhana Chaimontree, Katie Atkinson, Frans Coenen
Auton. Agents Multi Agent Syst.3
2012 Multi-agent based classification using argumentation from experience
Maya Wardeh, Frans Coenen, Trevor J. M. Bench-Capon
Auton. Agents Multi Agent Syst.2
2012 PISA: A framework for multiagent classification using argumentation
Maya Wardeh, Frans Coenen, Trevor J. M. Bench-Capon
Data Knowl. Eng.2
2012 Data mining techniques for the screening of age-related macular degeneration
Mohd. Hanafi Ahmad Hijazi, Frans Coenen, Yalin Zheng
Knowl. Based Syst.2
2012 Finding "interesting" trends in social networks using frequent pattern mining and self organizing maps
Puteri Nor Ellyza binti Nohuddin, Frans Coenen, Rob M. Christley, Christian Setzkorn, Shane Williams
Knowl. Based Syst.2
2011 Time Series Case Based Reasoning for Image Categorisation
Ashraf Elsayed, Mohd. Hanafi Ahmad Hijazi, Frans Coenen, Marta García-Fiñana, Vanessa Sluming, Yalin Zheng
ICCBR3
2011 Multi-agent Based Classification Using Argumentation from Experience
Maya Wardeh, Frans Coenen, Trevor J. M. Bench-Capon, Adam Z. Wyner
PAKDD (2)2
2011 Image Classification for Age-related Macular Degeneration Screening Using Hierarchical Image Decompositions and Graph Mining
Mohd. Hanafi Ahmad Hijazi, Chuntao Jiang, Frans Coenen, Yalin Zheng
ECML/PKDD (2)3
2010 Best Clustering Configuration Metrics: Towards Multiagent Based Clustering
Santhana Chaimontree, Katie Atkinson, Frans Coenen
ADMA (1)3
2010 Classification Inductive Rule Learning with Negated Features
Stephanie Chua, Frans Coenen, Grant Malcolm
ADMA (1)2
2010 Finding Frequent Subgraphs in Longitudinal Social Network Data Using a Weighted Graph Mining Approach
Chuntao Jiang, Frans Coenen, Michele Zito 0001
ADMA (1)2
2010 Frequent Pattern Trend Analysis in Social Networks
Puteri Nor Ellyza binti Nohuddin, Rob M. Christley, Frans Coenen, Christian Setzkorn, Shane Williams
ADMA (1)3
2010 Arguing in Groups
abstract
We have previously introduced the notion of arguing from experience, whereby agents debate a classification problem using arguments based on association rules mined “on the fly” from their individual datasets. In this paper we extend PISA, which allows for n agents to argue about cases which have n possible classifications. By allowing any number of agents to participate all the agents supporting a given classification can form a collaborative group for the purposes of the dialogue. We describe how the system is organised, give an example, and report results which suggest that allowing groups in this way has a beneficial effect on the quality of the result.
Maya Wardeh, Frans Coenen, Trevor J. M. Bench-Capon
COMMA2
2010 Region of Interest Based Image Categorization
Ashraf Elsayed, Frans Coenen, Marta García-Fiñana, Vanessa Sluming
DaWak2
2010 Frequent Sub-graph Mining on Edge Weighted Graphs
Chuntao Jiang, Frans Coenen, Michele Zito 0001
DaWak2
2010 Region of Interest Based Image Classification using time series analysis
abstract
An approach to Region Of Interest Based Image Classification (ROIBIC), based on a time series analysis approach, is described. The focus of the approach is the classification of MRI brain scan data according to the nature of the corpus callosum (a feature within such scans), however the approach also has general applicability. The advocated approach combines a number of image processing techniques combined with time series analysis, specifically dynamic time warping. Of note is the mechanism used to generate the desired time series. The application of the time series based ROIBIC demonstrates that the proposed approach performs both efficiently and effectively, obtaining a classification accuracy of over 98% in the case of the given application. Comparisons are also presented with a graph based ROIBIC approach.
Ashraf Elsayed, Frans Coenen, Marta García-Fiñana, Vanessa Sluming
IJCNN2
2010 Retinal image classification using a histogram based approach
abstract
An approach to classifying retinal images using a histogram based representation is described. More specifically, a two stage Case Based Reasoning (CBR) approach is proposed, to be applied to histogram represented retina images to identify Age-related Macular Degeneration (AMD). To measure the similarity between histograms, a time series analysis technique, Dynamic Time Warping (DTW), is employed. The advocated approach utilises two “case bases” for the classification process. The first case base consists of green and saturation histograms with retinal blood vessels removed. The second case base comprises the same histograms, but with the Optic Disc (OD) removed as well. The reported experiments demonstrate that the proposed two stage classification process outperforms the single stage classification process with respect to a number of evaluation metrics: specificity, sensitivity and accuracy.
Mohd. Hanafi Ahmad Hijazi, Frans Coenen, Yalin Zheng
IJCNN2
2010 Corpus callosum MR image classification
Ashraf Elsayed, Frans Coenen, Chuntao Jiang, Marta García-Fiñana, Vanessa Sluming
Knowl. Based Syst.2
2010 Text classification using graph mining-based feature extraction
Chuntao Jiang, Frans Coenen, Robert Sanderson, Michele Zito 0001
Knowl. Based Syst.2
2010 A sliding windows based dual support framework for discovering emerging trends from temporal data
M. Sulaiman Khan, Frans Coenen, David C. Reid, R. Patel, L. Archer
Knowl. Based Syst.2
2009 A Hybrid Statistical Data Pre-processing Approach for Language-Independent Text Classification
Yanbo J. Wang, Frans Coenen, Robert Sanderson
ADMA2
2009 Arguing from Experience to Classifying Noisy Data
Maya Wardeh, Frans Coenen, Trevor J. M. Bench-Capon
DaWaK2
2009 EMADS: An extendible multi-agent data miner
Kamal Ali Albashiri, Frans Coenen, Paul H. Leng
Knowl. Based Syst.2
2009 Improved methods for extracting frequent itemsets from interim-support trees
abstract
Abstract Mining association rules in relational databases is a significant computational task with lots of applications. A fundamental ingredient of this task is the discovery of sets of attributes (itemsets) whose frequency in the data exceeds some threshold value. In this paper we describe two algorithms for completing the calculation of frequent sets using a tree structure for storing partial supports, called interim‐support (IS) tree. The first of our algorithms (T‐Tree‐First (TTF)) uses a novel tree pruning technique, based on the notion of (fixed‐prefix) potential inclusion, which is specially designed for trees that are implemented using only two pointers per node. This allows to implement the IS tree in a space‐efficient manner. The second algorithm (P‐Tree‐First (PTF)) explores the idea of storing the frequent itemsets in a second tree structure, called the total support tree (T‐tree); the main innovation lies in the use of multiple pointers per node, which provides rapid access to the nodes of the T‐tree and makes it possible to design a new, usually faster, method for updating them. Experimental comparison shows that these techniques result in considerable speedup for both algorithms compared with earlier approaches that also use IS trees (Principles of Data Mining and Knowledge Discovery, Proceedings of the 5th European Conference, PKDD, 2001, Freiburg, September 2001 (Lecture Notes in Artificial Intelligence, vol. 2168). Springer: Berlin, Heidelberg, 54–66; Journal of Knowledge‐Based Syst. 2000; 13:141–149). Further comparison between the two new algorithms, shows that the PTF is generally faster on instances with a large number of frequent itemsets, provided that they are relatively short, whereas TTF is more appropriate whenever there exist few or quite long frequent itemsets; in addition, TTF behaves well on instances in which the densities of the items of the database have a high variance. Copyright © 2008 John Wiley & Sons, Ltd.
Frans Coenen, Paul H. Leng, Aris Pagourtzis, Wojciech Rytter, Dora Souliou
Softw. Pract. Exp.1
2008 Arguments from Experience: The PADUA Protocol
Maya Wardeh, Trevor J. M. Bench-Capon, Frans Coenen
COMMA3
2008 Document-Base Extraction for Single-Label Text Classification
Yanbo J. Wang, Robert Sanderson, Frans Coenen, Paul H. Leng
DaWaK3
2008 Argument Based Moderation of Benefit Assessment
abstract
Error rates in the assessment of routine claims for welfare benefits have been found to be very high in Netherlands, USA and UK. This is a significant problem both in terms of quality of service and financial loss through over payments. These errors also present challenges for machine learning programs using the data. In this paper we propose a way of addressing this problem by using a process of moderation, in which agents argue about the classification on the basis of data from distinct groups of assessors. Our agents employ an argument based dialogue protocol (PADUA) in which the agents produce arguments directly from a database of cases, with each agent having their own separate database. We describe the protocol and report encouraging results from a series of experiments comparing PADUA with other classifiers, and assessing the effectiveness of the moderation process.
Maya Wardeh, Trevor J. M. Bench-Capon, Frans Coenen
JURIX3
2007 PADUA Protocol: Strategies and Tactics
Maya Wardeh, Trevor J. M. Bench-Capon, Frans Coenen
ECSQARU3
2007 The effect of threshold values on association rule based classification accuracy
Frans Coenen, Paul H. Leng
Data Knowl. Eng.1
2006 Tree-based partitioning of date for association rule mining
Shakil Ahmed 0002, Frans Coenen, Paul H. Leng
Knowl. Inf. Syst.2
2006 The Knowledge Bazaar
Brian Craker, Frans Coenen
Knowl. Based Syst.2
2005 Obtaining Best Parameter Values for Accurate Classification
abstract
In this paper we examine the effect that the choice of support and confidence thresholds has on the accuracy of classifiers obtained by classification association rule mining. We show that accuracy can almost always be improved by a suitable choice of threshold values, and we describe a method for finding the best values. We present results that demonstrate this approach can obtain higher accuracy without the need for coverage analysis of the training data.
Frans Coenen, Paul H. Leng
ICDM1
2005 Threshold Tuning for Improved Classification Association Rule Mining
Frans Coenen, Paul H. Leng, Lu Zhang 0023
PAKDD1
2004 A Tree Partitioning Method for Memory Management in Association Rule Mining
Shakil Ahmed 0002, Frans Coenen, Paul H. Leng
DaWaK2
2004 An Evaluation of Approaches to Classification Rule Selection
abstract
In this paper a number of classification rule evaluation measures are considered. In particular the authors review the use of a variety of selection techniques used to order classification rules contained in a classifier, and a number of mechanisms used to classify unseen data. The authors demonstrate that rule ordering founded on the size of antecedent works well given certain conditions.
Frans Coenen, Paul H. Leng
ICDM1
2004 Tree Structures for Mining Association Rules
Frans Coenen, Graham Goulbourne, Paul H. Leng
Data Min. Knowl. Discov.1
2004 Data Structure for Association Rule Mining: T-Trees and P-Trees
abstract
Two new structures for association rule mining (ARM), the T-tree, and the P-tree, together with associated algorithms, are described. The authors demonstrate that the structures and algorithms offer significant advantages in terms of storage and execution time.
Frans Coenen, Paul H. Leng, Shakil Ahmed 0002
IEEE Trans. Knowl. Data Eng.1
2003 T-Trees, Vertical Partitioning and Distributed Association Rule Mining
abstract
We consider a technique (DATA-VP) for distributed (and parallel) association rule mining that makes use of a vertical partitioning technique to distribute the input data, amongst processors. The proposed vertical partitioning is facilitated by a novel compressed set enumeration tree data structure (the T-tree), and an associated mining algorithm (Apriori-T), that allows for computationally effective distributed/parallel ARM when compared with existing approaches.
Frans Coenen, Paul H. Leng, Shakil Ahmed 0002
ICDM1
2003 Setting attribute weights for k-NN based binary classification via quadratic programming
Lu Zhang 0023, Frans Coenen, Paul H. Leng
Intell. Data Anal.2
2002 An Attribute Weight Setting Method for k-NN Based Binary Classification using Quadratic Programming
Lu Zhang 0023, Frans Coenen, Paul H. Leng
ECAI2
2002 Finding Association Rules with Some Very Frequent Attributes
Frans Coenen, Paul H. Leng
PKDD1
2002 Formalising optimal feature weight setting in case based diagnosis as linear programming problems
Lu Zhang 0023, Frans Coenen, Paul H. Leng
Knowl. Based Syst.2
2001 Computing Association Rules Using Partial Totals
Frans Coenen, Graham Goulbourne, Paul H. Leng
PKDD1
2001 Verification, validation, and integrity issues in expert and database systems: Two perspectives
abstract
This paper is directed at two central objectives. The first is to identify and to establish areas of overlap between the expert and database system domains. The second is to present a view of existing and ongoing work concerning the verification, validation, and integrity (VV&I) of rule base and database systems. The paper combines reviews from the two perspectives of the expert systems and database systems communities, with the express aim of identifying possibilities where VV&I knowhow of the one may also be of value to the other (and, vice versa), especially with respect to the identified areas of overlap. © 2001 John Wiley & Sons, Inc.
Frans Coenen, Barry Eaglestone, Mick J. Ridley
Int. J. Intell. Syst.1
2000 Algorithms for computing association rules using a partial-support tree
Graham Goulbourne, Frans Coenen, Paul H. Leng
Knowl. Based Syst.2
2000 Tesseral spatio-temporal reasoning for multi-dimensional data
Frans Coenen
Pattern Recognit.1
1999 Region Description and Comparative Analysis using a Tesseral Representation
abstract
Presents a region representation scheme and comparative analysis methods based on a tesseral addressing system. The proposed scheme is described in the context of a performance analysis of page segmentation methods, a document image analysis area that is particularly sensitive to both a successful region description scheme and efficient methods for comparative analysis. The proposed tesseral representation is more economical in storage than other Cartesian-based approaches and can be advantageous for comparative analysis.
Apostolos Antonacopoulos, Frans Coenen
ICDAR2
1998 Rulebase Checking Using a Spatial Representation
Frans Coenen
DEXA1
1998 KD in FM: Knowledge Discovery in Facilities Management Databases
Graham Goulbourne, Frans Coenen, Paul H. Leng
DEXA2
1998 Spatio-temporal Reasoning Using a Multi-dimensional Tesseral Representation
Frans Coenen, Bridget Beattie, Trevor J. M. Bench-Capon, Bernard M. Diaz, Michael J. R. Shave
ECAI1
1997 A Tesseral Approach to n-Dimensional Spatial Reasoning
Frans Coenen, Bridget Beattie, Trevor J. M. Bench-Capon, Bernard M. Diaz, Michael J. R. Shave
DEXA1
1996 An Ontology for Linear Spatial Reasoning
Frans Coenen, Bridget Beattie, Trevor J. M. Bench-Capon, Michael J. R. Shave, Bernard M. Diaz
DEXA1
1996 Temporal reasoning using tesseral addressing: towards an intelligent environmental impact assessment system
Frans Coenen, Bridget Beattie, Bernard M. Diaz, Trevor J. M. Bench-Capon, Michael J. R. Shave
Knowl. Based Syst.1
1995 Spatial Reasoning for GIS Using a Tesseral Data Representation
Bridget Beattie, Frans Coenen, Trevor J. M. Bench-Capon, Bernard M. Diaz, Michael J. R. Shave
DEXA2
1995 Developing Distributed Database Applications Using TSL
Frans Coenen, Ian Finch, Michael J. R. Shave, Trevor J. M. Bench-Capon
DEXA1
1995 Advanced binary encoded matrix representation for rule base verification
Frans Coenen
Knowl. Based Syst.1
1993 Representing Visual Conditions in a Legal knowledge Based System
abstract
Legal KBSs are based on knowledge contained in legal texts such as legislation, regulations and case histories and the practice of domain experts charged with operationalising this legislation. Legal texts and their opertionalisation can be analysed using textual analysis tools which lead to the production of a rule base which can be manipulated to establish a desired goal. In this paper we describe an approach to the development of legal KBSs where the legal texts include visual conditions which do not lend themselves to simple interpretation using textual analysis tools. The approach focuses on the use of preprocessors to generate descriptors derived from the geometrical interpretation of the visual data in question. These descriptors can then be used as direct input to a KBS without the need to include complex mathematics, to which KBS representations are not well suited, within individual rules. The approach was developed as part of a much larger project concerned with the production of a legal KBS to advise the navigators of ocean going vessels on how best to avoid collision with other vessels as prescribed by international maritime law. A fragment of the legislation on which this KBS is based is used as an example.
Frans Coenen, Trevor J. M. Bench-Capon, Peter Smeaton
ICAIL1
1992 Building Knowledge Based Systems for Maintainability
Frans Coenen, Trevor J. M. Bench-Capon
DEXA1
1992 Electronic Chart Representation and Interaction
Frans Coenen, Steve Fawcett, Peter Smeaton, Trevor J. M. Bench-Capon
DEXA1
1991 A Graphical Interactive Tool for KBS Maintenance
Frans Coenen, Trevor J. M. Bench-Capon
DEXA1
1991 Rule-based algorithms for geographic constraints in a marine knowledge-based system
Frans Coenen, Peter Smeaton
Knowl. Based Syst.1