VLDB 2026 Research / reviewers in the wild / expert
Anna Jurek-Loughrey
dblp:45/8120 · also Anna Jurek
· DBLP profile ↗
23ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DST-GAN: Dempster-Shafer Based Fusion for Multimodal Oversampling of Healthcare Data
Kevin Fee, Mohammed Hasanuzzaman, Suneil Jain, Ross G. Murphy, Anna Jurek-Loughrey |
AIME (2) | 5 |
| 2026 | Enhancing the interpretability of the mapper algorithmabstractAbstract The Mapper Algorithm is a powerful tool for representing the topology of a dataset’s structure as a similarity graph for the purposes of exploratory analysis. Despite Mapper’s ability to simplify complex high-dimensional data representations, interpreting the structure of the output graph remains a challenge. The conventional method of interpreting the Mapper graph by coloring nodes by features values is infeasible for high-dimensional data due to time limitations and the potential for subjectivity-related oversights. We present a novel method to enhance the interpretability of the Mapper algorithm. Specifically, we propose adapting eXplainable Artificial Intelligence techniques to determine feature importance, offering both local and global interpretations. Our approach can be used to assist domain experts in understanding functional differences across Mapper graphs, enabling them to draw meaningful conclusions from the graph’s structure. To validate our approach, we conducted experiments on five real-world medical datasets and the MNIST handwritten digit dataset. Our evaluation methods consist of a combination of visualization, classification tasks, and alignment of interpretations to existing literature. The results demonstrate our method’s effectiveness in providing a means to interpret Mapper graphs by highlighting the roles of specific features in the graph—such as pixel regions in MNIST and genes in TCGA datasets. Padraig Fitzpatrick, Anna Jurek-Loughrey, Pawel Dlotko |
Data Min. Knowl. Discov. | 2 |
| 2026 | Obstructive Sleep Apnea Prediction: A Comprehensive Review and Comparative StudyabstractAbstract Obstructive Sleep Apnea (OSA) is a highly prevalent sleep disorder linked to considerable public health burdens and comorbidities. However, its heterogeneous presentation and the limited accessibility of traditional diagnostic tools such as polysomnography (PSG) lead to widespread underdiagnosis. As a result, artificial intelligence (AI) approaches, including machine learning (ML) and deep learning (DL) models, have attracted attention as an alternative pathway to detection. This paper first provides a comprehensive review of AI-driven OSA diagnosis, covering different diagnosis problems, input-data types, data biases, pre-processing techniques, and model performance. We then leverage the largest clinical dataset used in OSA prediction to date, approximately 110,000 patients with 22,000 having complete entries for all 50 features, to systematically compare the performance of 39 ML/DL models. Our findings highlight the challenging nature of OSA prediction, with accuracies ranging from 29.66% to 46.9% for 4-class prediction and 46.04% to 87.18% for binary tasks. DL models such as DANet and GATE scored highest, whereas ensemble approaches such as LGBM and AdaBoost displayed more consistent performance across folds. However, as severe cases of OSA are easier to predict and over-represented in datasets, accuracy alone is insufficient for model evaluation and we explore a variety of metrics. Finally, imbalance correction and feature selection improved weaker models, but had only marginal effects on the best-performing models. Looking forwards, the development of more sophisticated and tailored DL models and large, high-quality datasets may help to break current performance barriers. We hope that our work can attract more attention to this challenging but interesting research problem. Huynh Thi Khanh Chi, Amonae Dabbs-Brown, Anna Jurek-Loughrey, James Mulhall, Tuan Dung Pham, Ngoc Phu Doan, Viet-Hung Tran, Zichi Zhang, Xuan Hoang Nguyen, Yimeng An, Peixin Li, Phi Hung Nguyen, Thi Linh Hoang, Xinming Shi, Hans Vandierendonck, Sébastien Bailly, Jean Louis Pépin, Son T. Mai |
Mach. Learn. | 3 |
| 2025 | Towards a Biological Evaluation Framework for Oversampling (BEFO) gene expression dataabstractMachine learning (ML) techniques are progressively being used in biomedical research to improve diagnostic and prognostic accuracy when used in conjunction with a clinician as a decision support system. However, many datasets used in biomedical research often suffer from severe class imbalance due to small population sizes, which causes machine learning models to become biased to majority class samples. Current oversampling methods primarily focus on balancing datasets without adequately validating the biological relevance of synthetic data, risking the clinical applicability of downstream model predictions. To address these shortcomings, we propose the Biological Evaluation Framework for Oversampling (BEFO) designed to ensure that synthetic gene expression samples accurately reflect the biological patterns present in original datasets. This innovation not only mitigates bias but enhances the trustworthiness of predictive models in clinical scenarios. We have developed a ranking method for synthetic samples based on this and evaluated each sample's inclusion based on its rank. This ranking method calculates the WGCNA gene co-expression clusters on the original dataset. Several random forests are constructed to assess the alignment of each synthetic sample to each cluster. Only synthetic samples more important than real samples are included in a study. The experimental results demonstrate that our proposed ML oversampling framework can improve the biological feasibility of oversampled datasets by an average of 11%, leading to improved classification performance by an average of 9% when compared against five state-of-the-art (SOTA) oversampling methods and ten classification algorithms across six real world gene expressions datasets. Thereby establishing a new standard for synthetic data evaluation in biomedical ML applications. Kevin Fee, Suneil Jain, Ross G. Murphy, Anna Jurek-Loughrey |
J. Biomed. Informatics | 4 |
| 2025 | New Automated Approach to Selection of Mapper Clustering ParametersabstractTopological methods have recently gained traction as powerful tools for extracting insights from high-dimensional data, forming the foundation of an approach known as Topological Data Analysis (TDA). Among the key developments in TDA is the Mapper algorithm, which constructs graph-based representations of complex datasets, capturing their topological structure at a user-defined resolution. The Mapper algorithm has shown promise across various applications, particularly in biomedical data analysis. However, its application requires careful selection of several parameters, especially the clustering algorithm and its settings. Without prior knowledge and a deep understanding of the data, these choices are non-trivial and can be a major barrier for researchers aiming to leverage Mapper effectively. In this work, we introduce enhancements to the Mapper algorithm to address this challenge. Specifically, we investigate the integration of ensemble learning (EL) techniques into Mapper’s graph construction to eliminate the need for arbitrary parameter selection. Additionally, we propose a data-driven criterion for selecting the clustering method best suited to the Mapper algorithm. Our experimental results demonstrate that the proposed approach enables the construction of Mapper graphs that accurately capture the underlying structure of the input data, all without manual parameter tuning. Padraig Fitzpatrick, Anna Jurek-Loughrey, Pawel Dlotko |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Extraction of Aquaculture Cages from High-Resolution Remote Sensing Images Based on Deep LearningabstractThe accurate recognition of the spatial distribution of aquaculture in coastal areas plays a crucial role in the management of natural resources and marine ecological environment protection. Using remote sensing detection method, the information of aquaculture areas can be quickly and accurately extracted from high-resolution remote sensing images. This work focuses on semantic segmentation and extraction of cage aquaculture regions using advanced deep learning algorithms. We selected Hainan Island in China as the experimental area and established the Hainan Island Offshore Cage Aquaculture Sources Dataset (HIOCASD) using high-resolution satellite remote sensing images from Gaofen-2 satellite. Six different deep learning models including DeepLabv3+, Segformer and U-Net architectures are evaluated with feature extraction via convolutional neural networks. The experimental results show that all models have excellent performance, especially U-Net model which uses VGG network as feature extractor. Through five cross-validations, its average F1-score is as high as 93.75%. Lu Bai 0006, Anna Jurek-Loughrey, Zhibao Wang |
IGARSS | 5 |
| 2023 | Ensemble Learning for Mapper Parameter OptimizationabstractThe Mapper algorithm is a technique from TDA used to create low-dimensional graph-based representations of high-dimensional data, proven effective in numerous exploratory data analysis tasks. The Mapper algorithm’s output depends on several user-chosen parameters, and selecting their values is a non-trivial choice, significantly narrowing its potential application in real-world scenarios. Research attempting to assist in selection of the parameters has been very limited to date. This paper is the first one to address the selection of Mapper’s three parameters simultaneously. The proposed idea incorporates the concept of Ensemble Learning into the Mapper algorithm. Using several datasets with known labels, we show that our method outperforms two baselines in recovering the dataset structure. Padraig Fitzpatrick, Anna Jurek-Loughrey, Pawel Dlotko, Jesús Martínez del Rincón |
ICTAI | 2 |
| 2022 | High-Value Token-Blocking: Efficient Blocking Method for Record LinkageabstractData integration is an important component of Big Data analytics. One of the key challenges in data integration is record linkage, that is, matching records that represent the same real-world entity. Because of computational costs, methods referred to as blocking are employed as a part of the record linkage pipeline in order to reduce the number of comparisons among records. In the past decade, a range of blocking techniques have been proposed. Real-world applications require approaches that can handle heterogeneous data sources and do not rely on labelled data. We propose high-value token-blocking (HVTB), a simple and efficient approach for blocking that is unsupervised and schema-agnostic, based on a crafted use of Term Frequency-Inverse Document Frequency. We compare HVTB with multiple methods and over a range of datasets, including a novel unstructured dataset composed of titles and abstracts of scientific papers. We thoroughly discuss results in terms of accuracy, use of computational resources, and different characteristics of datasets and records. The simplicity of HVTB yields fast computations and does not harm its accuracy when compared with existing approaches. It is shown to be significantly superior to other methods, suggesting that simpler methods for blocking should be considered before resorting to more sophisticated methods. Kevin O'Hare, Anna Jurek-Loughrey, Cassio P. de Campos |
ACM Trans. Knowl. Discov. Data | 2 |
| 2021 | Unsupervised Keyword Combination Query Generation from Online Health Related Content for Evidence-Based Fact CheckingabstractFalse information in the domain of online health related articles is of great concern, which can be witnessed in the current pandemic situation of Covid-19. It is markedly different from fake news in the political context as health information should be evaluated against the most recent and reliable medical resources such as scholarly repositories. However, one of the challenges with such an approach is the retrieval of the pertinent resources. In this work, we formulate a new unsupervised task of generating queries using keywords extracted from a health-related article which can be further applied to retrieve relevant authoritative and reliable medical content from scholarly repositories to assess the article’s veracity. We propose a three-step approach for it and illustrate that our method is able to generate effective queries. We also curate a new dataset to aid the evaluation for this task which will be made available upon request. Pritam Deka, Anna Jurek-Loughrey, Deepak P 0001 |
iiWAS | 2 |
| 2021 | The topology of data: opportunities for cancer researchabstractMOTIVATION: Topological methods have recently emerged as a reliable and interpretable framework for extracting information from high-dimensional data, leading to the creation of a branch of applied mathematics called Topological Data Analysis (TDA). Since then, TDA has been progressively adopted in biomedical research. Biological data collection can result in enormous datasets, comprising thousands of features and spanning diverse datatypes. This presents a barrier to initial data analysis as the fundamental structure of the dataset becomes hidden, obstructing the discovery of important features and patterns. TDA provides a solution to obtain the underlying shape of datasets over continuous resolutions, corresponding to key topological features independent of noise. TDA has the potential to support future developments in healthcare as biomedical datasets rise in complexity and dimensionality. Previous applications extend across the fields of neuroscience, oncology, immunology and medical image analysis. TDA has been used to reveal hidden subgroups of cancer patients, construct organizational maps of brain activity and classify abnormal patterns in medical images. The utility of TDA is broad and to understand where current achievements lie, we have evaluated the present state of TDA in cancer data analysis. RESULTS: This article aims to provide an overview of TDA in Cancer Research. A brief introduction to the main concepts of TDA is provided to ensure that the article is accessible to readers who are not familiar with this field. Following this, a focussed literature review on the field is presented, discussing how TDA has been applied across heterogeneous datatypes for cancer research. Ciara Frances Loughrey, Padraig Fitzpatrick, Nick Orr, Anna Jurek-Loughrey |
Bioinform. | 4 |
| 2021 | Novel deep learning-based solution for identification of prognostic subgroups in liver cancer (Hepatocellular carcinoma)abstractAbstract Background Liver cancer (Hepatocellular carcinoma; HCC) prevalence is increasing and with poor clinical outcome expected it means greater understanding of HCC aetiology is urgently required. This study explored a deep learning solution to detect biologically important features that distinguish prognostic subgroups. A novel architecture of an Artificial Neural Network (ANN) trained with a customised objective function (L RSC ) was developed. The ANN should discover new data representations, to detect patient subgroups that are biologically homogenous (clustering loss) and similar in survival (survival loss) while removing noise from the data (reconstruction loss). The model was applied to TCGA-HCC multi-omics data and benchmarked against baseline models that only use a reconstruction objective function (BCE, MSE) for learning. With the baseline models, the new features are then filtered based on survival information and used for clustering patients. Different variants of the customised objective function, incorporating only reconstruction and clustering losses (L RC ); and reconstruction and survival losses (L RS ) were also evaluated. Robust features consistently detected were compared between models and validated in TCGA and LIRI-JP HCC cohorts. Results The combined loss (L RSC ) discovered highly significant prognostic subgroups ( P -value = 1.55E−77) with more accurate sample assignment (Silhouette scores: 0.59–0.7) compared to baseline models (0.18–0.3). All L RSC bottleneck features (N = 100) were significant for survival, compared to only 11–21 for baseline models. Prognostic subgroups were not explained by disease grade or risk factors. Instead L RSC identified robust features including 377 mRNAs, many of which were novel (61.27%) compared to those identified by the other losses. Some 75 mRNAs were prognostic in TCGA, while 29 were prognostic in LIRI-JP also. L RSC also identified 15 robust miRNAs including two novel (hsa-let-7g; hsa-mir-550a-1) and 328 methylation features with 71% being prognostic. Gene-enrichment and Functional Annotation Analysis identified seven pathways differentiating prognostic clusters. Conclusions Combining cluster and survival metrics with the reconstruction objective function facilitated superior prognostic subgroup identification. The hybrid model identified more homogeneous clusters that consequently were more biologically meaningful. The novel and prognostic robust features extracted provide additional information to improve our understanding of a complex disease to help reveal its aetiology. Moreover, the gene features identified may have clinical applications as therapeutic targets. Alice R. Owens, Caitríona E. McInerney, Kevin M. Prise, Darragh G. McArt, Anna Jurek-Loughrey |
BMC Bioinform. | 5 |
| 2021 | Social network analysis of open source software: A review and categorisation
Kelvin McClean, Des Greer, Anna Jurek-Loughrey |
Inf. Softw. Technol. | 3 |
| 2020 | ReSCo-CC: Unsupervised Identification of Key Disinformation SentencesabstractDisinformation is often presented in long textual articles, especially when it relates to domains such as health, often seen in relation to COVID-19. These articles are typically observed to have a number of trustworthy sentences among which core disinformation sentences are scattered. In this paper, we propose a novel unsupervised task of identifying sentences containing key disinformation within a document that is known to be untrustworthy. We design a three-phase statistical NLP solution for the task which starts with embedding sentences within a bespoke feature space designed for the task. Sentences represented using those features are then clustered, following which the key sentences are identified through proximity scoring. We also curate a new dataset with sentence level disinformation scorings to aid evaluation for this task; the dataset is being made publicly available to facilitate further research. Based on a comprehensive empirical evaluation against techniques from related tasks such as claim detection and summarization, as well as against simplified variants of our proposed approach, we illustrate that our method is able to identify core disinformation effectively. Soumya Suvra Ghosal, Deepak P 0001, Anna Jurek-Loughrey |
iiWAS | 3 |
| 2020 | Siamese Neural Network for Unstructured Data LinkageabstractData integration is one of the key problems in the era of Big Data analytics. The key challenge of data integration is the identification of records representing the same entities (e.g. person). This task is referred to as Record Linkage. It is uncommon for different data sources to share a unique identifier hence the records must be matched by comparing their corresponding values. Most of the existing methods assume that records across different sources are structured and represented by the same set of attributes (e.g. name, date of birth). However, nowadays majority of the data comes without structure (e.g. social media sites). We propose a new approach to Record Linkage based on application of Siamese Neural Network. The model can be applied with structured, semi-structured and unstructured records and it does not assume a common format across different data sources. We demonstrate that the model performs on par with other approaches, which make constraining assumptions regarding the data. Anna Jurek-Loughrey |
iiWAS | 1 |
| 2019 | An unsupervised blocking technique for more efficient record linkage
Kevin O'Hare, Anna Jurek-Loughrey, Cassio P. de Campos |
Data Knowl. Eng. | 2 |
| 2018 | It Pays to Be Certain: Unsupervised Record Linkage via Ambiguity Minimization
Anna Jurek-Loughrey, Deepak P 0001 |
PAKDD (3) | 1 |
| 2018 | A new technique of selecting an optimal blocking method for better record linkage
Kevin O'Hare, Anna Jurek-Loughrey, Cassio P. de Campos |
Inf. Syst. | 2 |
| 2017 | Privacy preserving record linkage in the presence of missing values
Yuan Chi, Jun Hong 0001, Anna Jurek-Loughrey, Weiru Liu, Dermot O'Reilly |
Inf. Syst. | 3 |
| 2017 | A novel ensemble learning approach to unsupervised record linkage
Anna Jurek-Loughrey, Jun Hong 0001, Yuan Chi, Weiru Liu |
Inf. Syst. | 1 |
| 2014 | Sentiment Classification by Combining Triplet Belief Functions
Yaxin Bi, Maurice D. Mulvenna, Anna Jurek-Loughrey |
KSEM | 3 |
| 2014 | Clustering-Based Ensembles as an Alternative to StackingabstractOne of the most popular techniques of generating classifier ensembles is known as stacking which is based on a meta-learning approach. In this paper, we introduce an alternative method to stacking which is based on cluster analysis. Similar to stacking, instances from a validation set are initially classified by all base classifiers. The output of each classifier is subsequently considered as a new attribute of the instance. Following this, a validation set is divided into clusters according to the new attributes and a small subset of the original attributes of the instances. For each cluster, we find its centroid and calculate its class label. The collection of centroids is considered as a meta-classifier. Experimental results show that the new method outperformed all benchmark methods, namely Majority Voting, Stacking J48, Stacking LR, AdaBoost J48, and Random Forest, in 12 out of 22 data sets. The proposed method has two advantageous properties: it is very robust to relatively small training sets and it can be applied in semi-supervised learning problems. We provide a theoretical investigation regarding the proposed method. This demonstrates that for the method to be successful, the base classifiers applied in the ensemble should have greater than 50% accuracy levels. Anna Jurek-Loughrey, Yaxin Bi, Shengli Wu 0001, Chris D. Nugent |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | A Cluster-Based Classifier Ensemble as an Alternative to the Nearest Neighbor EnsembleabstractThe combination of multiple classifiers, commonly referred to as an ensemble, has previously demonstrated the ability to improve overall classification accuracy in many application domains. Some ensemble techniques, however, cannot easily improve the performance of stable classification methods. One such example of a stable classification method is the k Nearest Neighbor (kNN) Classifier. In this paper we propose an alternative to the kNN ensemble method through the use of a clustering technique applied for the purpose of selecting the neighborhood of a new instance. In addition, a novel combination function based on exponential support (ExSupp) has been introduced. The proposed approach exhibited improved classification results in 16 out 20 data sets which were considered in comparison with a single kNN and a kNN ensemble based approach. Besides higher classification accuracy the proposed method exhibited higher levels of efficiency in terms of classification time. Anna Jurek-Loughrey, Yaxin Bi, Shengli Wu 0001, Chris D. Nugent |
ICTAI | 1 |
| 2011 | Classification by Clusters Analysis - An Ensemble Technique in a Semi-supervised ClassificationabstractIn this work we adopt a previously introduced meta-learning classification method for semi-supervised learning problems. In our previous work we illustrated that the method is successful when applied in a supervised classification problem. In our current work the results demonstrate that following refinements made to the method it can be successfully applied to semi-supervised classification cases. Anna Jurek-Loughrey, Yaxin Bi, Shengli Wu 0001, Chris D. Nugent |
ICTAI | 1 |