Haiying Wang 0001

dblp:45/1696-1 · also Hai-Ying Wang 0001 · DBLP profile ↗
← Back
78ranked-venue papers
21as first author
14since 2021 · last 2026
0000-0001-8358-9065ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 65 · 18 first-author · 11 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Where and How: Mining Convertible Outlying Aspect for Outlier Interpretation
abstract
Outlier interpretation methods, based on outlying aspect mining, have been extensively utilized in diverse applications due to their effectiveness and interpretability. The primary objective of these methods is to identify an outlying feature subspace where a detected outlier deviates most significantly from the inliers. However, this subspace is typically not personalized and incomplete. In this article, we propose a novel convertible outlying aspect mining method named Mining convErtible ouTlying Aspect (META) to interpret the detected outlier. META not only identifies a personalized outlying feature subspace (i.e., where) that differentiates the detected outlier from inliers, but also quantifies the outlying direction and degree within this subspace (i.e., how), thereby offering actionable interpretative insights, rather than corrective actions, for understanding the outlier. Specifically, META defines a convertible outlying aspect with a convertible cost to convert the detected outlier into a converted instance, and employs a pretrained adversary to evaluate whether the instance is an inlier or not. Subsequently, we formulate an objective function that minimizes the convertible cost, ensures the successful conversion of the instance into an inlier, and minimizes the size of the outlying feature subspace. META leverages this objective function to learn an optimal convertible outlying aspect for the detected outlier. The optimal convertible outlying aspect provides the outlying feature subspace, outlying direction, outlying degree, and converted instance. Empirical results from experiments conducted on both real-world and synthetic datasets demonstrate that META outperforms state-of-the-art (SOTA) baselines.
Lili Guan, Lei Duan, Xinye Wang, Zhuoling Li, Haiying Wang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Review of Learning-Based Antibody Design: From Sequence to Structure
abstract
Improving antibodies' affinity and specificity has traditionally relied on iterative display selections or structure-based design, both costly and time-intensive. Recent advances in Deep Learning offer data-driven priors that effectively narrow the sequence space before expensive experiments. This paper provides an overview of the progress and challenges of learning -based antibody design. Adopting a pipeline-first perspective, this review organises current methods into three categories: (A) sequence-only protein language models (PLMs); (B) structure-aware strategies, including inverse folding and complex-aware optimisation; and (C) integrated AI-physics workflows. To avoid mixing endpoints, prospective wet-lab outcomes (e.g. hit rates, affinity gains) are reported separately from structure-linked sur-rogates (e.g. region recovery, refold root-mean-square deviation (RMSD), deep mutational scanning (DMS) correlation). Evidence indicates that sequence-only PLMs are effective for low-budget screening, inverse folding methods provide backbone-conditioned ranking and structure-preserving edits, and lightweight AI-physics overlays help prioritise manufacturable candidates. A concise method -selection guide is provided for different data availability scenarios.
Jialin Lyu, Ciaran Doherty, Hugh Morgan, Haiying Wang 0001, Huiru Zheng
BIBM4
2025 Consensus-Aware Balance Learning for Sexually Suggestive Video Classification
Haiying Wang 0001, Matthew Burns, Meng Liu 0006
CVM (1)3
2024 DDTExplainer: Mining Drug-Disease Therapeutic Mechanisms based on GNN Explainability
abstract
For a clinical prescription, clarifying the molecular mechanisms of actions (MMOAs) of the drug-disease interaction is helpful to optimize treatment, suggest possible side effects, and realize individualized treatment. Considering the relations among multiple biomedical entities, such as drugs, diseases, targets (genes), and pathways, what paths can be extracted connecting these biomedical entities to resemble the real mechanisms of a specific drug to a particular disease? Answering this question is crucial for understanding the underlying molecular mechanisms behind complex drug actions and identifying key pathways that can facilitate effective therapeutic interventions. In this paper, we propose an approach DDTExplainer that constructs a path-based graph neural network (GNN) explainer to mine the drug-disease therapeutic mechanisms. Technically, DDTExplainer transforms the drug-disease therapeutic mechanisms mining task into a GNN-based link prediction model explanation task. Firstly, a GNN-based drug-disease therapeutic prediction model is trained and joint-optimized with a translation-based graph embedding model. Secondly, mask learning is utilized to find the most prediction-influential edges and generate the path-based explanations with the shortest path algorithm. Finally, we assess the efficacy of DDTExplainer on a ground-truth dataset that consists of labeled entries describing the drug-disease therapeutic mechanisms. These labels are derived from well-established drug-target interactions, disease-target interactions, as well as target-pathway relationships, which were verified by wet experiments.
Yidan Zhang 0001, Lei Duan, Huiru Zheng, Haiying Wang 0001, Yongmei Lu
BIBM4
2024 Application of Smart Insoles for Recognition of Activities of Daily Living: A Systematic Review
abstract
Recent years have witnessed the increasing literature on using smart insoles in health and well-being, and yet, their capability of daily living activity recognition has not been reviewed. This paper addressed this need and provided a systematic review of smart insole-based systems in the recognition of Activities of Daily Living (ADLs). The review followed the PRISMA guidelines, assessing the sensing elements used, the participants involved, the activities recognised, and the algorithms employed. The findings demonstrate the feasibility of using smart insoles for recognising ADLs, showing their high performance in recognising ambulation and physical activities involving the lower body, ranging from 70% to 99.8% of Accuracy, with 13 studies over 95%. The preferred solutions have been those including machine learning. A lack of existing publicly available datasets has been identified, and the majority of the studies were conducted in controlled environments. Furthermore, no studies assessed the impact of different sampling frequencies during data collection, and a trade-off between comfort and performance has been identified between the solutions. In conclusion, real-life applications were investigated showing the benefits of smart insoles over other solutions and placing more emphasis on the capabilities of smart insoles.
Luigi D'Arco, Graham McCalmont, Haiying Wang 0001, Huiru Zheng
ACM Trans. Comput. Heal.3
2023 A Novel Mixed Effects Random Forest Approach for Predicting Dairy Cattle Methane Emissions
abstract
Methane (CH4) emissions produced by dairy cattle (DC) are a key contributor to global warming. To assess the effectiveness of strategies designed to mitigate CH4emissions, complex and expensive recording equipment is required. Therefore, the use of predictive models based on animal information provides a more accessible alternative. Traditionally, Statistical (SA) methods have been employed in the prediction of DC CH4emissions. However due to the smart farming revolution, the scale and variety of complex animal information now available for the prediction of DC CH4emissions has grown exponentially, and within them are likely to exist non-linear relationships which these traditional SA models may struggle to capture. Therefore, this research aims to explore if Machine Learning (ML) models are a viable alternative for the prediction of DC CH4emissions, as they can handle and extract these inevitable non-linear relationships present within today's large, heterogeneous datasets. In this research, we compared a traditional SA method, a Linear Mixed Effects (ME) model, with an original ML method, a Random Forest (RF) model, as well as a novel SA/ML hybrid method, a Mixed Effects Random Forest (MERF) model, in the prediction of CH4emissions (CH4g/d) produced by DC across 32 experiments. The ML RF model was able to challenge the traditional SA ME model in the prediction of DC CH4emissions, achieving a Root Mean Square Prediction Error (RMSPE) and Concordance Correlation Coefficient (CCC) of 52.73 CH4g/d and 0.70 respectively, compared to a ME model’s 53.90 CH4g/d and 0.71. When both the ME and RF models were combined within the novel SA/ML hybrid MERF model, a lower RMSPE and higher CCC were achieved than by each of its composite parts in isolation, 51.87 CH4g/d and 0.73 respectively. These results demonstrate the potential of ML in the prediction of DC CH4emissions, particularly when hybridised alongside traditional SA methods.
Stephen Ross, Tianhai Yan, Haiying Wang 0001, Masoud Shirali, Huiru Zheng
BIBM3
2023 DARE: Sequence-Structure Dual-Aware Encoder for RNA-Protein Binding Prediction
abstract
Predicting RNA-protein binding sites helps to explore the mechanisms of the interaction between RNA and proteins. Numerous deep learning methods have been applied to predict RNA-protein binding sites. Some of these methods use only sequence information for prediction which could lose information about the topology. And there may be a loss of important information if the secondary structure features are simply represented as one-hot matrices. Furthermore, existing deep learning methods are usually based on convolutional neural networks for feature extraction, which tend to focus on local features. As for the information of the whole sequence, existing methods usually ignore global features. Therefore, we propose a novel deep learning model called DARE for RNA-protein binding sites prediction using both sequence and secondary structure information of RNA. DARE employs the secondary structure feature extraction module to capture the features of the RNA secondary structure and learn the topological information. Therefore, we design a local feature extraction module and a global feature integration module to capture the whole information of RNA. Thus we can achieve the purpose of complementary information. Extensive experiments demonstrate that DARE outperforms baselines. Our analysis of the case study further confirm the effectiveness of DARE.
Luhan Shen, Chengxin He, Haiying Wang 0001, Yuening Qu, Lei Duan
BIBM3
2023 A Factor Graph Based Indoor Localization Approach for Healthcare
abstract
In healthcare facilities, indoor localization technology has a broad range of applications. Traditional Pedestrian Dead Reckoning (PDR) and WiFi fingerprint-based methods each have their limitations. To address these challenges, this study introduces a multi-source fusion indoor localization system that uses a Factor Graph to integrate inertial positioning algorithms with WiFi fingerprint-based localization. The system processes accelerometer and gyroscope data using a data-driven PDR algorithm. For WiFi localization, considering that the extensive data collection required is a significant barrier to the deployment of WiFi-based localization methods, the proposed approach applies Gaussian process regression techniques to limited WiFi fingerprint data, significantly reducing initial deployment costs and enhancing accuracy. Finally, the entire system employs a Factor Graph for the integration of the data-driven PDR and WiFi fingerprint localization results. Experimental results show that, compared to using only inertial or WiFi data for localization, this method significantly improves localization accuracy. The findings suggest that this approach could prompt the utilization of indoor localization technology in healthcare facilities.
Shiyu Zheng, Ao Peng, Lingxiang Zheng, Huiru Zheng, Haiying Wang 0001
BIBM8
2023 DeepHAR: a deep feed-forward neural network algorithm for smart insole-based human activity recognition
abstract
Abstract Health monitoring, rehabilitation, and fitness are just a few domains where human activity recognition can be applied. In this study, a deep learning approach has been proposed to recognise ambulation and fitness activities from data collected by five participants using smart insoles. Smart insoles, consisting of pressure and inertial sensors, allowed for seamless data collection while minimising user discomfort, laying the baseline for the development of a monitoring and/or rehabilitation system for everyday life. The key objective has been to enhance the deep learning model performance through several techniques, including data segmentation with overlapping technique (2 s with 50% overlap), signal down-sampling by averaging contiguous samples, and a cost-sensitive re-weighting strategy for the loss function for handling the imbalanced dataset. The proposed solution achieved an Accuracy and F1-Score of 98.56% and 98.57%, respectively. The Sitting activities obtained the highest degree of recognition, closely followed by the Spinning Bike class, but fitness activities were recognised at a higher rate than ambulation activities. A comparative analysis was carried out both to determine the impact that pre-processing had on the proposed core architecture and to compare the proposed solution with existing state-of-the-art solutions. The results, in addition to demonstrating how deep learning solutions outperformed those of shallow machine learning, showed that in our solution the use of data pre-processing increased performance by about 2%, optimising the handling of the imbalanced dataset and allowing a relatively simple network to outperform more complex networks, reducing the computational impact required for such applications.
Luigi D'Arco, Haiying Wang 0001, Huiru Zheng
Neural Comput. Appl.2
2022 A Rapid Detection of Parkinson's Disease using Smart Insoles: A Statistical and Machine Learning Approach
abstract
Determining whether a subject has a gait impairment due to a disease or to the loss of muscularity due to advancing age is fundamental for an early diagnosis of musculoskeletal diseases. Parkinson’s is the second most common neurodegenerative disease. The disease’s most prevalent symptom is slow movement or sluggish gait, which can adversely impact the individual’s quality of life. Generally, the gait analysis is carried out on long test sessions, which include for example long periods of walking, that cause inconvenience when the subjects under test have marked gait impairments. To help the diagnosis of Parkinson’s disease, in this study we investigated the classification of Parkinson’s disease by analysing only a few seconds of walking data using smart insoles, statistical analysis and machine learning techniques. The data from the smart insoles was assessed using correlation analysis. By creating pressure groups and analysing their values, it was found that the number of sensors could be reduced from 16 to 7. Furthermore, a feature vector representing the subject’s gait was created by applying on the data a time windowing segmentation of 5 seconds and extracting six statistical features (mean, variance, skewness, kurtosis, energy and entropy). Four different models have been compared in terms of classification performance, reaching an F1-Score in the classification of patients with Parkinson’s against healthy subjects, considering adult and elderly subjects as two separate classes, of 97.04% using the Random Forest. Such metric increased to 98.89%, using the K-Nearest Neighbours when healthy subjects were considered as a single class. The models’ performance for each experiment was determined to be statistically equivalent, demonstrating the potential of this approach to provide the groundwork for the rapid detection of Parkinson’s disease. Although the performance obtained is promising the number of subjects included in the study was fairly low, with a high bias towards the number of healthy subjects. Hence, in future work, the proposed solution will be tested on a larger cohort to ascertain its robustness.
Luigi D'Arco, Haiying Wang 0001, Huiru Zheng
BIBM2
2022 A New Phylogeny-Driven Random Forest-Based Classification Approach for Functional Metagenomics
abstract
Classifying microbial genes into their functional repertoire is an important task for metagenomic studies, where the research community is trying to develop Machine Learning (ML) based methods to achieve good classification performance. Random Forest (RF) has been proposed as one of the most favorable methods for such supervised analysis when applied over the abundance profiles of microbial genes mapping them to functional phenotypes. To further explore and make optimization in the existing RF model (based on the biological relationships between microbial features), a new classification method based on RF as guided by the evolutionary ancestry of microbial phylogeny, i.e. Phylogeny-RF, has been developed in this paper. This method facilitates to capture the effects of phylogenetic relatedness in a ML classifier itself. Closely related microbes by phylogeny are highly correlated and tend to have similar genetic and phenotypic traits. Such microbes behave similarly; and hence tend to be selected together or one of these could be dropped from the analysis, to make the ML process better. The proposed Phylogeny-RF algorithm has been compared with state-of-the-art classification methods including RF and the phylogeny-aware method of MetaPhyl, using 2 real-world 16S rRNA metagenomic data sets. It is observed that the proposed method performed better than the other phylogeny-driven benchmarks. For example, Phylogeny-RF attained a high AUC of 0.949 over soil microbiomes in comparison to other benchmarks.
Jyotsna Talreja Wassan, Haiying Wang 0001, Huiru Zheng
BIBM2
2022 An Enhanced Visual SLAM Supported by the Integration of Plane Features for the Indoor Environment
abstract
This paper presents an enhanced indoor RGB-D simultaneously localisation and mapping (SLAM) system based on the integration of plane and point features. A new method was proposed to register each point feature to a corresponding plane feature and then modify its position accordingly. The plane features are parallelly extracted from depth data sources and used jointly to solve the camera pose with point features. Both point and plane features are stored on the map and used for backend optimisation, where the weights associated with features can be dynamically updated. At the same time, the on-plane feature points are fixed during the optimisation. The proposed method has been tested with open-source benchmarks, including the scenarios with or without a structured environment. Experiment results demonstrated that the proposed algorithm performs better than other widely cited visual SLAM systems in some structured environments, in which the point features form plane features without introducing excessive errors.
Bingxin Zi, Haiying Wang 0001, Huiru Zheng
IPIN2
2021 A PDRVIO Loosely coupled Indoor Positioning System via Robust Particle Filter
abstract
In recent years, the Visual Inertial Odometry (VIO) technology has attracted attention as a support technology that improves medical experience and management efficiency. However, the performance of most VIO systems will drop drastically when the light intensity changes significantly or there are few texture features from the images. This paper designs a visual-inertial fusion-based navigation Indoor Positioning system to deal with the challenging scenario. It loosely coupled an inertial sensor-based pedestrian dead reckoning (PDR) model with the VIO model via a robust particle filter. The state estimation of the particle filter is based on the PDR model. The VIO model is used for the measurements of the particle filter. It compensates the gross errors of the VIO with a visual error propagation model which is established according to the posterior observation residuals of visual feature points. It is verified through experiments that the PDR/VIO fusion indoor positioning system based on the robust particle filter implemented in this paper has improved positioning accuracy and strong ability to deal with complex scenes.
Xinwei Hu, Weilong Huang, Lingxiang Zheng, Ao Peng, Huiru Zheng, Haiying Wang 0001
BIBM8
2021 Deep learning in systems medicine
abstract
Systems medicine (SM) has emerged as a powerful tool for studying the human body at the systems level with the aim of improving our understanding, prevention and treatment of complex diseases. Being able to automatically extract relevant features needed for a given task from high-dimensional, heterogeneous data, deep learning (DL) holds great promise in this endeavour. This review paper addresses the main developments of DL algorithms and a set of general topics where DL is decisive, namely, within the SM landscape. It discusses how DL can be applied to SM with an emphasis on the applications to predictive, preventive and precision medicine. Several key challenges have been highlighted including delivering clinical impact and improving interpretability. We used some prototypical examples to highlight the relevance and significance of the adoption of DL in SM, one of them is involving the creation of a model for personalized Parkinson's disease. The review offers valuable insights and informs the research in DL and SM.
Haiying Wang 0001, Estelle Pujos-Guillot, Blandine Comte, João Luís de Miranda, Vojtech Spiwok, Ivan Chorbev, Filippo Castiglione, Paolo Tieri, Steven Watterson, Roisin McAllister, Tiago De Melo Malaquias, Massimiliano Zanin, Taranjit Singh Rai, Huiru Zheng
Briefings Bioinform.1
2020 Bayesian Network Approach to Modelling Nitrogen Utilization Efficiency of Dairy Cows
abstract
Losses of nitrogen (N) from dairy cattle farming system cause environmental pollution and impact human health. A number of statistical models for predicting manure N excretion from lactating dairy cows have been developed based on regression analysis. In this study, we proposed a Bayesian network-based approach to modelling relationships among factors influencing manure N excretion of lactating dairy cows using a dataset collated from total diet digestibility studies undertaken at Agri-Food and Biosciences Institute in Northern Ireland. The preliminary results indicate that Bayesian network model can be used to capture relationships among factors that influence N utilization efficiency and can be used to establish causal influence among predictors. These may provide an effective tool for optimizing the management of feed N resource for current dairy production and developing strategies to reduce N excretion in dairy production systems.
Xianjiang Chen, Huiru Zheng, Haiying Wang 0001, Tianhai Yan
BIBM3
2020 A multilayer co-occurrence network reveals the systemic difference of diet-based rumen microbiome associated with methane yield phenotype
abstract
Reducing rumen methane emissions based on the basal diet is one primary approach for livestock production. However, microbiological mechanisms of rumen methane emissions under different diets are very complex and lack comprehensive understanding. This research developed an innovative multilayer network framework to investigate relationships of rumen microbes, transformations of metabolites and functions of rumen microbial genes. The multilayer network model explored the diets-based variation in rumen microbial system differences. The topological structure of the multilayer network reflects volatile fatty acid (VFA) fermentation characteristics between diets. Propionate is the highest degree node of the concentrate-diet (CONC) multilayer network while acetate has the highest degree in the forage-diet (FOR) multilayer network. The microbial communities under FOR diet are more diverse and have more complicated interactions. Methanobrevibacter copresented with Methanosphaera only in the CONC network. There are 33 microbial functional genes in the FOR network that are enriched in the methane synthesize functional module, including the CO2 reduction and the enzymatic reaction of acetyl phosphate to synthesize acetyl-CoA. The microbial genes of the CONC network are found to be associated with the serine biosynthetic and the biochemical process of acetate convert to acetyl-CoA and acetyl phosphate. This study verified the variations of diet-based methane phenotypic are results of systemic changing in the metabolism of the entire rumen microbiome.
Haiying Wang 0001, Huiru Zheng, Richard J. Dewhurst, Rainer Roehe
BIBM2
2020 Improving the Inference of Co-Occurrence Networks in the Bovine Rumen Microbiome
abstract
The importance of the composition and signature of rumen microbial communities has gained increasing attention. One of the key techniques was to infer co-abundance networks through correlation analysis based on relative abundances. While substantial insights and progress have been made, it has been found that due to the compositional nature of data, correlation analysis derived from relative abundance could produce misleading results and spurious associations. In this study, we proposed the use of a framework including a compendium of two correlation measures and three dissimilarity metrics in an attempt to mitigate the compositional effect in the inference of significant associations in the bovine rumen microbiome. We tested the framework on rumen microbiome data including both 16S rRNA and KEGG genes associated with methane production in cattle. Based on the identification of significant positive and negative associations supported by multiple metrics, two co-occurrence networks, e.g., co-presence and mutual-exclusion networks, were constructed. Significant modules associated with methane emissions were identified. In comparison to previous studies, our analysis demonstrates that deriving microbial associations based on the correlations between relative abundances may not only lead to missing information but also produce spurious associations. To bridge together different co-presence and mutual-exclusion relations, a multiplex network model has been proposed for integrative analysis of co-occurrence networks which has great potential to support the prediction of animal phytotypes and to provide additional insights into biological mechanisms of the microbiome associated with the traits.
Huiru Zheng, Haiying Wang 0001, Richard J. Dewhurst, Rainer Roehe
IEEE ACM Trans. Comput. Biol. Bioinform.2
2019 Automated Object Tracking for Animal Behaviour Studies
abstract
It can be an arduous and time consuming task to manually curate video data for tracking the movement of an object. This is particularly challenging when tracking animals whose behaviour is sporadic and unpredictable, but it can provide pertinent information when monitoring health, conditions and treatments. Here, we evaluate the use of machine learning retroactively applied to automatically track specific steers in hours of video with minimal manual interaction. This is approached using the Faster R-CNN Object Detection algorithm with VGG-16 acting as a feature extractor. Performance on a number of video segments is presented and discussed, and the issues encountered are outlined. This highlights a number of guidelines that should be taken under consideration when generating video data to improve object detection performance, and helps define the applicability of the approach to pre-existing data.
Timmy Manning, Miguel Somarriba, Rainer Roehe, Simon Turner, Haiying Wang 0001, Huiru Zheng, Brian Kelly, Jennifer Lynch, Paul Walsh
BIBM5
2019 Is there an Optimal Technology to Provide Personal Supportive Feedback in Prevention of Obesity?
abstract
Obesity is a global challenge that affects health and wellbeing worldwide. In this position paper, we review the digital technology used in prevention of obesity and present the proposed STOP project that integrates state-of-the-art wearable technology, chatbot, gamification data fusion, and machine learning with the aim to provide personalised supportive feedback for preventing obesity and maintaining healthy weight. Implication of sensitive data with General Data Protection Regulation (GDPR) is discussed. We conclude that machine learning plays an important role in data fusion, analytics, and providing optimal messaging tailored design to support healthy weight.
Simone Sandri, Matthias L. Hemmje, Huiru Zheng, Felix Engel 0002, Anne Moorhead, Haiying Wang 0001, Raymond R. Bond, Michael F. McTear, Andrea Molinari, Paolo Bouquet
BIBM6
2019 A knowledge driven mutual information-based analytical framework for the identification of rumen metabolites
abstract
Metabolites are the final product of biochemical reactions in the rumen micro-ecological system and very sensitive to changes of microbial genes. However, limited by the spectra library and the computational techniques of structure identification, the identification of metabolites from non-targeted metabolomics is time-consuming and inefficient. The absence of specific information about metabolites makes the biological interpretation of the quantitative analysis of metabolomics meaningless. Based on the nonlinear association between microbial genes and metabolites, combined with knowledge of metabolic pathways from the KEGG database, this study developed a knowledge driven mutual information-based analytical framework for identifying metabolites associated with integrals derived from NMR analysis results. In this study, one known metabolite and three sets of integrals with unknow metabolites were identified within the novel framework. The results showed that this mutual information-based framework could very efficiently target metabolites that may correspond to integrals from NMR spectra.
Huiru Zheng, Haiying Wang 0001, Richard J. Dewhurst, Rainer Roehe
BIBM3
2019 A Phylogeny-aware Feature Ranking for Classification of Cattle Rumen Microbiome
abstract
Metagenomics is proliferating for studying environmental microbial communities and their role in animal functions. This paper aims to study the role of functions of microbial communities present in cattle (Bos taurus) and their relation to dietary supplement usage. The functional study was conducted as part of the EU H2020 MetaPlat project11MetaPlat, http://www.metaplat.eu. In this research, we proposed a novel phylogeny-driven approach to classify 16S rRNA samples from cattle rumen microbiome and relate them to the functional phenotype of diet (referred to as functional analysis). Phylogeny covers biological relationships from different taxonomical levels combined with their respective evolutionary measures. We performed this analysis by proposing a novel method based on phylogeny-adjusted distance-based indices. These indices are used in ranking microbial feature space derived from the topology of the phylogenetic tree. The integrative approach incorporating phylogeny into feature engineering as part of machine learning (ML) modeling, achieved high predictive performance with Accuracy of 0.962 and Kappa of 0.950 for classifying cattle microbiome into the phenotype of a diet supplemented with oil, nitrate, combined (with oil and nitrate) and controls.
Jyotsna Talreja Wassan, Huiru Zheng, Haiying Wang 0001, Fiona Browne, Paul Walsh, Timmy Manning, Richard J. Dewhurst, Rainer Roehe
BIBM3
2019 A practical computerized decision support system for predicting the severity of Alzheimer's disease of an individual
Magda Bucholc, Xuemei Ding, Haiying Wang 0001, David H. Glass, Hui Wang 0001, Girijesh Prasad, Liam P. Maguire, Anthony J. Bjourson, Paula L. McClean, Stephen Todd, David P. Finn, KongFatt Wong-Lin
Expert Syst. Appl.3
2019 A Comprehensive Study on Predicting Functional Role of Metagenomes Using Machine Learning Methods
abstract
"Metagenomics" is the study of genomic sequences obtained directly from environmental microbial communities with the aim to linking their structures with functional roles. The field has been aided in the unprecedented advancement through high-throughput omics data sequencing. The outcome of sequencing are biologically rich data sets. Metagenomic data consisting of microbial species which outnumber microbial samples, lead to the "curse of dimensionality" in datasets. Hence, the focus in metagenomics studies has moved towards developing efficient computational models using Machine Learning (ML), reducing the computational cost. In this paper, we comprehensively assessed various ML approaches to classifying high-dimensional human microbiota effectively into their functional phenotypes. We propose the application of embedded feature selection methods, namely, Extreme Gradient Boosting and Penalized Logistic Regression to determine important microbial species. The resultant feature set enhanced the performance of one of the most popular state-of-the-art methods, Random Forest (RF) over metagenomic studies. Experimental results indicate that the proposed method achieved best results in terms of accuracy, area under the Receiver Operating Characteristic curve (ROC-AUC), and major improvement in processing time. It outperformed other feature selection methods of filters or wrappers over RF and classifiers such as Support Vector Machine (SVM), Extreme Learning Machine (ELM), and k- Nearest Neighbors (k-NN).
Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Huiru Zheng
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 Phylogeny-Aware Deep 1-Dimensional Convolutional Neural Network for the Classification of Metagenomes
Timmy Manning, Jyotsna Talreja Wassan, Cintia C. Palu, Haiying Wang 0001, Fiona Browne, Huiru Zheng, Brian Kelly, Paul Walsh
BIBM4
2018 eZiGait: Toward an AI Gait Analysis And Sssistant System
Graham McCalmont, Philip J. Morrow, Huiru Zheng, Anas Samara, Sara Yasaei, Haiying Wang 0001, Sally I. McClean
BIBM6
2018 Correlation Model Analysis of Nitrogen Addition and Tan Sheep Grazing Effects on Soil Bacterial Community in the Loess Plateau, China
Fujiang Hou, Huiru Zheng, Haiying Wang 0001
BIBM4
2018 PAAM-ML: A novel Phylogeny and Abundance aware Machine Learning Modelling Approach for Microbiome Classification
Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Huiru Zheng
BIBM2
2018 An Integrative Framework for Functional Analysis of Cattle Rumen Microbiomes
Jyotsna Talreja Wassan, Huiru Zheng, Fiona Browne, Jenna Bowen, Paul Walsh, Rainer Roehe, Richard J. Dewhurst, Cintia C. Palu, Brian Kelly, Haiying Wang 0001
BIBM10
2018 Synthetic Optimisation Techniques for Epidemic Disease Prediction Modelling
Terence Fusco, Yaxin Bi, Haiying Wang 0001, Fiona Browne
DATA3
2018 CASNMF: A Converged Algorithm for symmetrical nonnegative matrix factorization
Liping Tian 0001, Ping Luo 0003, Haiying Wang 0001, Huiru Zheng, Fang-Xiang Wu
Neurocomputing3
2017 The modularity of microbial interaction network in healthy human saliva: Stability and specificity
abstract
The human oral cavity is an important habitat of microbes in the human body. It includes the colonization of various microorganisms such as bacteria, archaea, fungi, protozoa and viruses. Although oral diseases have been studied for decades, we have limited understanding of the boundaries of a healthy oral ecosystem and ecological shift toward dysbiosis. Here, we analyzed salivary microbiomes from 268 healthy adults after overnight fasting. The microbiome data set is firstly divided into five sample clusters based on the similarity pattern of microbial abundance. For each cluster, the correlation networks among salivary bacteria are constructed based on an ensemble of six correlations and two dissimilarity measure. The stability and specificity of modularity in the five microbial networks are investigated. The existences of conserved and changing modules were found across five microbial correlation networks.
Xingpeng Jiang, Huiru Zheng, Haiying Wang 0001, Tingting He 0003, Xiaohua Hu 0001
BIBM5
2017 Participatory design-based requirements elicitation involving people living with dementia towards a home-based platform to monitor emotional wellbeing
abstract
We are living in an ageing population with an escalation in chronic illnesses including dementia and other age related diseases. People living with dementia often continue to live at home and are supported by caregivers and next of kin. It is often important to monitor the wellbeing of people living with dementia in order to measure their level of independence and to provide proper support at the time of need as well as supporting their quality of life. Some researchers have focused on monitoring physical wellbeing and activities of daily living (ADL). However, there has been a paucity of research focussed on monitoring mood, affect and the emotional wellbeing of people living with dementia, despite these people experiencing frustration, agitation, depression and social isolation to name but a few known effects. As a result, the SenseCare project aims to build an affective computing platform that uses sensors placed in the home environment to monitor moods, affect and the emotional wellbeing of people living with dementia. This platform is being iteratively designed and will likely use plug-n-play sensors such as passive infrared, wearables and camera technologies to infer emotions from facial expressions, voice intonations and physical behaviour and other modalities. However, it is important to interact iteratively with people living with dementia and their caregivers in order to understand their profound needs. In this study, we report on two focus groups that were conducted to elicit user stories and eventual requirements for the SenseCare platform. Since participatory design involving people living with dementia could bring about unique challenges, we adopted a dyad approach where a caregiver and the person living with dementia participate together in the focus group. This ensures that their needs are fully represented and that consent is fully transparent. In this paper, we report the personal stories elicited during these discussions which will ultimately inform the implementation of the SenseCare platform.
Maurice D. Mulvenna, Huiru Zheng, Raymond R. Bond, Patrick McAllister, Haiying Wang 0001, Ruben Riestra
BIBM5
2017 A metagenomics analysis of rumen microbiome
abstract
Climate change and food security are significant global challenges facing society. The dairy industry is inextricably linked to these challenges as it is concerned with the economies of food production, while acknowledging that it is a major contributor to greenhouse gas production. Action by microbial communities in the rumen is responsible for efficient breakdown of plant matter for food conversion, but a by-product of this action is substantial methane production. Insight into food conversion and methane production in rumen microbiota is possible through metagenomics analysis, which is the analysis of microbial communities and their interactions with the environment. However, metagenomic analysis is hampered by the sheer volume and complexity of data that needs to be processed. This paper presents a bioinformatics pipeline and visualisation platform that facilitates deep analysis of microbial communities, under various conditions in cattle rumen, with the aim of leading to significant impact on probiotic supplement usage, methane production and feed conversion efficiency. This pipeline was developed as part of the EU H2020 MetaPlat project and will pave the way for a more optimal usage of metagenomic datasets, thus reducing the number of animals necessary to be engaged in such studies. This will ensure better and more economic animal welfare, better use of resources and lessen the impact of the dairy industry on climate change.
Paul Walsh, Cintia C. Palu, Brian Kelly, Brendan Lawlor, Jyotsna Talreja Wassan, Huiru Zheng, Haiying Wang 0001
BIBM7
2017 Microbial co-presence and mutual-exclusion networks in the Bovine rumen microbiome
abstract
The recognized significance of rumen microbiome has inspired efforts to examine the composition of rumen microbial communities in a large scale. One of the key research areas is to infer association and dependencies between members of rumen microbial communities through correlation analysis. However, it has been found that due to the compositional nature of data, simply applying correlation-based techniques to the analysis of relative abundance of microbial genes may produce artefactual correlation and loss of information. In an attempt to mitigate the compositional effect on the analysis of rumen microbiome data, this study applied a framework including a compendium of two correlation measures and three dissimilarity metrics that are intrinsically robust to compositionality. Based on the inference of significant positive and negative associations, co-presence and mutual-exclusion networks were constructed. The corresponding modules associated with methane production were identified. The modules are highly enriched with microbial genes associated with methane emissions and encoding enzymes involved in the methane methanogensis pathway. In comparisons to previous studies, our analysis demonstrates that deriving microbial associations based on the correlations between relative abundances may not only lead to missing information but also produce spurious associations.
Haiying Wang 0001, Huiru Zheng, Richard J. Dewhurst, Rainer Roehe
BIBM1
2017 Microbial abundance analysis and phylogenetic adoption in functional metagenomics
abstract
Metagenomics is an unobtrusive science of studying uncultivated microbes sampled directly from an environment, e.g. soil, ocean, air, human body, or animals, etc. Functional metagenomics particularly deals with linking microbes to environmental derivations, such as classifying the role of human gut microbiome into a diseased or non-diseased state. Ongoing research in this area includes analyzing the structure of microbial communities, and relate it to functional analysis. We present an integrative experimental framework for functional metagenomics, including data driven (abundance count of microbial species) and knowledge driven (phylogenetic tree structure) contexts. Our related experiments, indicate that i) feature selection improves the performance of classifying human microbiome samples, ii) the classification of human microbiome remains a challenging problem while incorporating phylogenetic structures. For example, our best accuracy attained on the Costello body site (CBH) dataset with forehead and external ear as body sites, is 89.13 % with a non-phylogenetic model, and 78.26 % with a phylogenetic model. This forms a potential research direction of further exploration of space for incorporating phylogeny in microbial analysis and hence developing integrative computational models for deriving functional phenotypes, based on metagenomic sequencing data.
Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Huiru Zheng
CIBCB2
2017 An Integrative Approach for the Functional Analysis of Metagenomic Studies
Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Paul Walsh, Brian Kelly, Cintia C. Palu, Nina Konstantinidou, Rainer Roehe, Richard J. Dewhurst, Huiru Zheng
ICIC (2)2
2017 Analysis of Organization of the Interactome Using Dominating Sets: A Case Study on Cell Cycle Interaction Networks
abstract
In this study, a minimum dominating set based approach was developed and implemented as a Cytoscape plugin to identify critical and redundant proteins in a protein interaction network. We focused on the investigation of the properties associated with critical proteins in the context of the analysis of interaction networks specific to cell cycle in both yeast and human. A total of 132 yeast genes and 129 human proteins have been identified as critical nodes while 950 in yeast and 980 in human have been categorized as redundant nodes. A clear distinction between critical and redundant proteins was observed when examining their topological parameters including betweenness centrality, suggesting a central role of critical proteins in the control of a network. The significant differences in terms of gene coexpression and functional similarity were observed between the two sets of proteins in yeast. Critical proteins were found to be enriched with essential genes in both networks and have a more deleterious effect on the network integrity than their redundant counterparts. Furthermore, we obtained statistically significant enrichments of proteins that govern human diseases including cancer-related and virus-targeted genes in the corresponding set of critical proteins.
Huiru Zheng, Haiying Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 A network analysis of methane and feed conversion genes in the rumen microbial community
abstract
Metagenomics involves the genetic analysis of microbial DNA extracted from communities in an environment sample. Advent and falling costs of next-generation sequencing technologies has accelerated metagenomics research providing an improved understanding of microbial communities. In this study we investigate if the traits methane production and feed conversion rates in the rumen microbial community overlap with top genes ranked by topological metrics in a co-abundance network. A co-abundance network was constructed from abundance values of 1570 microbial genes in rumen samples of 8 cattle identified in a metagenomics study at the Beef and Sheep Research Centre of Scotland's Rural College. We used 4 different topological measures: Degree Centrality, Betweenness Centrality, Bonacich Power Centrality and PageRank to the network. Using permutation testing, we discovered, methane production trait genes significantly overlapped with top ranked genes obtained using the metrics PageRank and Bonacich Power Centrality. Feed conversion trait genes overlapped with top ranked genes using Bonacich Power Centrality and Betweenness. Furthermore, we observed the top ranked genes from PageRank and Bonacich Power Centrality significantly overlapped with genes involved in the KEGG methane metabolism pathway and ranked highly key methanogenesis genes such as mcrA and fmdB. Identified functional clusters containing most methane and feed conversion genes were also analyzed in terms of overlap with top ranked genes from topological metrics.
Fiona Browne, Haiying Wang 0001, Huiru Zheng, Rainer Roehe, Richard J. Dewhurst, Paul Walsh
BIBM2
2016 Analysis of rumen microbial community in cattle through the integration of metagenomic and network-based approaches
abstract
A better understanding of the composition of rumen microbial communities and the association between host genetic and microbial activities has important applications and implication in bioscience. Being capable of revealing the full extent of microbial gene diversity, metagenomics-based approaches hold great promises in this endeavor. This study investigates the rumen microbial community in cattle through the integration of metagenomic and network-based approaches. Based on the relative abundance of 1570 microbial genes identified in a metagenomics analysis, the co-abundance network was constructed and functional modules of microbial genes were identified. One of the main contributions of this study is to develop a random matrix theory-based approach to automatically determine the correlation threshold used to construct the co-abundance network. It has been shown that the network exhibits a highly modular structure with each of the three main modules well separated. The involvement of KEGG pathways in each module was analysed. A close look at the abundance profiles highlights that Module B is strongly associated with methane emissions while Module C is highly enriched with microbial genes associated with feed conversion efficiency.
Haiying Wang 0001, Huiru Zheng, Fiona Browne, Rainer Roehe, Richard J. Dewhurst, Felix Engel 0002, Matthias L. Hemmje, Paul Walsh
BIBM1
2016 Modelling enteric methane emissions from milking dairy cows with Bayesian networks
abstract
As one of the potent greenhouse gases, methane emission from ruminants has been intensively studied over the past decades. Various regression-based models have been applied to examine factors affecting enteric methane emission. Based on Bayesian networks, this paper proposes an alternative network-based approach to model the relationship among factors affecting enteric methane emissions from milking cows. It was evaluated on the dataset consisting of 934 milking dairy cows collected at Agri-Food and Biosciences Institute, Northern Ireland. The preliminary results demonstrated that the proposed model has a great potential to capture the complex relationship among factors and establish causal influence among predictors. To the best of our knowledge, this is the first study to use Bayesian networks to model causal influence among factors associated with enteric methane emission from milking cows.
Huiru Zheng, Haiying Wang 0001, Tianhai Yan
BIBM2
2015 Assessment of gait patterns of chronic low back pain patients: A smart mobile phone based approach
abstract
Chronic low back pain is a common and costly condition and has been shown to affect gait. This paper describes the use of gait analysis as measured by a smart phone in a group of chronic low back pain subjects. Reliability of features extracted from the smart phone sensors was investigated using a mutual information based minimum redundancy and maximum relevance feature selection method to identify a key feature set related to lower back pain. This analysis was carried out using a KStar classification model. Results indicate the feasibility of reducing gait features to 6 key components while still achieving very promising classification accuracy (92.50%). The results also demonstrated that it is feasible to use a smart mobile phone in gait tele-monitoring and tele-assessment suggesting potential as both a prognostic and potential treatment outcome. In addition, we show that predicting context such as age and gender using smart mobile phones is achievable, which has potential to provide personalised services and context-related monitoring and intervention.
Herman Chan, Huiru Zheng, Haiying Wang 0001, Dave Newell
BIBM3
2015 Integrating omics data for identifying disease subtypes: A multiplex network-based approach
abstract
It is widely acknowledged that technologies centred on the integration of omics data could play an important role in capturing heterogeneity of phenotypes and identifying disease subtypes. This paper proposed a multiplex network-based approach for integrative analysis of heterogeneous omics data. It represents a useful alternative network-based solution to the problem and a significant step forward to the methods in which each type of data is treated independently. It has been tested on the identification of the subtypes of glioblastoma multiforme. Results obtained have shown that it can achieve comparable performance in comparison to state-of-the-art techniques. The proposed methodology has several advantages. It provides a flexible platform to integrate different types of patient data, potentially from multiple sources, allowing discovering complex disease patterns with multiple facets.
Haiying Wang 0001, Huiru Zheng
BIBM1
2015 Thermal sensor based multi-occupancy motion tracking and visualisation in smart environments
abstract
A smart environment is a physical space where smart devices/technologies are applied to continuously sense the occupant's daily living activities or health condition. Recent years have witnessed the application of smart environments to support independent living for elderly people or for people with chronic conditions. Nevertheless, most of the projects have been exclusively designed to support instances of single occupancy. In an effort to move beyond the scenario of a single occupant within a smart environment, this paper investigates the feasibility of using thermal sensors to detect the presence of multi-occupants and to track and visualise their motions. Results indicate that the use of thermals sensor can detect multi-occupancy, including moving subjects in addition to static subjects. The system developed demonstrated the ability to track up to 5 occupants, although mistrackings were found to occur when the individuals were apart after they stayed closely together.
Huiru Zheng, Haiying Wang 0001, Jonathan Synnott, Chris D. Nugent, Paul Jeffers 0001
BIBM3
2015 Identification of Protein Complexes from Tandem Affinity Purification/Mass Spectrometry Data via Biased Random Walk
abstract
Systematic identification of protein complexes from protein-protein interaction networks (PPIs) is an important application of data mining in life science. Over the past decades, various new clustering techniques have been developed based on modelling PPIs as binary relations. Non-binary information of co-complex relations (prey/bait) in PPIs data derived from tandem affinity purification/mass spectrometry (TAP-MS) experiments has been unfairly disregarded. In this paper, we propose a Biased Random Walk based algorithm for detecting protein complexes from TAP-MS data, resulting in the random walk with restarting baits (RWRB). RWRB is developed based on Random walk with restart. The main contribution of RWRB is the incorporation of co-complex relations in TAP-MS PPI networks into the clustering process, by implementing a new restarting strategy during the process of random walk. Through experimentation on un-weighted and weighted TAP-MS data sets, we validated biological significance of our results by mapping them to manually curated complexes. Results showed that, by incorporating non-binary, co-membership information, significant improvement has been achieved in terms of both statistical measurements and biological relevance. Better accuracy demonstrates that the proposed method outperformed several state-of-the-art clustering algorithms for the detection of protein complexes in TAP-MS data.
Bingjing Cai, Haiying Wang 0001, Huiru Zheng, Hui Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 An integrative network-driven pipeline for the prioritization of Alzheimer's disease genes
abstract
Large-scale, high-throughput technologies and genome-wide studies have been pivotal in the identification of disease-gene candidates from patient cohorts. Output from these studies often result in gene candidate lists which are large in size. Therefore, there is a pressing need for computational tools to integrate heterogeneous data and prioritize disease-gene candidates for further experimental investigation. To address this need, we propose a computational pipeline for the prioritization of disease-gene candidates. Our pipeline integrates diverse heterogeneous data including: gene-expression, protein-protein interaction network, ontology-based similarity and betweenness measures. Furthermore, we incorporate tissue-specific gene expression data into the evaluation section of our approach. The pipeline was applied to prioritize Alzheimer's Disease (AD) genes, whereby a list of 31 prioritized genes was generated. This approach correctly identified key AD susceptible genes: INPP5D and PSEN1. Biological process enrichment analysis revealed the prioritized genes are modulated in AD pathogenesis including: regulation of neurogenesis and generation of neurons. KEGG pathway analysis identified significant hub involvement in the Neurotrophin signaling and Huntington Disease pathways. Furthermore, our evaluation demonstrated a relatively high predictive performance (AUC: 0.73) when classifying AD and normal gene expression profiles from individuals using leave-one-out cross validation. This work provides a foundation for future investigation of diverse heterogeneous data integration for disease-gene prioritization.
Fiona Browne, Haiying Wang 0001, Huiru Zheng
BIBM2
2014 Minimum dominating sets in cell cycle specific protein interaction networks
abstract
Recently, scientists start to examine the dynamics of biological networks from a control theory perspective. Based on the determination of minimum dominating sets (MDSets), this paper investigated the properties associated with MDSet proteins in the context of the analysis of protein interaction networks specific to the yeast cell cycle. Statistically significant differences between MDSet and non-MDSet proteins were observed in terms of topological features, Gene Ontology-driven semantic similarities, and the number of protein domains associated with each protein. However, unlike previous studies, MDSet proteins were found to be enriched with essential genes. Furthermore, we constructed and analyzed a PPI network specific to the human cell cycle and highlighted that the distinction between MDSet and non-MDSet proteins is far more complex than that observed in yeast. The system used to determine a minimum dominating set in a protein interaction network was implemented as a user-friendly Java-based plugin for Cytoscape.
Haiying Wang 0001, Huiru Zheng, Fiona Browne
BIBM1
2014 Night optimised care technology for users needing assisted lifestyles
abstract
There is growing interest in the development of ambient assisted living services to increase the quality of life of the increasing proportion of the older population. We report on the Night Optimised Care Technology for UseRs Needing Assisted Lifestyles project, which provides specialised night time support to people at early stages of dementia. This article explains the technical infrastructure, the intelligent software behind the decision-making driving the system, the software development process followed, the interfaces used to interact with the user, and the findings and lessons of our user-centred approach.
Juan Carlos Augusto, Maurice D. Mulvenna, Huiru Zheng, Haiying Wang 0001, Suzanne Martin, Paul J. McCullagh, Jonathan G. Wallace
Behav. Inf. Technol.4
2014 Organized Modularity in the Interactome: Evidence from the Analysis of Dynamic Organization in the Cell Cycle
abstract
The organization of global protein interaction networks (PINs) has been extensively studied and heatedly debated. We revisited this issue in the context of the analysis of dynamic organization of a PIN in the yeast cell cycle. Statistically significant bimodality was observed when analyzing the distribution of the differences in expression peak between periodically expressed partners. A close look at their behavior revealed that date and party hubs derived from this analysis have some distinct features. There are no significant differences between them in terms of protein essentiality, expression correlation and semantic similarity derived from gene ontology (GO) biological process hierarchy. However, date hubs exhibit significantly greater values than party hubs in terms of semantic similarity derived from both GO molecular function and cellular component hierarchies. Relating to three-dimensional structures, we found that both single- and multi-interface proteins could become date hubs coordinating multiple functions performed at different times while party hubs are mainly multi-interface proteins. Furthermore, we constructed and analyzed a PPI network specific to the human cell cycle and highlighted that the dynamic organization in human interactome is far more complex than the dichotomy of hubs observed in the yeast cell cycle.
Haiying Wang 0001, Huiru Zheng
IEEE ACM Trans. Comput. Biol. Bioinform.1
2013 Detection of protein complexes in protein interaction networks is improved through network-driven functional homogeneity analysis
abstract
The detection of biologically meaningful clusters in protein interaction networks is crucial in systems biology. Among its applications, it can enable the identification of protein complexes. Notwithstanding significant advances, the detection of meaningful clusters faces important challenges, including the need to aid researchers in the prioritization of hundreds or even thousands of clusters. To address this need, we developed a method for the prioritization of network clusters based on the analysis of their functional homogeneity, Horn. Based on Horn scores, clustering results can be statistically ranked and attention directed toward clusters that are more likely to be biologically meaningful. We tested it on a global human protein-protein interaction network and four network clustering algorithms. Our method substantially reduced the space of potentially spurious clusters. Furthermore, we evaluated its protein complex detection capability on an independent reference dataset of protein complexes. Irrespectively of clustering approach, our approach improved protein complex identification capacity.
Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
BIBM1
2013 Correlating adverse drug reactions with biological pathways in humans
abstract
It has been well recognized that adverse drug reactions (ADRs) are a significant cause of morbidity and mortality. There is a growing interest in investigating biological pathways involved in cellular response to drugs. Based on examining the co-occurrence of drugs in pathway activity and ADR profiles, in this paper, we propose a new method to explore the relationship between biological pathways and ADRs at a large scale. Using sparse canonical correlation analysis of 495 drugs with two profiles for 173 pathways and 1385 ADRs, a total of 80 correlated sets of pathways and ADRs were extracted. To evaluate the performance of our method, extracted correlated components were used to retrieve known ADR profiles from drug pathway profiles using a 5-fold cross validation. A relatively high prediction performance (AUC: 0.881) was achieved. This work provides a foundation for future investigation of ADRs in the context of biological pathways under different conditions.
Huiru Zheng, Haiying Wang 0001, Hua Xu 0001, Zhongming Zhao, Francisco Azuaje
BIBM2
2013 Assessing Gait Patterns of Healthy Adults Climbing Stairs Employing Machine Learning Techniques
abstract
So far, stair climbing has not been studied as extensively as gait has, although the significance of the prevention of falling on stairs has been well recognized. Based on acceleration data taken from 25 healthy subjects climbing up and down a set of 13 stairs with an accelerometer placed on the lumbo-sacral joint, this paper aims to assess gait patterns of younger and older adults climbing stairs using a machine learning approach. A total of 14 gait features were extracted and analyzed. The performance of six representative classification models: Multilayer Perceptron (MLP), KStar, Support Vector Machine (SVM), Naïve Bayesian (NB), C4.5 Decision Trees, and Random Forests were evaluated in terms of their ability to discriminate between younger and older adults climbing up- and downstairs. MLP was found to provide the highest accuracy for classification. Accuracy of 95.7% was found for classifying a subject walking either up or down the stairs and an accuracy of 80.6% for classifying whether the subject was younger or older. An evaluation of individual features showed poor performance of classification for younger and older subjects climbing up- and downstairs, and in most cases failed to distinguish between the two classes. To access which set of features derived from a triaxial accelerometer can better describe the performance differences between younger and older adults climbing up- and downstairs, two feature selection algorithms, sequential feature selection and correlation-based feature selection, were implemented. Results show that 10 features derived from correlation-based feature selection were able to produce a 96.8% accuracy for classification between subjects climbing up and down. A subset of seven features achieved a performance of 84.9% accuracy for classification between younger and older subjects.
Herman Chan, Mingjing Yang 0001, Haiying Wang 0001, Huiru Zheng, Sally I. McClean, Roy Sterritt, Ruth E. Mayagoitia
Int. J. Intell. Syst.3
2013 Editorial: Special Issue on "Recent Advances in Intelligent Techniques"
Yongmin Li 0001, Ning Xiong 0001, Haiying Wang 0001, Lipo Wang 0001
Int. J. Intell. Syst.3
2012 Incorporating semantic similarity into clustering process for identifying protein complexes from Affinity Purification/Mass Spectrometry data
abstract
This paper presents a framework for incorporating semantic similarities in the detection of protein complexes from Affinity Purification/Mass Spectrometry (AP-MS) data. AP-MS data is modeled as a bipartite network, where one set of nodes consist of bait proteins and the other set are prey proteins. Pair-wise similarities of bait proteins are computed by combining similarities based on topological features and functional semantic similarities. A hierarchical clustering algorithm is then applied to obtain `seed clusters' consisting of bait proteins. Starting from these `seed' clusters, an expansion process is developed to recruit prey proteins which are significantly associated with bait proteins, to produce final sets of identified protein complexes. In the application to real AP-MS datasets, we validate biological significance of predicted protein complexes by using curated protein complexes. Six statistical metrics have been applied. Results show that by integrating semantic similarities into the clustering process, the accuracy of identifying complexes has been greatly improved. Meanwhile, clustering results obtained by the proposed framework are better than those from several existent clustering methods.
Bingjing Cai, Haiying Wang 0001, Huiru Zheng, Hui Wang 0001
BIBM2
2012 Drug-target network in myocardial infarction: A structural analysis
abstract
The identification of drug-target interactions is a crucial step in the drug-discovery process. It has been suggested that drug-target interactions are driven by drug-domain interactions. Based on the integration of two recently published datasets, i.e., Drug-target interactions in myocardial infarction (My-DTome) and drug-domain interaction network, this paper reports the association between drugs and protein domains in the context of myocardial infarction (MI). A MI drug-domain interaction network, My-DDome, was constructed. The functional similarity between domains based on their Gene Ontology (GO) annotations was estimated. The association between domains and therapeutic effects was investigated. Lists of GO annotations and Anatomical Therapeutic Chemical classification (ATC) codes highly enriched in My-DDome were identified. We show that drugs acting on blood and blood forming organs (ATC code B) and sensory organs (ATC code S) are significantly enriched in My-DDome (p <; 0.000001). Top enriched GO terms include GO:0003824 (catalytic activity), GO:0008152 (metabolic process) and GO:0030170 (pyridoxal phosphate binding). By incorporating protein domain information into My-DTome, more detailed insights into the interplay between drugs, their known targets and seemingly unrelated proteins are provided.
Haiying Wang 0001, Huiru Zheng, Francisco Azuaje, Xing-Ming Zhao
BIBM1
2011 An improved random walk based clustering algorithm for community detection in complex networks
abstract
In recent years, there is an increasing interest in the research community in finding community structure in complex networks. The networks are usually represented as graphs, and the task is usually cast as a graph clustering problem. Traditional clustering algorithms and graph partitioning algorithms have been applied to this problem. New graph clustering algorithms have also been proposed. Random walk based clustering, in which the similarities between pairs of nodes in a graph are usually estimated using random walk with restart (RWR) algorithm, is one of the most popular graph clustering methods. Most of these clustering algorithms only find disjoint partitions in networks; however, communities in many real-world networks often overlap to some degree. In this paper, we propose an efficient clustering method based on random walks for discovering communities in graphs. The proposed method makes use of network topology and edge weights, and is able to discover overlapping communities. We analyze the effect of parameters in the proposed method on clustering results. We evaluate the proposed method on real world social networks that are well documented in the literature, using both topological-based and knowledge-based evaluation methods. We compare the proposed method to other clustering methods including recently published Repeated Random Walks, and find that the proposed method achieves better precision and accuracy values in terms of six statistical measurements including both data-driven and knowledge-driven evaluation metrics.
Bingjing Cai, Haiying Wang 0001, Huiru Zheng, Hui Wang 0001
SMC2
2011 Systems-based biological concordance and predictive reproducibility of gene set discovery methods in cardiovascular disease
Francisco Azuaje, Huiru Zheng, Anyela Camargo, Haiying Wang 0001
J. Biomed. Informatics4
2011 Recent progress in natural computation and knowledge discovery: an ICNC'09-FSKD'09 special issue
Haiying Wang 0001, Yixin Chen 0001, Hepu Deng, Lipo Wang 0001
Soft Comput.1
2010 seGOsa: Software environment for gene ontology-driven similarity assessment
abstract
In recent years there has been a growing trend towards the adoption of ontologies to support comprehensive, large-scale functional genomics research. This paper introduces seGOsa, a user-friendly cross-platform system to support large-scale assessment of Gene Ontology (GO)-driven similarity among gene products. Using information-theoretic approaches, the system exploits both topological features of the GO (i.e., between-term relationships in the hierarchy) and statistical features of the model organism databases annotated to the GO (i.e., term frequency) to assess functional similarity among gene products. Based on the assumption that the more information two terms share in common, the more similar they are, three GO-driven similarity measures (Resnik's, Lin's and Jiang's metrics) have been implemented to measure between-term similarity within each of the GO hierarchies. Meanwhile, seGOsa offers two approaches (simple and highest average similarity) to assessing the similarity between gene products based on the aggregation of between-term similarities. The program is freely available for non-profit use on request from the authors.
Huiru Zheng, Francisco Azuaje, Haiying Wang 0001
BIBM3
2010 Ontology- and graph-based similarity assessment in biological networks
abstract
SUMMARY: A standard systems-based approach to biomarker and drug target discovery consists of placing putative biomarkers in the context of a network of biological interactions, followed by different 'guilt-by-association' analyses. The latter is typically done based on network structural features. Here, an alternative analysis approach in which the networks are analyzed on a 'semantic similarity' space is reported. Such information is extracted from ontology-based functional annotations. We present SimTrek, a Cytoscape plugin for ontology-based similarity assessment in biological networks. AVAILABILITY: http://rosalind.infj.ulst.ac.uk/SimTrek.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
Bioinform.1
2010 Integration of Gene Ontology-based similarities for supporting analysis of protein-protein interaction networks
Haiying Wang 0001, Huiru Zheng, Fiona Browne, David H. Glass, Francisco Azuaje
Pattern Recognit. Lett.1
2009 Assessing the impact of network depth on the analysis of PPI networks: A case study
abstract
Recent years have seen a growing interest in the incorporation of protein-protein interaction (PPI) networks to support functional genomic research. Often a default depth is assumed by network inference software. This case study considers the impact of network depth on the analysis of PPI networks using seven proteins known to be relevant to heart failure as inputs into the analysis. This paper analyses how the characteristics of a PPI network vary according to the level examined, suggesting that the investigation of network topology is an essential first step in PPI analysis. The classification of nodes, in terms of degree and betweenness centrality, within the network is also considered. The effect of network depth is also proved to be significant in the identification of potentially essential proteins with large connectivity and/or high betweenness centrality values.
Jaine K. Blayney, Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
CIBCB2
2008 Signature genes in human heart failure based on gene expression analysis: Can we identify a unique set?
abstract
Dilated Cardiomyopathy is one of leading courses of heart failure. Recent advances in microarray technology have promised significant advantages in understanding the molecular mechanisms underlying dilated cardiomyopathy and heart failure. Several microarray studies have successfully yielded a set of signature genes associated with heart failure. However, it has been found that the overlap of these heart failure associated genes derived from different experiments is very small. Based on the analysis of two publicly available microarray datasets associated with heart failure with three types of machine learning and statistical prediction models, this paper explores this phenomenon. We found that there is no unique set of genes associated with heart failure. Many sets of genes can achieve very high prediction accuracy. In order to identify biomarkers in human heart failure, it may not be sufficient to just focus a certain number of top genes. Such main candidates should be chosen from the much longer list of genes.
Haiying Wang 0001, Huiru Zheng
BIBE1
2008 Reassessing the limit of data integration for the prediction of protein-protein interactions in Saccharomyces cerevisiae
abstract
This paper investigates the integration of functional genomic data for the prediction of protein-protein interactions (PPI) in Saccharomyces cerevisiae. A previous benchmark study observed a marginal increase in predictive power when integrating diverse features. Classification performance was evaluated using the Receiver Operating Characteristic (ROC) curve. In this study we propose the implementation of a likelihood ratio based Bayesian classifier to reassess the limits of genomic integration. The classifier combines seven genomic features ranging from co-expression to essentiality. Due to the imbalance of the dataset in this study, ROC curves may present an overly optimistic view of the classification performance. We use the true positive/false positive (TP/FP) rate and sensitivity as comparative predictive measures to the ROC curve. Predicted interactions are verified using a Gold Standard constructed from the Munich Database of Interacting Proteins Complex Catalogue. Using the measures TP/FP and sensitivity, a clear increase in classification performance was observed with the integration of features. This framework could be extended to the analysis of PPI in more complex organisms such as Drosophila melanogaster and Homo sapiens.
Fiona Browne, Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
CIBCB2
2008 An Improved Support Vector Machine for the Classification of Imbalanced Biological Datasets
Haiying Wang 0001, Huiru Zheng
ICIC (1)1
2008 Poisson-Based Self-Organizing Neural Networks for Pattern Discovery
Haiying Wang 0001, Huiru Zheng
ICIC (1)1
2008 Improving Pattern Discovery and Visualization of SAGE Data Through Poisson-Based Self-Adaptive Neural Networks
abstract
Serial analysis of gene expression (SAGE) allows a detailed, simultaneous analysis of thousands of genes without the need for prior, complete gene sequence information. However, due to its inherent complexity and the lack of complete structural and function knowledge, mining vast collections of SAGE data to extract useful knowledge poses great challenges to traditional analytical techniques. Moreover, SAGE data are characterized by a specific statistical model that has not been incorporated into traditional data analysis techniques. The analysis of SAGE data requires advanced, intelligent computational techniques, which consider the underlying biology and the statistical nature of SAGE data. By addressing the statistical properties demonstrated by SAGE data, this paper presents a new self-adaptive neural network, Poisson-based growing self-organizing map (PGSOM), which implements novel weight adaptation and neuron growing strategies. An empirical study of key dynamic mechanisms of PGSOM is presented. It was tested on three datasets, including synthetic and experimental SAGE data. The results indicate that, in comparison to traditional techniques, the PGSOM offers significant advantages in the context of pattern discovery and visualization in SAGE data. The pattern discovery and visualization platform discussed in this paper can be applied to other problem domains where the data are better approximated by a Poisson distribution.
Huiru Zheng, Haiying Wang 0001, Francisco Azuaje
IEEE Trans. Inf. Technol. Biomed.2
2008 Integration of Genomic Data for Inferring Protein Complexes from Global Protein-Protein Interaction Networks
abstract
Protein-protein interactions (PPIs) play crucial roles in virtually every aspect of cellular function within an organism. One important objective of modern biology is the extraction of functional modules, such as protein complexes from global protein interaction networks. This paper describes how seven genomic features and four experimental interaction data sets were combined using a Bayesian-networks-based data integration approach to infer PPI networks in yeast. Greater coverage and higher accuracy were achieved than in previous high-throughput studies of PPI networks in yeast. A Markov clustering algorithm was then used to extract protein complexes from the inferred protein interaction networks. The quality of the computed complexes was evaluated using the hand-curated complexes from the Munich Information Center for Protein Sequences database and gene-ontology-driven semantic similarity. The results indicated that, by integrating multiple genomic information sources, a better clustering result was obtained in terms of both statistical measures and biological relevance.
Huiru Zheng, Haiying Wang 0001, David H. Glass
IEEE Trans. Syst. Man Cybern. Part B2
2007 Supervised Statistical and Machine Learning Approaches to Inferring Pairwise and Module-Based Protein Interaction Networks
abstract
This paper evaluates three classification techniques: Naive Bayesian (NB), multilayer perceptron (MLP) and K-nearest neighbour (KNN) that integrate diverse, large-scale functional data to infer pairwise (PW) and module-based (MB) interaction networks in Saccharomyces cerevisiae. Existing multi-source functional data from S. cerevisiae were merged and transformed to construct MB datasets. The results indicate that selection of a classifier depends upon the specific PPI classification problem. Feature integration and encoding methods proposed significantly impact the predictive performance of the classifiers. Generation of PPI maps for S. cerevisiae and beyond will be improved with new, high-quality, large-scale datasets with increased interactome coverage and the integration of classification methods.
Fiona Browne, Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
BIBE2
2007 Cluster Analysis of Regulatory Sequences with a Log Likelihood Ratio Statistics-based Similarity Measure
abstract
Upstream regions in the DNA sequence are characterized by the presence of short regulatory motifs, which function as target binding sites for transcription factors. Finding two genes with common motifs in their regulatory regions may aid users in identifying co-regulated genes or inferring regulatory modules. By modelling pattern occurrences in the regulatory regions with Poisson statistics, this paper presents a log likelihood ratio statistics-based distance measure to calculate pair-wise similarities between sequences. To perform cluster analysis of regulatory sequences, this paper introduces two clustering algorithms on the basis of the incorporation of the log likelihood ratio statistics-based distance into hierarchical clustering and Self-Organizing Map. The proposed approach has been tested on a synthetic dataset and a real biological example. The results indicate that, in comparison to traditional distance functions, the log likelihood ratio statistics-based similarity measure offers considerable improvements in the process of regulatory sequence-based gene classification.
Huiru Zheng, Haiying Wang 0001, Jinglu Hu
BIBE2
2007 homeML - An Open Standard for the Exchange of Data Within Smart Environments
Chris D. Nugent, Dewar D. Finlay, Richard J. Davies, Haiying Wang 0001, Huiru Zheng, Josef Hallberg, Kåre Synnes, Maurice D. Mulvenna
ICOST4
2007 Self-Adaptive Neural Networks Based on a Poisson Approach for Knowledge Discovery
Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
IJCAI1
2007 Poisson-Based Self-Organizing Feature Maps and Hierarchical Clustering for Serial Analysis of Gene Expression Data
abstract
Serial analysis of gene expression (SAGE) is a powerful technique for global gene expression profiling, allowing simultaneous analysis of thousands of transcripts without prior structural and functional knowledge. Pattern discovery and visualization have become fundamental approaches to analyzing such large-scale gene expression data. From the pattern discovery perspective, clustering techniques have received great attention. However, due to the statistical nature of SAGE data (i.e., underlying distribution), traditional clustering techniques may not be suitable for SAGE data analysis. Based on the adaptation and improvement of Self-Organizing Maps and hierarchical clustering techniques, this paper presents two new clustering algorithms, namely, PoissonS and PoissonHC, for SAGE data analysis. Tested on synthetic and experimental SAGE data, these algorithms demonstrate several advantages over traditional pattern discovery techniques. The results indicate that, by incorporating statistical properties of SAGE data, PoissonS and PoissonHC, as well as a hybrid approach (neuro-hierarchical approach) based on the combination of PoissonS and PoissonHC, offer significant improvements in pattern discovery and visualization for SAGE data. Moreover, a user-friendly platform, which may improve and accelerate SAGE data mining, was implemented. The system is freely available on request from the authors for nonprofit use.
Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
IEEE ACM Trans. Comput. Biol. Bioinform.1
2006 Computational Approaches to Supporting Large-Scale Analysis of Photoreceptor-Enriched Gene Expression
abstract
Retinal photoreceptor cells are responsible for light detection and phototransduction. The understanding of molecular mechanisms regulating photoreceptor gene expression during retinal development may have important implications in clinical neuroscience. Using self-adaptive neural networks and pattern validation statistical tools, this paper explores large-scale analysis of photoreceptor gene expression. Based on the analysis of data generated by serial analysis of gene expression (SAGE) in the developing mouse retina, significant expression patterns for the in silico detection of photoreceptor-enriched genes were revealed. This study demonstrates how machine learning and statistical techniques may be effectively combined to detect key complex relationships encoded in SAGE data. Such approaches may support inexpensive functional predictions prior to the application of experimental methodologies.
Haiying Wang 0001, Huiru Zheng, Francisco Azuaje
CBMS1
2006 Machine learning approaches to supporting the identification of photoreceptor-enriched genes based on expression data
abstract
BACKGROUND: Retinal photoreceptors are highly specialised cells, which detect light and are central to mammalian vision. Many retinal diseases occur as a result of inherited dysfunction of the rod and cone photoreceptor cells. Development and maintenance of photoreceptors requires appropriate regulation of the many genes specifically or highly expressed in these cells. Over the last decades, different experimental approaches have been developed to identify photoreceptor enriched genes. Recent progress in RNA analysis technology has generated large amounts of gene expression data relevant to retinal development. This paper assesses a machine learning methodology for supporting the identification of photoreceptor enriched genes based on expression data. RESULTS: Based on the analysis of publicly-available gene expression data from the developing mouse retina generated by serial analysis of gene expression (SAGE), this paper presents a predictive methodology comprising several in silico models for detecting key complex features and relationships encoded in the data, which may be useful to distinguish genes in terms of their functional roles. In order to understand temporal patterns of photoreceptor gene expression during retinal development, a two-way cluster analysis was firstly performed. By clustering SAGE libraries, a hierarchical tree reflecting relationships between developmental stages was obtained. By clustering SAGE tags, a more comprehensive expression profile for photoreceptor cells was revealed. To demonstrate the usefulness of machine learning-based models in predicting functional associations from the SAGE data, three supervised classification models were compared. The results indicated that a relatively simple instance-based model (KStar model) performed significantly better than relatively more complex algorithms, e.g. neural networks. To deal with the problem of functional class imbalance occurring in the dataset, two data re-sampling techniques were studied. A random over-sampling method supported the implementation of the most powerful prediction models. The KStar model was also able to achieve higher predictive sensitivities and specificities using random over-sampling techniques. CONCLUSION: The approaches assessed in this paper represent an efficient and relatively inexpensive in silico methodology for supporting large-scale analysis of photoreceptor gene expression by SAGE. They may be applied as complementary methodologies to support functional predictions before implementing more comprehensive, experimental prediction and validation methods. They may also be combined with other large-scale, data-driven methods to facilitate the inference of transcriptional regulatory networks in the developing retina. Furthermore, the methodology assessed may be applied to other data domains.
Haiying Wang 0001, Huiru Zheng, David Simpson, Francisco Azuaje
BMC Bioinform.1
2005 An Ontology-Driven Clustering Method for Supporting Gene Expression Analysis
abstract
The gene ontology (GO) is an important knowledge resource for biologists and bioinformaticians. This paper explores the integration of similarity information derived from GO into clustering-based gene expression analysis. A system that integrates GO annotations, similarity patterns and expression data in yeast is assessed. In comparison with a clustering model based only on expression data correlation, the proposed framework not only produces consistent results, but also it offers alternative, potentially meaningful views of the biological problem under study. Moreover, it provides the basis for developing other automated, knowledge-driven data mining systems in this and related application areas.
Haiying Wang 0001, Francisco Azuaje, Olivier Bodenreider
CBMS1
2005 Non-linear mapping for exploratory data analysis in functional genomics
abstract
BACKGROUND: Several supervised and unsupervised learning tools are available to classify functional genomics data. However, relatively less attention has been given to exploratory, visualisation-driven approaches. Such approaches should satisfy the following factors: Support for intuitive cluster visualisation, user-friendly and robust application, computational efficiency and generation of biologically meaningful outcomes. This research assesses a relaxation method for non-linear mapping that addresses these concerns. Its applications to gene expression and protein-protein interaction data analyses are investigated. RESULTS: Publicly available expression data originating from leukaemia, round blue-cell tumours and Parkinson disease studies were analysed. The method distinguished relevant clusters and critical analysis areas. The system does not require assumptions about the inherent class structure of the data, its mapping process is controlled by only one parameter and the resulting transformations offer intuitive, meaningful visual displays. Comparisons with traditional mapping models are presented. As a way of promoting potential, alternative applications of the methodology presented, an example of exploratory data analysis of interactome networks is illustrated. Data from the C. elegans interactome were analysed. Results suggest that this method might represent an effective solution for detecting key network hubs and for clustering biologically meaningful groups of proteins. CONCLUSION: A relaxation method for non-linear mapping provided the basis for visualisation-driven analyses using different types of data. This study indicates that such a system may represent a user-friendly and robust approach to exploratory data analysis. It may allow users to gain better insights into the underlying data structure, detect potential outliers and assess assumptions about the cluster composition of the data.
Francisco Azuaje, Haiying Wang 0001, Alban Chesneau
BMC Bioinform.2
2004 Gene expression correlation and gene ontology-based similarity: an assessment of quantitative relationships
abstract
Genome Database were analyzed to calculate functional similarity of gene products. Three methods for measuring similarity (including a distance-based approach) were implemented. Significant, quantitative relationships between similarity and expression correlation of pairs of genes were detected. Using a known gene expression dataset in yeast, this study compared more than three million pairs of gene products on the basis of these functional properties. Highly correlated genes exhibit strong similarity based on information originating from the gene ontology taxonomies. Such a similarity is significantly stronger than that observed between weakly correlated genes. This study supports the feasibility of applying gene ontology-driven similarity methods to functional prediction tasks, such as the validation of gene expression analyses and the identification of false positives in protein interaction studies.
Haiying Wang 0001, Francisco Azuaje, Olivier Bodenreider, Joaquín Dopazo
CIBCB1
2004 An integrative and interactive framework for improving biomedical pattern discovery and visualization
abstract
Recent progress in medical sciences has led to an explosive growth of data. Due to its inherent complexity and diversity, mining such volumes of data to extract relevant knowledge represents an enormous challenge and opportunity. Interactive pattern discovery and visualization systems for biomedical data mining have received relatively little attention. Emphasis has been traditionally placed on automation and supervised classification problems. Based on self-adaptive neural networks and pattern-validation statistical tools, this paper presents a user-friendly platform to support biomedical pattern discovery and visualization. It has been tested on several types of biomedical data, such as dermatology and cardiology data sets. The results indicate that in comparison to traditional techniques, such as Kohonen Maps, this platform may significantly improve the effectiveness and efficiency of pattern discovery and classification tasks, including problems described by several classes. Furthermore, this study shows how the combination of graphical and statistical tools may make these patterns more meaningful.
Haiying Wang 0001, Francisco Azuaje, Norman D. Black
IEEE Trans. Inf. Technol. Biomed.1