VLDB 2026 Research / reviewers in the wild / expert
Huiru Zheng
dblp:27/2027 · also Hui-Ru Zheng, Huiru Jane Zheng
· DBLP profile ↗
124ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0001-7648-8709ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 106 · 9 first-author · 20 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AURORA: An Adaptive Multi-granularity Graph Learning Framework for Drug Repositioning
Yidan Zhang 0001, Lei Duan, Huiru Zheng, Jiaxuan Xu 0001 |
DASFAA (3) | 3 |
| 2026 | D.R.E.A.M: diabetes risk via explainable AI modelingabstractAbstract Most machine learning models for diabetes prediction rely on small, homogeneous datasets and fixed thresholds, producing binary outputs with limited clinical utility. These approaches lack generalizability, probabilistic awareness, and interpretability, which are essential for real-world healthcare adoption. We present Diabetes Risk via Explainable AI Modeling (D.R.E.A.M.), a framework for Type 2 diabetes mellitus (T2DM) risk prediction that delivers continuous, calibrated probabilities with transparent explanations. D.R.E.A.M. integrates two complementary datasets (PIMA and BRFSS 2015) after excluding gestational diabetes cases, applies clinically guided feature engineering and class balancing, and trains ensemble models (Random Forest, XGBoost, LightGBM). Decision thresholds are optimized using precision–recall curve analysis rather than default cutoffs, enabling clinically meaningful stratification. Model interpretability is achieved through SHapley Additive exPlanations (SHAP), providing both global and patient-level insights. All models achieved Area Under the Curves above 0.83 and F1-scores of 0.78, with Random Forest offering the best balance of sensitivity (recall = 0.89 at an optimized threshold of 0.389) and interpretability. SHAP confirmed the contribution of both physiological and behavioral factors, including glucose, BMI, blood pressure, cholesterol, and physical activity. Accessible via a lightweight web interface, D.R.E.A.M. provides real-time, explainable risk scores to support personalized preventive strategies. In summary, D.R.E.A.M. advances beyond conventional post-hoc explainability by integrating calibrated probabilistic predictions, PRC-based thresholding, and direct clinician-facing deployment. This combination transforms it from a research prototype into a transparent and clinically actionable decision support system. Domenico Rossi, Alessia Auriemma Citarella, Fabiola De Marco, Luigi Di Biasi, Huiru Zheng, Genny Tortora |
Multim. Tools Appl. | 5 |
| 2026 | From ECG to identity recognition: a scalable, image-based approach to biometric authenticationabstractAbstract Biometric identification based on the electrocardiogram (ECG) is gaining attention as a secure and reliable approach to healthcare authentication, employing the unique physiological patterns detected in the ECG signal. Traditional approaches often depend on raw waveform analysis or the extraction of fiducial points, both of which are computationally intensive and challenging to implement in real-time systems. This work presents CardioIdNet, a lightweight convolutional neural network designed to perform biometric identification directly from ECG images, eliminating the need for complex signal preprocessing steps. ECG recordings from 21 subjects in the MIT-BIH arrhythmia database were segmented and converted to grayscale waveform plots, generating a comprehensive well-suited dataset for image-based deep learning classification. The CardioIdNet architecture consists of convolutional and pooling layers for hierarchical feature extraction, followed by fully connected layers for subject classification. Training was carried out using sparse categorical cross-entropy and the Adam optimizer. The dataset was split 80/20 for training and testing, and early stopping was applied to prevent overfitting and improve generalization. The results show that CardioIdNet achieves excellent performance, with accuracy of 99%, precision, recall, and F1-score of 98.18%, an AUC of 99%, and a false negative rate of 1.85%. CardioIdNet suggests to be a promising solution for biometric authentication in healthcare real-time settings, offering a balance of simplicity, interpretability, and efficiency through image-based deep learning. Domenico Rossi, Alessia Auriemma Citarella, Fabiola De Marco, Genny Tortora, Huiru Zheng |
Multim. Tools Appl. | 5 |
| 2025 | Assessing a Smart-Insole-Based System in Stroke Gait Pattern RecognitionabstractStroke commonly leads to long-term gait impairments, underscoring the need for objective and continuous functional assessment during rehabilitation. This study employs machine learning methods to assess a smart insole-based system in stroke gait recognition. Data were collected from stroke survivors and healthy control participants during Walk and Timed-Up-and-Go tasks. After preprocessing, group differences were quantified using Hedges' g, and multiple machine learning models were applied to classify two participant groups. Support Vector Machine and KNN achieved the best performance, with accuracies of 0.88. The results demonstrate that sensor-based gait features can be used to distinguish stroke gait patterns from control gait patterns, highlighting the potential of this approach for future homebased monitoring and personalised rehabilitation. Yu-Huan Chien, Chun-Chung Chang, Jia-Yu Li, Kai-Li Fang, Wen-Yuan Lee, Mi-Hsuan Lin, Yi-Yin Lai, Luigi D'Arco, Alastair Martin, Katy Pedlow, Haying Wang, Huiru Zheng, Che-Lun Hung |
BIBM | 12 |
| 2025 | Review of Learning-Based Antibody Design: From Sequence to StructureabstractImproving antibodies' affinity and specificity has traditionally relied on iterative display selections or structure-based design, both costly and time-intensive. Recent advances in Deep Learning offer data-driven priors that effectively narrow the sequence space before expensive experiments. This paper provides an overview of the progress and challenges of learning -based antibody design. Adopting a pipeline-first perspective, this review organises current methods into three categories: (A) sequence-only protein language models (PLMs); (B) structure-aware strategies, including inverse folding and complex-aware optimisation; and (C) integrated AI-physics workflows. To avoid mixing endpoints, prospective wet-lab outcomes (e.g. hit rates, affinity gains) are reported separately from structure-linked sur-rogates (e.g. region recovery, refold root-mean-square deviation (RMSD), deep mutational scanning (DMS) correlation). Evidence indicates that sequence-only PLMs are effective for low-budget screening, inverse folding methods provide backbone-conditioned ranking and structure-preserving edits, and lightweight AI-physics overlays help prioritise manufacturable candidates. A concise method -selection guide is provided for different data availability scenarios. Jialin Lyu, Ciaran Doherty, Hugh Morgan, Haiying Wang 0001, Huiru Zheng |
BIBM | 5 |
| 2025 | Cross-Attentioned Dynamic Hierarchical Representation Learning Utilizing Mamba Fusion for Brain Network AnalysisabstractFunctional brain networks (FBNs) derived from resting-state functional Magnetic Resonance Imaging (rs-fMRI) have become pivotal tools for the auxiliary diagnosis of brain disorders, including Alzheimer's Disease (AD) and Major Depressive Disorder (MDD). However, most existing FBN analysis methods struggle to effectively capture both the temporal dynamics of rs-fMRI data and the hierarchical topological structures within FBNs, due to the inherent spatiotemporal complexity of brain activities. To address these challenges, we propose a novel framework for brain network analysis, called Crossattentioned Dynamic Hierarchical representation learning with Mamba fusion (CDHM). CDHM includes four principal phases: (1) dynamic FBNs construction via overlapping sliding windows, (2) cross-attentioned spatial encoding to capture local-to-global spatial interactions, (3) hierarchical graph pooling to distill multiscale FBN organisation, and (4) Mamba-based fusion modelling long-range temporal dependency and adaptive information flow regulation. Our framework pioneers synergistic integration of spatial cross-attention mechanisms with Mamba-based temporal fusion, enabling joint learning of transient neural dynamics and multi-scale hierarchical brain representations. We validate our approach through comprehensive experiments on two publicly available datasets, demonstrating superior performance compared to existing methods. Junze Wang, Idongesit Ekerete, Huiru Zheng, Lishan Qiao |
BIBM | 5 |
| 2025 | Visualization of local wind field based forest-fire's forecast modeling for transportation planning
Lvqing Yang, Huiru Zheng |
Multim. Tools Appl. | 4 |
| 2024 | SCOP: A Sequence-Structure Contrast-Aware Framework for Protein Function PredictionabstractImproving the ability to predict protein function can potentially facilitate research in the fields of drug discovery and precision medicine. Technically, the properties of proteins are directly or indirectly reflected in their sequence and structure information, especially as the protein function is largely determined by its spatial properties. Existing approaches mostly focus on protein sequences or topological structures, while rarely exploiting the spatial properties and ignoring the relevance between sequence and structure information. Moreover, obtaining annotated data to improve protein function prediction is often time-consuming and costly. To this end, this work proposes a novel contrast-aware pre-training framework, called SCOP, for protein function prediction. We first design a simple yet effective encoder to integrate the protein topological and spatial features under the structure view. Then a convolutional neural network is utilized to learn the protein features under the sequence view. Finally, we pretrain SCOP by leveraging two types of auxiliary supervision to explore the relevance between these two views and thus extract informative representations to better predict protein function. Experimental results on four benchmark datasets and one self-built dataset demonstrate that SCOP provides more specific results, while using less pre-training data. Chengxin He, Huiru Zheng, Xinye Wang, Yidan Zhang 0001, Lei Duan |
BIBM | 3 |
| 2024 | Comparative Analysis of Machine Learning Approaches for Emotion Recognition Using EEG and ECG SignalsabstractEmotions significantly influence human behaviour and decision-making, particularly in a digital era dominated by human-computer interactions (HCIs). Emotion can be expressed in various forms, including facial expressions, textual descriptions, and physiological responses. The main objective of this study is to comparatively analyze the performance of various machine learning (ML) classifiers to accurately recognize human emotional states using electroencephalogram (EEG) and electrocardiogram (ECG) signals. This study uses the DREAMER dataset and classifies emotional state in four different ways according to valence, arousal, and dominance (VAD) values – binary emotions, positive-neutral-negative (PNN) emotions, two-dimensional valence-arousal emotional space, and three-dimensional VAD emotional space. An ML pipeline has been developed to detect human emotions with EEG and ECG signals. Without removing outliers and balancing the dataset, the classifier that achieved the best performance was the ensemble classifier (SVM + random forest). If emotion is defined as a binary state, our experimental results show that both the SVM and the ensemble classifiers strike a good performance with approximately 80% accuracy; however, they perform poorly with the non-binary emotional models. The multinomial logistic regression (MLR) classifier and the random forest (RF) classifier consistently achieve a good performance for both the binary and the non-binary emotion models with 80% - 90% accuracy, its accuracy is higher than the accuracy in the original DREAMER experimental results. Our study experimentally confirmed this obvious finding. Jing-Hua Ye, Ian Cleland, Huiru Zheng, Patrick McAllister |
BIBM | 3 |
| 2024 | DDTExplainer: Mining Drug-Disease Therapeutic Mechanisms based on GNN ExplainabilityabstractFor a clinical prescription, clarifying the molecular mechanisms of actions (MMOAs) of the drug-disease interaction is helpful to optimize treatment, suggest possible side effects, and realize individualized treatment. Considering the relations among multiple biomedical entities, such as drugs, diseases, targets (genes), and pathways, what paths can be extracted connecting these biomedical entities to resemble the real mechanisms of a specific drug to a particular disease? Answering this question is crucial for understanding the underlying molecular mechanisms behind complex drug actions and identifying key pathways that can facilitate effective therapeutic interventions. In this paper, we propose an approach DDTExplainer that constructs a path-based graph neural network (GNN) explainer to mine the drug-disease therapeutic mechanisms. Technically, DDTExplainer transforms the drug-disease therapeutic mechanisms mining task into a GNN-based link prediction model explanation task. Firstly, a GNN-based drug-disease therapeutic prediction model is trained and joint-optimized with a translation-based graph embedding model. Secondly, mask learning is utilized to find the most prediction-influential edges and generate the path-based explanations with the shortest path algorithm. Finally, we assess the efficacy of DDTExplainer on a ground-truth dataset that consists of labeled entries describing the drug-disease therapeutic mechanisms. These labels are derived from well-established drug-target interactions, disease-target interactions, as well as target-pathway relationships, which were verified by wet experiments. Yidan Zhang 0001, Lei Duan, Huiru Zheng, Haiying Wang 0001, Yongmei Lu |
BIBM | 3 |
| 2024 | Application of Smart Insoles for Recognition of Activities of Daily Living: A Systematic ReviewabstractRecent years have witnessed the increasing literature on using smart insoles in health and well-being, and yet, their capability of daily living activity recognition has not been reviewed. This paper addressed this need and provided a systematic review of smart insole-based systems in the recognition of Activities of Daily Living (ADLs). The review followed the PRISMA guidelines, assessing the sensing elements used, the participants involved, the activities recognised, and the algorithms employed. The findings demonstrate the feasibility of using smart insoles for recognising ADLs, showing their high performance in recognising ambulation and physical activities involving the lower body, ranging from 70% to 99.8% of Accuracy, with 13 studies over 95%. The preferred solutions have been those including machine learning. A lack of existing publicly available datasets has been identified, and the majority of the studies were conducted in controlled environments. Furthermore, no studies assessed the impact of different sampling frequencies during data collection, and a trade-off between comfort and performance has been identified between the solutions. In conclusion, real-life applications were investigated showing the benefits of smart insoles over other solutions and placing more emphasis on the capabilities of smart insoles. Luigi D'Arco, Graham McCalmont, Haiying Wang 0001, Huiru Zheng |
ACM Trans. Comput. Heal. | 4 |
| 2023 | SIDE: Sequence-Interaction-Aware Dual Encoder for Predicting circRNA Back-Splicing EventsabstractCircular RNAs (circRNAs) play a critical role in gene regulation and association with diseases due to their specialized structure, which is formed as a closed loop structure during a non-canonical splicing process where the donor site back-spliced to an upstream acceptor site. As fundamental work to clarify their functions and mechanisms, a large number of computational methods for predicting circRNA formation have been proposed, among which, in particular, deep learning is utilized to capture relevant patterns from raw RNA sequences and model their interactions to facilitate prediction. However, these methods fail to fully utilize the important characteristics of back-splicing events, i.e., the positional information of the splice sites and the interaction features of its flanking sequences, for prediction. To this end, we hereby propose a novel approach called SIDE for predicting circRNA back-splicing events using only nucleotide sequences. Our model employs a dual encoder to capture global and interactive features of the sequence, and then a decoder designed by the contrastive learning to fuse out discriminative features improving the prediction of circRNAs formation. Empirical results on three real-world datasets have shown the effectiveness of SIDE. Our code is publicly available at https://github.com/scu-kdde/Bioinfo-SIDE-2023. Chengxin He, Lei Duan, Huiru Zheng, Yuening Qu, Zhenyang Yu |
BIBM | 3 |
| 2023 | A Novel Mixed Effects Random Forest Approach for Predicting Dairy Cattle Methane EmissionsabstractMethane (CH4) emissions produced by dairy cattle (DC) are a key contributor to global warming. To assess the effectiveness of strategies designed to mitigate CH4emissions, complex and expensive recording equipment is required. Therefore, the use of predictive models based on animal information provides a more accessible alternative. Traditionally, Statistical (SA) methods have been employed in the prediction of DC CH4emissions. However due to the smart farming revolution, the scale and variety of complex animal information now available for the prediction of DC CH4emissions has grown exponentially, and within them are likely to exist non-linear relationships which these traditional SA models may struggle to capture. Therefore, this research aims to explore if Machine Learning (ML) models are a viable alternative for the prediction of DC CH4emissions, as they can handle and extract these inevitable non-linear relationships present within today's large, heterogeneous datasets. In this research, we compared a traditional SA method, a Linear Mixed Effects (ME) model, with an original ML method, a Random Forest (RF) model, as well as a novel SA/ML hybrid method, a Mixed Effects Random Forest (MERF) model, in the prediction of CH4emissions (CH4g/d) produced by DC across 32 experiments. The ML RF model was able to challenge the traditional SA ME model in the prediction of DC CH4emissions, achieving a Root Mean Square Prediction Error (RMSPE) and Concordance Correlation Coefficient (CCC) of 52.73 CH4g/d and 0.70 respectively, compared to a ME model’s 53.90 CH4g/d and 0.71. When both the ME and RF models were combined within the novel SA/ML hybrid MERF model, a lower RMSPE and higher CCC were achieved than by each of its composite parts in isolation, 51.87 CH4g/d and 0.73 respectively. These results demonstrate the potential of ML in the prediction of DC CH4emissions, particularly when hybridised alongside traditional SA methods. Stephen Ross, Tianhai Yan, Haiying Wang 0001, Masoud Shirali, Huiru Zheng |
BIBM | 5 |
| 2023 | A smart chicken farming platform for chicken behavior identification and feed residual estimationabstractIt is very potential to develop digital villages for promoting smart agriculture. As one of the important research fields of smart agriculture, smart chicken farms encounter management problems such as difficulties in quickly and accurately warning of sick and dead chickens and estimating feed residuals. Therefore, this study not only respectively proposed CKTrack and FRCM to detect sick and dead chickens and estimate feed residuals, but also developed a smart chicken farming platform for automagical management. Our main results include (1) the proposed CKTrack method can effectively identify sick and dead chickens under the condition of limited data volume and computing capacity; (2) the proposed FRCM method can accurately estimate the feed residuals; and (3) the smart chicken farming platform developed can provide farmers with functions such as early warning of sick and dead chickens, visualization of the chicken quantity inventory, and feed residual estimation. Jiezhi Yang, Antong Zhou, Chaochao Qu, Kuo Zhao, Linjing Wei, Le Zhang 0004, Zirong Liu, Wenjing Tao, Kangzhe Ma, Huiru Zheng |
BIBM | 14 |
| 2023 | MGDTI: Graph Transformer with Meta-Learning for Drug-Target Interaction PredictionabstractDrug-target interaction (DTI) prediction is of great importance for drug discovery and development. With the rapid development of biological and chemical technologies, computational methods for DTI prediction are becoming a promising strategy. However, there are few methods which explore solving the cold-start problem in DTI prediction scenarios due to most of existing methods require modeling under the existing interaction that can’t effectively capture information from new drugs and new targets which have few interactions in existing literature. In this paper, we propose a graph transformer method based on meta-learning named MGDTI to fill the gap. In particular, we employ drug-drug similarity and target-target similarity as additional information for network to mitigate the scarcity of interactions. Besides, we trained our model via meta-learning to be adaptive to cold-start tasks. Moreover, we introduced graph transformer to prevent over-smoothing by capturing long-range dependencies. Comparison results on the benchmark dataset demonstrate that our proposed MGDTI is effective in the DTI prediction. Chengxin He, Yuening Qu, Huiru Zheng, Lei Duan, Jie Zuo |
BIBM | 4 |
| 2023 | A Factor Graph Based Indoor Localization Approach for HealthcareabstractIn healthcare facilities, indoor localization technology has a broad range of applications. Traditional Pedestrian Dead Reckoning (PDR) and WiFi fingerprint-based methods each have their limitations. To address these challenges, this study introduces a multi-source fusion indoor localization system that uses a Factor Graph to integrate inertial positioning algorithms with WiFi fingerprint-based localization. The system processes accelerometer and gyroscope data using a data-driven PDR algorithm. For WiFi localization, considering that the extensive data collection required is a significant barrier to the deployment of WiFi-based localization methods, the proposed approach applies Gaussian process regression techniques to limited WiFi fingerprint data, significantly reducing initial deployment costs and enhancing accuracy. Finally, the entire system employs a Factor Graph for the integration of the data-driven PDR and WiFi fingerprint localization results. Experimental results show that, compared to using only inertial or WiFi data for localization, this method significantly improves localization accuracy. The findings suggest that this approach could prompt the utilization of indoor localization technology in healthcare facilities. Shiyu Zheng, Ao Peng, Lingxiang Zheng, Huiru Zheng, Haiying Wang 0001 |
BIBM | 7 |
| 2023 | DeepHAR: a deep feed-forward neural network algorithm for smart insole-based human activity recognitionabstractAbstract Health monitoring, rehabilitation, and fitness are just a few domains where human activity recognition can be applied. In this study, a deep learning approach has been proposed to recognise ambulation and fitness activities from data collected by five participants using smart insoles. Smart insoles, consisting of pressure and inertial sensors, allowed for seamless data collection while minimising user discomfort, laying the baseline for the development of a monitoring and/or rehabilitation system for everyday life. The key objective has been to enhance the deep learning model performance through several techniques, including data segmentation with overlapping technique (2 s with 50% overlap), signal down-sampling by averaging contiguous samples, and a cost-sensitive re-weighting strategy for the loss function for handling the imbalanced dataset. The proposed solution achieved an Accuracy and F1-Score of 98.56% and 98.57%, respectively. The Sitting activities obtained the highest degree of recognition, closely followed by the Spinning Bike class, but fitness activities were recognised at a higher rate than ambulation activities. A comparative analysis was carried out both to determine the impact that pre-processing had on the proposed core architecture and to compare the proposed solution with existing state-of-the-art solutions. The results, in addition to demonstrating how deep learning solutions outperformed those of shallow machine learning, showed that in our solution the use of data pre-processing increased performance by about 2%, optimising the handling of the imbalanced dataset and allowing a relatively simple network to outperform more complex networks, reducing the computational impact required for such applications. Luigi D'Arco, Haiying Wang 0001, Huiru Zheng |
Neural Comput. Appl. | 3 |
| 2023 | Window-Adjusted Common Spatial Pattern for Detecting Error-Related Potentials in P300 BCIabstractAbstract Under certain task conditions, error-related potential (ErrP) will be elicited, meaning that the subject is perceiving an error, responding to an external error, or engaging in a cognitive process of reinforcement learning. The detection of ErrP on a single trial basis has been studied and applied to improve all kinds of brain–computer interfaces (BCIs). However, the performance of this kind of detection is not currently good enough. In the paper, we proposed a novel method, called window-adjusted common spatial pattern (WACSP), for detecting ErrP in P300 BCI. In this method, the coefficient of determination was introduced to measure the difference of Electroencephalogram (EEG) signals on a channel at a moment and to guide the search of time windows in which EEG differences are significant, and common spatial pattern (CSP) was further used to capture the stable spatial patterns of EEG differences between correct and incorrect responses in each time window. WACSP and the commonly used methods were tested on the data sets that were built using the EEG signals acquired during the P300 BCI experiments with different feedback. The comparisons of accuracy, area under receiver operating characteristics curve (AUC) and F-measure show that WACSP significantly outperforms the commonly used methods. The proposed method can improve ErrP detection based on a single trial. Minghong Li, Wenming Zheng, Huiru Zheng |
Neural Process. Lett. | 6 |
| 2023 | An Integrative Disease Information Network Approach to Similar Disease DetectionabstractDisease similarity analysis impacts significantly in pathogenesis revealing, treatment recommending, and disease-causing genes predicting. Previous works study the disease similarity based on the semantics obtaining from biomedical ontologies (e.g., disease ontology) or the function of disease-causing molecules. However, such methods almost focus on a single perspective for obtaining disease features, which may lead to biased results for similar disease detection. To address this issue, we propose a disease information network-based integrative approach named MISSION for detecting similar diseases. By leveraging the associations between diseases and other biomedical entities, the disease information network is established first. Then, the disease similarity features extracted from the aspects of disease taxonomy, attributes, literature, and annotations are integrated into the disease information network. Finally, the top-k similar disease query is performed based on the integrative disease information. The experiments conducted on real-world datasets demonstrate that MISSION is effective and useful in similar disease detection. Wuli Xu, Lei Duan, Huiru Zheng, Jesse Li-Ling, Yidan Zhang 0001, Tingting Wang 0009, Ruiqi Qin 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | A Rapid Detection of Parkinson's Disease using Smart Insoles: A Statistical and Machine Learning ApproachabstractDetermining whether a subject has a gait impairment due to a disease or to the loss of muscularity due to advancing age is fundamental for an early diagnosis of musculoskeletal diseases. Parkinson’s is the second most common neurodegenerative disease. The disease’s most prevalent symptom is slow movement or sluggish gait, which can adversely impact the individual’s quality of life. Generally, the gait analysis is carried out on long test sessions, which include for example long periods of walking, that cause inconvenience when the subjects under test have marked gait impairments. To help the diagnosis of Parkinson’s disease, in this study we investigated the classification of Parkinson’s disease by analysing only a few seconds of walking data using smart insoles, statistical analysis and machine learning techniques. The data from the smart insoles was assessed using correlation analysis. By creating pressure groups and analysing their values, it was found that the number of sensors could be reduced from 16 to 7. Furthermore, a feature vector representing the subject’s gait was created by applying on the data a time windowing segmentation of 5 seconds and extracting six statistical features (mean, variance, skewness, kurtosis, energy and entropy). Four different models have been compared in terms of classification performance, reaching an F1-Score in the classification of patients with Parkinson’s against healthy subjects, considering adult and elderly subjects as two separate classes, of 97.04% using the Random Forest. Such metric increased to 98.89%, using the K-Nearest Neighbours when healthy subjects were considered as a single class. The models’ performance for each experiment was determined to be statistically equivalent, demonstrating the potential of this approach to provide the groundwork for the rapid detection of Parkinson’s disease. Although the performance obtained is promising the number of subjects included in the study was fairly low, with a high bias towards the number of healthy subjects. Hence, in future work, the proposed solution will be tested on a larger cohort to ascertain its robustness. Luigi D'Arco, Haiying Wang 0001, Huiru Zheng |
BIBM | 3 |
| 2022 | A New Phylogeny-Driven Random Forest-Based Classification Approach for Functional MetagenomicsabstractClassifying microbial genes into their functional repertoire is an important task for metagenomic studies, where the research community is trying to develop Machine Learning (ML) based methods to achieve good classification performance. Random Forest (RF) has been proposed as one of the most favorable methods for such supervised analysis when applied over the abundance profiles of microbial genes mapping them to functional phenotypes. To further explore and make optimization in the existing RF model (based on the biological relationships between microbial features), a new classification method based on RF as guided by the evolutionary ancestry of microbial phylogeny, i.e. Phylogeny-RF, has been developed in this paper. This method facilitates to capture the effects of phylogenetic relatedness in a ML classifier itself. Closely related microbes by phylogeny are highly correlated and tend to have similar genetic and phenotypic traits. Such microbes behave similarly; and hence tend to be selected together or one of these could be dropped from the analysis, to make the ML process better. The proposed Phylogeny-RF algorithm has been compared with state-of-the-art classification methods including RF and the phylogeny-aware method of MetaPhyl, using 2 real-world 16S rRNA metagenomic data sets. It is observed that the proposed method performed better than the other phylogeny-driven benchmarks. For example, Phylogeny-RF attained a high AUC of 0.949 over soil microbiomes in comparison to other benchmarks. Jyotsna Talreja Wassan, Haiying Wang 0001, Huiru Zheng |
BIBM | 3 |
| 2022 | An Enhanced Visual SLAM Supported by the Integration of Plane Features for the Indoor EnvironmentabstractThis paper presents an enhanced indoor RGB-D simultaneously localisation and mapping (SLAM) system based on the integration of plane and point features. A new method was proposed to register each point feature to a corresponding plane feature and then modify its position accordingly. The plane features are parallelly extracted from depth data sources and used jointly to solve the camera pose with point features. Both point and plane features are stored on the map and used for backend optimisation, where the weights associated with features can be dynamically updated. At the same time, the on-plane feature points are fixed during the optimisation. The proposed method has been tested with open-source benchmarks, including the scenarios with or without a structured environment. Experiment results demonstrated that the proposed algorithm performs better than other widely cited visual SLAM systems in some structured environments, in which the point features form plane features without introducing excessive errors. Bingxin Zi, Haiying Wang 0001, Huiru Zheng |
IPIN | 4 |
| 2022 | An explainable framework for drug repositioning from disease information networkabstractExploring efficient and high-accuracy computational drug repositioning methods has become a popular and attractive topic in drug development. This technology can systematically identify potential drug-disease interactions, which could greatly alleviate the pressures from the high cost and long period taken by traditional drug research and discovery. However, plenty of current computational drug repositioning approaches lack interpretability in predicting drug-disease associations, which will not be friendly to their subsequent in-depth research. To this end, we hereby propose a novel computational framework, called EDEN, for exploring explainable drug repositioning from the disease information network (DIN). EDEN is a graph neural network framework that learns the local semantics and global structure of the DIN, and models the drug-disease associations into the DIN by maximizing the mutual information of both and an end-to-end manner. In this way, the learned biomedical entity and link embeddings are enabled to retain the ability to drug repositioning with the semantical structure of external knowledge, thereby making interpretation possible. Meanwhile, we also propose a matching score based on the final embeddings to generate the predictive drug repositioning explanation. Empirical results on the real-world dataset show that EDEN outperforms other state-of-the-art baselines on most of the metrics. Further studies reveal the effectiveness of the explainability of our approach. Chengxin He, Lei Duan, Huiru Zheng, Linlin Song, Menglin Huang |
Neurocomputing | 3 |
| 2022 | Ergonomic Assessment Method of Risk Factors for Musculoskeletal Disorders Associated with Sitting PosturesabstractMusculoskeletal disorders (MSDs) are associated with sitting postures. The assessment and prevention of risk factors for workplace exposure are indispensable aspects of reducing the occurrence of MSDs. This paper proposes an ergonomic assessment method of risk factors for MSDs associated with sitting postures in the actual working conditions. A Kinect sensor with the RULA method was primarily used to collect the data and evaluate the relevant postures. The results obtained were compared with the evaluation results by a human expert. Additionally, we verified the capability and effectiveness of this method. A program system for human joint recognition and acquisition was implemented. The results indicated that the Kinect joint data is generally accurate and can adequately complete the RULA evaluation table. The results from the front and right-hand side obtained by the Kinect were consistent with the results of the expert evaluation, and no significant difference was observed between them ([Formula: see text]). However, when the participants faced the Kinect, the sensor performed better, and the evaluation result was more accurate. A high consistency was observed between the evaluation results obtained from the front and the expert (proportion agreement [Formula: see text], Cohen’s [Formula: see text]). Only a slight consistency was observed between the evaluation results obtained from the right-hand side and the expert (proportion agreement [Formula: see text], Cohen’s [Formula: see text]). This research created a new ergonomic method for the risk assessment of MSDs associated with sitting postures. The combination of theory and practice is crucial in the risk assessment of sitting postures in workplaces. Sihan Huang, Faming Wang, Sixi Chen, Huiru Zheng |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2022 | Mining Similar Aspects for Gene Similarity Explanation Based on Gene Information NetworkabstractAnalysis of gene similarity not only can provide information on the understanding of the biological roles and functions of a gene, but may also reveal the relationships among various genes. In this paper, we introduce a novel idea of mining similar aspects from a gene information network, i.e., for a given gene pair, we want to know in which aspects (meta paths) they are most similar from the perspective of the gene information network. We defined a similarity metric based on the set of meta paths connecting the query genes in the gene information network and used the rank of similarity of a gene pair in a meta path set to measure the similarity significance in that aspect. A minimal set of gene meta paths where the query gene pair ranks the highest is a similar aspect, and the similar aspect of a query gene pair is far from trivial. We proposed a novel method, SCENARIO, to investigate minimal similar aspects. Our empirical study on the gene information network, constructed from six public gene-related databases, verified that our proposed method is effective, efficient, and useful. Yidan Zhang 0001, Lei Duan, Huiru Zheng, Jesse Li-Ling, Ruiqi Qin 0001, Chengxin He, Tingting Wang 0009 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | A PDRVIO Loosely coupled Indoor Positioning System via Robust Particle FilterabstractIn recent years, the Visual Inertial Odometry (VIO) technology has attracted attention as a support technology that improves medical experience and management efficiency. However, the performance of most VIO systems will drop drastically when the light intensity changes significantly or there are few texture features from the images. This paper designs a visual-inertial fusion-based navigation Indoor Positioning system to deal with the challenging scenario. It loosely coupled an inertial sensor-based pedestrian dead reckoning (PDR) model with the VIO model via a robust particle filter. The state estimation of the particle filter is based on the PDR model. The VIO model is used for the measurements of the particle filter. It compensates the gross errors of the VIO with a visual error propagation model which is established according to the posterior observation residuals of visual feature points. It is verified through experiments that the PDR/VIO fusion indoor positioning system based on the robust particle filter implemented in this paper has improved positioning accuracy and strong ability to deal with complex scenes. Xinwei Hu, Weilong Huang, Lingxiang Zheng, Ao Peng, Huiru Zheng, Haiying Wang 0001 |
BIBM | 7 |
| 2021 | Deep learning in systems medicineabstractSystems medicine (SM) has emerged as a powerful tool for studying the human body at the systems level with the aim of improving our understanding, prevention and treatment of complex diseases. Being able to automatically extract relevant features needed for a given task from high-dimensional, heterogeneous data, deep learning (DL) holds great promise in this endeavour. This review paper addresses the main developments of DL algorithms and a set of general topics where DL is decisive, namely, within the SM landscape. It discusses how DL can be applied to SM with an emphasis on the applications to predictive, preventive and precision medicine. Several key challenges have been highlighted including delivering clinical impact and improving interpretability. We used some prototypical examples to highlight the relevance and significance of the adoption of DL in SM, one of them is involving the creation of a model for personalized Parkinson's disease. The review offers valuable insights and informs the research in DL and SM. Haiying Wang 0001, Estelle Pujos-Guillot, Blandine Comte, João Luís de Miranda, Vojtech Spiwok, Ivan Chorbev, Filippo Castiglione, Paolo Tieri, Steven Watterson, Roisin McAllister, Tiago De Melo Malaquias, Massimiliano Zanin, Taranjit Singh Rai, Huiru Zheng |
Briefings Bioinform. | 14 |
| 2021 | An Ontology-Independent Representation Learning for Similar Disease Detection Based on Multi-Layer Similarity NetworkabstractTo identify similar diseases has significant implications for revealing the etiology and pathogenesis of diseases and further research in the domain of biomedicine. Currently, most methods for the measurement of disease similarity utilize either associations of ontological disease concepts or functional interactions between disease-related genes. These methods are heavily dependent on the ontology, which are not always available, and the selection of datasets. Moreover, many methods suffer from a drawback that they only use a single metric to evaluate disease similarity from an individual data source, which may result in biased conclusions without consideration of other aspects. In this study, we proposed a novel ontology-independent framework, namely RADAR, for learning representations for diseases to deduce their similarities from an integrative perspective. By leveraging the associations between diseases and disease-related biomedical entities, a disease similarity network was built under various metrics. Then, a multi-layer disease similarity network was constructed by integrating multiple disease similarity networks derived from multiple data sources, where the representation learning was derived to provide a comprehensive evaluation of disease similarities. The performance of RADAR was assessed by a benchmark disease set and 100 random disease sets. Experimental results demonstrated that RADAR can detect similar diseases effectively. Ruiqi Qin 0001, Lei Duan, Huiru Zheng, Jesse Li-Ling, Kaiwen Song, Yidan Zhang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Bayesian Network Approach to Modelling Nitrogen Utilization Efficiency of Dairy CowsabstractLosses of nitrogen (N) from dairy cattle farming system cause environmental pollution and impact human health. A number of statistical models for predicting manure N excretion from lactating dairy cows have been developed based on regression analysis. In this study, we proposed a Bayesian network-based approach to modelling relationships among factors influencing manure N excretion of lactating dairy cows using a dataset collated from total diet digestibility studies undertaken at Agri-Food and Biosciences Institute in Northern Ireland. The preliminary results indicate that Bayesian network model can be used to capture relationships among factors that influence N utilization efficiency and can be used to establish causal influence among predictors. These may provide an effective tool for optimizing the management of feed N resource for current dairy production and developing strategies to reduce N excretion in dairy production systems. Xianjiang Chen, Huiru Zheng, Haiying Wang 0001, Tianhai Yan |
BIBM | 2 |
| 2020 | DRAMA: Discovering Disease-related circRNA-miRNA-mRNA Axes from Disease-RNA Information NetworkabstractNon-coding RNAs are gaining prominence in biology and medicine, as they play major roles in cellular homeostasis and disease. A large number of computational methods have been recently developed for the prediction of the relationship between ncRNAs and diseases, which can alleviate the time-consuming and labor-intensive exploration among biological experiments. However, such methods have mainly focused on the association between the disease and certain types of ncRNAs such as miRNA or circRNA, thereby ignoring the impact of the interactions among ncRNAs on the diseases. We hereby propose a novel approach called DRAMA for discovering disease-related circRNA-miRNA-mRNA axes from the disease-RNA information network we constructed. Our method, using graph convolutional network, learns the characteristic representation of each biological entity by propagating and aggregating local neighbor information based on the global structure of the network. And then we design a favorable measurement to infer disease-related circRNA-miRNA-mRNA axes based on the learned embeddings. To evaluate the effectiveness of DRAMA, we conduct experiments on real-world datasets. Further analysis reveals that DRAMA outperforms other state-of-the-art baselines on most of the metrics. Chengxin He, Lei Duan, Huiru Zheng, Jesse Li-Ling, Longhai Li |
BIBM | 3 |
| 2020 | A multilayer co-occurrence network reveals the systemic difference of diet-based rumen microbiome associated with methane yield phenotypeabstractReducing rumen methane emissions based on the basal diet is one primary approach for livestock production. However, microbiological mechanisms of rumen methane emissions under different diets are very complex and lack comprehensive understanding. This research developed an innovative multilayer network framework to investigate relationships of rumen microbes, transformations of metabolites and functions of rumen microbial genes. The multilayer network model explored the diets-based variation in rumen microbial system differences. The topological structure of the multilayer network reflects volatile fatty acid (VFA) fermentation characteristics between diets. Propionate is the highest degree node of the concentrate-diet (CONC) multilayer network while acetate has the highest degree in the forage-diet (FOR) multilayer network. The microbial communities under FOR diet are more diverse and have more complicated interactions. Methanobrevibacter copresented with Methanosphaera only in the CONC network. There are 33 microbial functional genes in the FOR network that are enriched in the methane synthesize functional module, including the CO2 reduction and the enzymatic reaction of acetyl phosphate to synthesize acetyl-CoA. The microbial genes of the CONC network are found to be associated with the serine biosynthetic and the biochemical process of acetate convert to acetyl-CoA and acetyl phosphate. This study verified the variations of diet-based methane phenotypic are results of systemic changing in the metabolism of the entire rumen microbiome. Haiying Wang 0001, Huiru Zheng, Richard J. Dewhurst, Rainer Roehe |
BIBM | 3 |
| 2020 | MISSION: Multimodal-Information-Aided Similar Disease Detection Based on Disease Information NetworkabstractTo detect similar diseases is meaningful for revealing pathogenesis, and predicting therapeutic drugs. Previous methods measure disease similarity almost according to the semantic on biomedical ontology or the function of disease-causing molecules. However, such methods mostly describe diseases from single information, which may lead to a biased description of the relationships among diseases. In this paper, we propose a novel approach, called MISSION, for measuring the disease similarity based on multimodal-information. MISSION enhances similar disease detection based on disease information network from three aspects, including disease ontology, attribute, and literature, therefore providing a comprehensive evaluation for disease similarity. Through experiments on real-world datasets, we demonstrate that MISSION is effective, efficient, and potentially useful. Further analysis shows that MISSION has the ability to detect similar diseases with varying degrees of rich information. Wuli Xu, Lei Duan, Huiru Zheng, Jesse Li-Ling, Menglin Huang, Yidan Zhang 0001 |
BIBM | 3 |
| 2020 | Developing the novel bioinformatics algorithms to systematically investigate the connections among survival time, key genes and proteins for Glioblastoma multiformeabstractBACKGROUND: Glioblastoma multiforme (GBM) is one of the most common malignant brain tumors and its average survival time is less than 1 year after diagnosis. RESULTS: Firstly, this study aims to develop the novel survival analysis algorithms to explore the key genes and proteins related to GBM. Then, we explore the significant correlation between AEBP1 upregulation and increased EGFR expression in primary glioma, and employ a glioma cell line LN229 to identify relevant proteins and molecular pathways through protein network analysis. Finally, we identify that AEBP1 exerts its tumor-promoting effects by mainly activating mTOR pathway in Glioma. CONCLUSIONS: We summarize the whole process of the experiment and discuss how to expand our experiment in the future. Yujie You, Xufang Ru, Wanjing Lei, Ming Xiao 0002, Huiru Zheng, Yujie Chen 0010, Le Zhang 0004 |
BMC Bioinform. | 6 |
| 2020 | Improving the Inference of Co-Occurrence Networks in the Bovine Rumen MicrobiomeabstractThe importance of the composition and signature of rumen microbial communities has gained increasing attention. One of the key techniques was to infer co-abundance networks through correlation analysis based on relative abundances. While substantial insights and progress have been made, it has been found that due to the compositional nature of data, correlation analysis derived from relative abundance could produce misleading results and spurious associations. In this study, we proposed the use of a framework including a compendium of two correlation measures and three dissimilarity metrics in an attempt to mitigate the compositional effect in the inference of significant associations in the bovine rumen microbiome. We tested the framework on rumen microbiome data including both 16S rRNA and KEGG genes associated with methane production in cattle. Based on the identification of significant positive and negative associations supported by multiple metrics, two co-occurrence networks, e.g., co-presence and mutual-exclusion networks, were constructed. Significant modules associated with methane emissions were identified. In comparison to previous studies, our analysis demonstrates that deriving microbial associations based on the correlations between relative abundances may not only lead to missing information but also produce spurious associations. To bridge together different co-presence and mutual-exclusion relations, a multiplex network model has been proposed for integrative analysis of co-occurrence networks which has great potential to support the prediction of animal phytotypes and to provide additional insights into biological mechanisms of the microbiome associated with the traits. Huiru Zheng, Haiying Wang 0001, Richard J. Dewhurst, Rainer Roehe |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | WeightMentor, bespoke chatbot for weight loss maintenance: Needs assessment & DevelopmentabstractObesity and overweight are significant health risks. Effective communication can improve weight loss maintenance. There are technologies to support communication in weight management, such as text messaging that has been shown to benefit self-reporting and motivation but has its limitations. Having a conversation is an emotional experience, and chatbots are conversation-driven intelligent systems. Chatbots have been reported to increase compliance with health interventions. The aim of this study was to identify the needs of individuals who are maintaining weight loss in order to design and develop WeightMentor, a weight loss maintenance chatbot. This was a needs assessment and technology (chatbot) development study. The needs assessment identified the needs adults aged 18+ who were maintaining weight loss using semi-structure interviews. These needs were used to inform the design and development of a chatbot. Data were analyzed using thematic analysis. Findings identified five key themes: (1) Weight loss maintenance is challenging; (2) Social contact is beneficial but may also reinforce unhealthy habits; (3) Apps should be convenient and support progress tracking; (4) Personal messages should be specific and relevant; (5) Chatbots have potential for weight loss maintenance. Chatbot, WeightMentor, was designed and developed based on findings from the needs assessment and a review of the most popular nutrition apps. WeightMentor operates within Facebook messenger, which is currently one of the most popular social media platform. WeightMentor's purpose is to positively influence the user's emotions while maintaining weight loss. Its functions include self-reporting of diet and physical activity, viewing previous self-reporting trends, and motivational messaging. Chatbots such as WeightMentor has the potential to aid weight loss maintenance. Further research is required to test the usability and effectiveness of this WeightMentor. Samuel Holmes, Anne Moorhead, Raymond R. Bond, Huiru Zheng, Vivien Coates, Michael F. McTear |
BIBM | 4 |
| 2019 | Automated Object Tracking for Animal Behaviour StudiesabstractIt can be an arduous and time consuming task to manually curate video data for tracking the movement of an object. This is particularly challenging when tracking animals whose behaviour is sporadic and unpredictable, but it can provide pertinent information when monitoring health, conditions and treatments. Here, we evaluate the use of machine learning retroactively applied to automatically track specific steers in hours of video with minimal manual interaction. This is approached using the Faster R-CNN Object Detection algorithm with VGG-16 acting as a feature extractor. Performance on a number of video segments is presented and discussed, and the issues encountered are outlined. This highlights a number of guidelines that should be taken under consideration when generating video data to improve object detection performance, and helps define the applicability of the approach to pre-existing data. Timmy Manning, Miguel Somarriba, Rainer Roehe, Simon Turner, Haiying Wang 0001, Huiru Zheng, Brian Kelly, Jennifer Lynch, Paul Walsh |
BIBM | 6 |
| 2019 | Is there an Optimal Technology to Provide Personal Supportive Feedback in Prevention of Obesity?abstractObesity is a global challenge that affects health and wellbeing worldwide. In this position paper, we review the digital technology used in prevention of obesity and present the proposed STOP project that integrates state-of-the-art wearable technology, chatbot, gamification data fusion, and machine learning with the aim to provide personalised supportive feedback for preventing obesity and maintaining healthy weight. Implication of sensitive data with General Data Protection Regulation (GDPR) is discussed. We conclude that machine learning plays an important role in data fusion, analytics, and providing optimal messaging tailored design to support healthy weight. Simone Sandri, Matthias L. Hemmje, Huiru Zheng, Felix Engel 0002, Anne Moorhead, Haiying Wang 0001, Raymond R. Bond, Michael F. McTear, Andrea Molinari, Paolo Bouquet |
BIBM | 3 |
| 2019 | A knowledge driven mutual information-based analytical framework for the identification of rumen metabolitesabstractMetabolites are the final product of biochemical reactions in the rumen micro-ecological system and very sensitive to changes of microbial genes. However, limited by the spectra library and the computational techniques of structure identification, the identification of metabolites from non-targeted metabolomics is time-consuming and inefficient. The absence of specific information about metabolites makes the biological interpretation of the quantitative analysis of metabolomics meaningless. Based on the nonlinear association between microbial genes and metabolites, combined with knowledge of metabolic pathways from the KEGG database, this study developed a knowledge driven mutual information-based analytical framework for identifying metabolites associated with integrals derived from NMR analysis results. In this study, one known metabolite and three sets of integrals with unknow metabolites were identified within the novel framework. The results showed that this mutual information-based framework could very efficiently target metabolites that may correspond to integrals from NMR spectra. Huiru Zheng, Haiying Wang 0001, Richard J. Dewhurst, Rainer Roehe |
BIBM | 2 |
| 2019 | A Phylogeny-aware Feature Ranking for Classification of Cattle Rumen MicrobiomeabstractMetagenomics is proliferating for studying environmental microbial communities and their role in animal functions. This paper aims to study the role of functions of microbial communities present in cattle (Bos taurus) and their relation to dietary supplement usage. The functional study was conducted as part of the EU H2020 MetaPlat project11MetaPlat, http://www.metaplat.eu. In this research, we proposed a novel phylogeny-driven approach to classify 16S rRNA samples from cattle rumen microbiome and relate them to the functional phenotype of diet (referred to as functional analysis). Phylogeny covers biological relationships from different taxonomical levels combined with their respective evolutionary measures. We performed this analysis by proposing a novel method based on phylogeny-adjusted distance-based indices. These indices are used in ranking microbial feature space derived from the topology of the phylogenetic tree. The integrative approach incorporating phylogeny into feature engineering as part of machine learning (ML) modeling, achieved high predictive performance with Accuracy of 0.962 and Kappa of 0.950 for classifying cattle microbiome into the phenotype of a diet supplemented with oil, nitrate, combined (with oil and nitrate) and controls. Jyotsna Talreja Wassan, Huiru Zheng, Haiying Wang 0001, Fiona Browne, Paul Walsh, Timmy Manning, Richard J. Dewhurst, Rainer Roehe |
BIBM | 2 |
| 2019 | SCENARIO: Discovery of Similar Aspects for Gene Similarity Explanation from Gene Information NetworkabstractGene similarity analysis not only provides information on understanding the biological roles and functions of a gene, but also reveals the relationships among different genes. In this paper, we identify the novel idea of mining similar aspects from gene information network, i.e., given a pair of genes, we want to know, in which aspects (meta paths) the two genes are mostly similar from the perspective of gene information network? We define a similarity metric based on the set of meta paths connecting the query genes in the gene information network, and use the rank of the similarity of a gene pair in a meta path set to measure the similarity significance in the aspect. A minimal set of meta paths where the query gene pair is ranked the best is a similar aspect. Computing the similar aspects of a query gene pair is far from trivial. In this paper, we propose a novel heuristic based-mining method, SCENARIO, to investigate minimal similar aspects. Our empirical study on the gene information network, constructed from seven public gene-related databases, verified that our proposed method is effective, efficient, and useful. Yidan Zhang 0001, Lei Duan, Huiru Zheng, Jesse Li-Ling, Ruiqi Qin 0001, Chengxin He |
BIBM | 3 |
| 2019 | Review of applications of high-throughput sequencing in personalized medicine: barriers and facilitators of future progress in research and clinical applicationabstractThere has been an exponential growth in the performance and output of sequencing technologies (omics data) with full genome sequencing now producing gigabases of reads on a daily basis. These data may hold the promise of personalized medicine, leading to routinely available sequencing tests that can guide patient treatment decisions. In the era of high-throughput sequencing (HTS), computational considerations, data governance and clinical translation are the greatest rate-limiting steps. To ensure that the analysis, management and interpretation of such extensive omics data is exploited to its full potential, key factors, including sample sourcing, technology selection and computational expertise and resources, need to be considered, leading to an integrated set of high-performance tools and systems. This article provides an up-to-date overview of the evolution of HTS and the accompanying tools, infrastructure and data management approaches that are emerging in this space, which, if used within in a multidisciplinary context, may ultimately facilitate the development of personalized medicine. Gaye Lightbody, Valeriia Haberland, Fiona Browne, Laura Taggart, Huiru Zheng, Eileen Parkes, Jaine K. Blayney |
Briefings Bioinform. | 5 |
| 2019 | Community effort endorsing multiscale modelling, multiscale data science and multiscale computing for systems medicineabstractSystems medicine holds many promises, but has so far provided only a limited number of proofs of principle. To address this road block, possible barriers and challenges of translating systems medicine into clinical practice need to be identified and addressed. The members of the European Cooperation in Science and Technology (COST) Action CA15120 Open Multiscale Systems Medicine (OpenMultiMed) wish to engage the scientific community of systems medicine and multiscale modelling, data science and computing, to provide their feedback in a structured manner. This will result in follow-up white papers and open access resources to accelerate the clinical translation of systems medicine. Massimiliano Zanin, Ivan Chorbev, Blaz Stres, Egils Stalidzans, Julio Vera, Paolo Tieri, Filippo Castiglione, Derek Groen, Huiru Zheng, Jan Baumbach, Johannes A. Schmid, José Basilio, Peter Klimek, Natasa Debeljak, Damjana Rozman, Harald H. H. W. Schmidt |
Briefings Bioinform. | 9 |
| 2019 | A Comprehensive Study on Predicting Functional Role of Metagenomes Using Machine Learning Methodsabstract"Metagenomics" is the study of genomic sequences obtained directly from environmental microbial communities with the aim to linking their structures with functional roles. The field has been aided in the unprecedented advancement through high-throughput omics data sequencing. The outcome of sequencing are biologically rich data sets. Metagenomic data consisting of microbial species which outnumber microbial samples, lead to the "curse of dimensionality" in datasets. Hence, the focus in metagenomics studies has moved towards developing efficient computational models using Machine Learning (ML), reducing the computational cost. In this paper, we comprehensively assessed various ML approaches to classifying high-dimensional human microbiota effectively into their functional phenotypes. We propose the application of embedded feature selection methods, namely, Extreme Gradient Boosting and Penalized Logistic Regression to determine important microbial species. The resultant feature set enhanced the performance of one of the most popular state-of-the-art methods, Random Forest (RF) over metagenomic studies. Experimental results indicate that the proposed method achieved best results in terms of accuracy, area under the Receiver Operating Characteristic curve (ROC-AUC), and major improvement in processing time. It outperformed other feature selection methods of filters or wrappers over RF and classifiers such as Support Vector Machine (SVM), Extreme Learning Machine (ELM), and k- Nearest Neighbors (k-NN). Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Huiru Zheng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | Bootstrapping analysis of crowdsourced non-expert estimates of the number of calories in photographs of meals
Raymond R. Bond, Anne Moorhead, Huiru Zheng, Patrick McAllister |
BIBM | 3 |
| 2018 | SenseCare: Using Automatic Emotional Analysis to Provide Effective Tools for Supporting
Ryan Donovan, Michael Healy, Huiru Zheng, Felix Engel 0002, Michael Fuchs 0002, Paul Walsh, Matthias L. Hemmje, Paul Mc Kevitt |
BIBM | 3 |
| 2018 | A Machine Learning Emotion Detection Platform to Support Affective Well Being
Michael Healy, Ryan Donovan, Paul Walsh, Huiru Zheng |
BIBM | 4 |
| 2018 | Phylogeny-Aware Deep 1-Dimensional Convolutional Neural Network for the Classification of Metagenomes
Timmy Manning, Jyotsna Talreja Wassan, Cintia C. Palu, Haiying Wang 0001, Fiona Browne, Huiru Zheng, Brian Kelly, Paul Walsh |
BIBM | 6 |
| 2018 | eZiGait: Toward an AI Gait Analysis And Sssistant System
Graham McCalmont, Philip J. Morrow, Huiru Zheng, Anas Samara, Sara Yasaei, Haiying Wang 0001, Sally I. McClean |
BIBM | 3 |
| 2018 | RADAR: Representation Learning across Disease Information Networks for Similar Disease Detection
Ruiqi Qin 0001, Lei Duan, Huiru Zheng, Jesse Li-Ling, Kaiwen Song, Xuan Lan |
BIBM | 3 |
| 2018 | Correlation Model Analysis of Nitrogen Addition and Tan Sheep Grazing Effects on Soil Bacterial Community in the Loess Plateau, China
Fujiang Hou, Huiru Zheng, Haiying Wang 0001 |
BIBM | 3 |
| 2018 | PAAM-ML: A novel Phylogeny and Abundance aware Machine Learning Modelling Approach for Microbiome Classification
Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Huiru Zheng |
BIBM | 4 |
| 2018 | An Integrative Framework for Functional Analysis of Cattle Rumen Microbiomes
Jyotsna Talreja Wassan, Huiru Zheng, Fiona Browne, Jenna Bowen, Paul Walsh, Rainer Roehe, Richard J. Dewhurst, Cintia C. Palu, Brian Kelly, Haiying Wang 0001 |
BIBM | 2 |
| 2018 | Multiscale Computing in Systems Medicine: a Brief Reflection
Huiru Zheng, Jyotsna Talreja Wassan, Mihnea Alexandru Moisescu, Lacramioara Stoicu-Tivadar, João Miranda, Mihaela Marcella Vida, Ioan Stefan Sacala, Almir Badnjevic, Ivan Chorbev, Boro Jakimovski |
BIBM | 1 |
| 2018 | CASNMF: A Converged Algorithm for symmetrical nonnegative matrix factorization
Liping Tian 0001, Ping Luo 0003, Haiying Wang 0001, Huiru Zheng, Fang-Xiang Wu |
Neurocomputing | 4 |
| 2018 | Three-Dimensional Dynamic Simulation System for Forest Surface Fire Spreading PredictionabstractForest fire is one of the most frequent, fast spreading and destructive natural disasters. Many countries have developed their own fire prediction model and computational systems to predict the fire spreading, however, the user interaction, display effect and prediction accuracy have not yet met the requirements for firefighting in real forest fire events. The forest fire spreading is a complex process affected by multi-factors. Understanding the relationships between these multi-factors and the forest fire spreading trend is vital to predicting the fire spreading promptly and accurately to make the strategy in extinguishing the forest fire. In this paper, we propose and develop a three-dimensional (3D) forest fire spreading simulation system, FFSimulator, to visualize the impact of multi-factors to the fire spread. FFSimultor integrates the multi-factor analysis approach with the FARSITE prediction model to improve the prediction. The FFSimulator developed applies 3D scene organization, template-based vector data mapping and overlaps visualization techniques to provide a 3D dynamic visualization of large-scale forest fire. The 3D multi-factors superposition analysis simulates the impacts of individual factor and multi-factors on the trend of surface fire spreading, which can be used to identify the key sites for the prevention and the control of forest fires. The system has been tested and evaluated using real data of Shanghan forest fire. Chongchen Chen, Huiru Zheng, Naiyuan Liu |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2017 | Assessing impacts of data volume and data set balance in using deep learning approach to human activity recognitionabstractOver the past decade, deep learning developed rapidly and had significant impact on a variety of application domains. It has been applied to the field of human activity recognition to substitute for well-established analysis techniques that rely on handcrafted feature extraction and classification methods in recent years. However, less attentions have been paid to the influence of training data on recognition accuracy. In this paper, we assessed the influence factors of data volume and data balance in human activity recognition when using deep learning approaches. We evaluated the relationship between data volumes of training dataset and predict accuracy of deep learning algorithms. Given the impact of the data balance between activity categories on the recognition accuracy, we modified the SMOTE algorithm so that it can be applied to human activity recognition. Results show that when the data volume is small (<;4M), the recognition accuracy increased quickly with the increase of the quantity of training data. However, the growth trend of recognition accuracy slows down when the data quantity reaches 4 million. Further increase the data volume does not significantly improve the activity recognition performance. So we can conclude that 4 million data volume can ensure a sufficient accuracy for human activity recognition. Meanwhile, the data set balance operation can not only improve the recognition accuracy of minority categories, but also helps to increase the overall accuracy. Fuhai Xiong, Dihong Wu, Lingxiang Zheng, Ao Peng, Xuemin Hong, Biyu Tang, Haibin Shi, Huiru Zheng |
BIBM | 10 |
| 2017 | Machine learning approaches for cyanobacteria bloom prediction using metagenomic sequence data, a case studyabstractCyanobacteria bloom is a serious public health threat and a global challenge. Literature on the bloom prediction and forecasting has been accumulating and the emphasis appears to have been on the relation between the blooms and environmental factors, whilst the complexity of the bloom mechanism makes it difficult to reach adequate output of the models. Rapid development of next generation sequencing techniques provides a way in which comprehensive and quick examination of the microbial community can be achieved, especially for the bloom community structure. This facilitates using of merely the sequence data along with the machine learning techniques to predict and forecast the bloom occurrence. But there has been rare report on this theme in the literature. In this case study, machine learning approaches were applied with the metagenomic data as the only input (rather than with environmental data) to predict the Cyanobacteria blooms. k-NN classification, SVM classification and k-means clustering were applied and their efficiencies were evaluated using relevant indices. Feature selection was performed and the yielded sub datasets were worked on seriatim. In the predicting experiment with k-NN approach, the final year's data among the 8 years OTU time series were used as target data and various combination of the preceding years' data were used as predictor data; the output came with the best values of 1.00 and 100% for the evaluation indices F1 score and sensitivity, specificity, precision, and accuracy, for the 7 preceding years' predictor input, among the experiment results. This case study demonstrated the feasibility of using machine learning approaches in the Cyanobacteria bloom prediction with only metagenomic sequence data, and the importance of feature selection processing in obtaining better output of the machine learning approaches. The metagenomic data based machine learning approaches are efficient, economic, and faster, possessing the advantage and potential for being adopted as a promising means in the bloom prediction practice. Jiandong Huang, Huiru Zheng, Haying Wang, Xingpeng Jiang |
BIBM | 2 |
| 2017 | The modularity of microbial interaction network in healthy human saliva: Stability and specificityabstractThe human oral cavity is an important habitat of microbes in the human body. It includes the colonization of various microorganisms such as bacteria, archaea, fungi, protozoa and viruses. Although oral diseases have been studied for decades, we have limited understanding of the boundaries of a healthy oral ecosystem and ecological shift toward dysbiosis. Here, we analyzed salivary microbiomes from 268 healthy adults after overnight fasting. The microbiome data set is firstly divided into five sample clusters based on the similarity pattern of microbial abundance. For each cluster, the correlation networks among salivary bacteria are constructed based on an ensemble of six correlations and two dissimilarity measure. The stability and specificity of modularity in the five microbial networks are investigated. The existences of conserved and changing modules were found across five microbial correlation networks. Xingpeng Jiang, Huiru Zheng, Haiying Wang 0001, Tingting He 0003, Xiaohua Hu 0001 |
BIBM | 3 |
| 2017 | Automated adjustment of crowdsourced calorie estimations for accurate food image loggingabstractObesity is increasing globally and is a risk factor for many chronic conditions such as such as heart disease, sleep apnea, type-2 diabetes, and some cancers. Research shows that food logging is beneficial in promoting weight loss. Crowdsourcing has also been used in promoting dietary feedback for food logging. This work investigates the feasibility of crowdsourcing to provide support in accurately determining calories in meal images. Two groups, 1. experts and 2. non-experts, completed a calorie estimation survey consisting of 15 meal images. Descriptive statistics were used to analyse the performance of each group. Collectively, non-experts could determine which meals had larger amounts of calories and analysis showed that meals with greater calories resulted in greater standard deviations of non-expert estimates. Secondary experiments were completed that used crowdsourcing to adjust user calorie estimations using non-expert calorie estimations. Five-fold cross validation was used and results from the calorie adjustment process show a reduced overall mean calorie difference in each fold and the mean error percentage decreased from 40.85% to 25.52% in comparing original mean estimations against adjusted mean estimations. As such, there is credibility in adjusting calorie estimates from a crowd as opposed to simply taking a central measure such as the mean. Patrick McAllister, Anne Moorhead, Raymond R. Bond, Huiru Zheng |
BIBM | 4 |
| 2017 | Participatory design-based requirements elicitation involving people living with dementia towards a home-based platform to monitor emotional wellbeingabstractWe are living in an ageing population with an escalation in chronic illnesses including dementia and other age related diseases. People living with dementia often continue to live at home and are supported by caregivers and next of kin. It is often important to monitor the wellbeing of people living with dementia in order to measure their level of independence and to provide proper support at the time of need as well as supporting their quality of life. Some researchers have focused on monitoring physical wellbeing and activities of daily living (ADL). However, there has been a paucity of research focussed on monitoring mood, affect and the emotional wellbeing of people living with dementia, despite these people experiencing frustration, agitation, depression and social isolation to name but a few known effects. As a result, the SenseCare project aims to build an affective computing platform that uses sensors placed in the home environment to monitor moods, affect and the emotional wellbeing of people living with dementia. This platform is being iteratively designed and will likely use plug-n-play sensors such as passive infrared, wearables and camera technologies to infer emotions from facial expressions, voice intonations and physical behaviour and other modalities. However, it is important to interact iteratively with people living with dementia and their caregivers in order to understand their profound needs. In this study, we report on two focus groups that were conducted to elicit user stories and eventual requirements for the SenseCare platform. Since participatory design involving people living with dementia could bring about unique challenges, we adopted a dyad approach where a caregiver and the person living with dementia participate together in the focus group. This ensures that their needs are fully represented and that consent is fully transparent. In this paper, we report the personal stories elicited during these discussions which will ultimately inform the implementation of the SenseCare platform. Maurice D. Mulvenna, Huiru Zheng, Raymond R. Bond, Patrick McAllister, Haiying Wang 0001, Ruben Riestra |
BIBM | 2 |
| 2017 | A metagenomics analysis of rumen microbiomeabstractClimate change and food security are significant global challenges facing society. The dairy industry is inextricably linked to these challenges as it is concerned with the economies of food production, while acknowledging that it is a major contributor to greenhouse gas production. Action by microbial communities in the rumen is responsible for efficient breakdown of plant matter for food conversion, but a by-product of this action is substantial methane production. Insight into food conversion and methane production in rumen microbiota is possible through metagenomics analysis, which is the analysis of microbial communities and their interactions with the environment. However, metagenomic analysis is hampered by the sheer volume and complexity of data that needs to be processed. This paper presents a bioinformatics pipeline and visualisation platform that facilitates deep analysis of microbial communities, under various conditions in cattle rumen, with the aim of leading to significant impact on probiotic supplement usage, methane production and feed conversion efficiency. This pipeline was developed as part of the EU H2020 MetaPlat project and will pave the way for a more optimal usage of metagenomic datasets, thus reducing the number of animals necessary to be engaged in such studies. This will ensure better and more economic animal welfare, better use of resources and lessen the impact of the dairy industry on climate change. Paul Walsh, Cintia C. Palu, Brian Kelly, Brendan Lawlor, Jyotsna Talreja Wassan, Huiru Zheng, Haiying Wang 0001 |
BIBM | 6 |
| 2017 | Microbial co-presence and mutual-exclusion networks in the Bovine rumen microbiomeabstractThe recognized significance of rumen microbiome has inspired efforts to examine the composition of rumen microbial communities in a large scale. One of the key research areas is to infer association and dependencies between members of rumen microbial communities through correlation analysis. However, it has been found that due to the compositional nature of data, simply applying correlation-based techniques to the analysis of relative abundance of microbial genes may produce artefactual correlation and loss of information. In an attempt to mitigate the compositional effect on the analysis of rumen microbiome data, this study applied a framework including a compendium of two correlation measures and three dissimilarity metrics that are intrinsically robust to compositionality. Based on the inference of significant positive and negative associations, co-presence and mutual-exclusion networks were constructed. The corresponding modules associated with methane production were identified. The modules are highly enriched with microbial genes associated with methane emissions and encoding enzymes involved in the methane methanogensis pathway. In comparisons to previous studies, our analysis demonstrates that deriving microbial associations based on the correlations between relative abundances may not only lead to missing information but also produce spurious associations. Haiying Wang 0001, Huiru Zheng, Richard J. Dewhurst, Rainer Roehe |
BIBM | 2 |
| 2017 | Microbial abundance analysis and phylogenetic adoption in functional metagenomicsabstractMetagenomics is an unobtrusive science of studying uncultivated microbes sampled directly from an environment, e.g. soil, ocean, air, human body, or animals, etc. Functional metagenomics particularly deals with linking microbes to environmental derivations, such as classifying the role of human gut microbiome into a diseased or non-diseased state. Ongoing research in this area includes analyzing the structure of microbial communities, and relate it to functional analysis. We present an integrative experimental framework for functional metagenomics, including data driven (abundance count of microbial species) and knowledge driven (phylogenetic tree structure) contexts. Our related experiments, indicate that i) feature selection improves the performance of classifying human microbiome samples, ii) the classification of human microbiome remains a challenging problem while incorporating phylogenetic structures. For example, our best accuracy attained on the Costello body site (CBH) dataset with forehead and external ear as body sites, is 89.13 % with a non-phylogenetic model, and 78.26 % with a phylogenetic model. This forms a potential research direction of further exploration of space for incorporating phylogeny in microbial analysis and hence developing integrative computational models for deriving functional phenotypes, based on metagenomic sequencing data. Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Huiru Zheng |
CIBCB | 4 |
| 2017 | An Integrative Approach for the Functional Analysis of Metagenomic Studies
Jyotsna Talreja Wassan, Haiying Wang 0001, Fiona Browne, Paul Walsh, Brian Kelly, Cintia C. Palu, Nina Konstantinidou, Rainer Roehe, Richard J. Dewhurst, Huiru Zheng |
ICIC (2) | 10 |
| 2017 | Research on multiple gait and 3D indoor positioning systemabstractHigh accuracy in indoor navigation with foot-mounted sensors attracts a lot of researchers in the last decades. Most indoor positioning schemes based on strap-down inertial navigation can only be used for normal walking. This paper present a 3D foot-mounted inertial navigation system, which can meet the challenge of the multi-gaits. During walking, the foot will have a contact with the ground in every step, in which time, the velocity of foot is zero. The correctness of zero velocity detection is important for drift removing in pedestrian dead-reckoning based inertial pedestrian indoor position systems. Previous algorithm of zero velocity detection is hard to handle the gaits variety. In this paper, by analyzing the inertial data from different modes of motion, a heuristic zero-velocity detection algorithm is designed. The algorithm can accurately detect the zero-velocity time of pedestrians among a variety of gaits. Then the speed and the displacement are updated in the Kalman Filter. Moreover, the barometer is fused with accelerometer for the calculation of height and achievement the 3D trajectory tracking. The experimental results show that the average distance error is 2.59%, the average distance error is 5.78% during running and the average height error is about 0.2m when the pedestrian is going stairs. Rongxin Wang, Lingxiang Zheng, Dihong Wu, Ao Peng, Biyu Tang, Haibin Shi, Huiru Zheng |
IPIN | 8 |
| 2017 | A smart-phone based hand-held indoor tracking systemabstractA smart-phone based hand-held indoor positioning system is presented in this paper. The system collects data using the accelerometers, gyroscopes, barometers and gravity sensors embedded in the smart-phone. The accelerometer and gravity data are used for zero-velocity detection and calculating the vertical displacement of each walking step, and then the inverted pendulum model is applied to calculate the step length of every step. The angle of direction is estimated by processing gyroscope data with the quaternion method. The step length and the direction angle of each step are combined to determine the coordinates of each step. The barometer is used for measuring the height information. A Kalman filter is used in zero-velocity-update (ZUPT) to reduce the vertical speed offset caused by accelerometer drift errors. Wifi is also fused in our system. In order to guarantee the accuracy, map information and magnetic field information are used in the navigation systems. The experiment results show that we obtained high precision results with common hand-held smart-phone often seen on market. Dihong Wu, Ao Peng, Lingxiang Zheng, Zhenyang Wu, Biyu Tang, Haibin Shi, Huiru Zheng |
IPIN | 9 |
| 2017 | Analysis of Organization of the Interactome Using Dominating Sets: A Case Study on Cell Cycle Interaction NetworksabstractIn this study, a minimum dominating set based approach was developed and implemented as a Cytoscape plugin to identify critical and redundant proteins in a protein interaction network. We focused on the investigation of the properties associated with critical proteins in the context of the analysis of interaction networks specific to cell cycle in both yeast and human. A total of 132 yeast genes and 129 human proteins have been identified as critical nodes while 950 in yeast and 980 in human have been categorized as redundant nodes. A clear distinction between critical and redundant proteins was observed when examining their topological parameters including betweenness centrality, suggesting a central role of critical proteins in the control of a network. The significant differences in terms of gene coexpression and functional similarity were observed between the two sets of proteins in yeast. Critical proteins were found to be enriched with essential genes in both networks and have a more deleterious effect on the network integrity than their redundant counterparts. Furthermore, we obtained statistically significant enrichments of proteins that govern human diseases including cancer-related and virus-targeted genes in the corresponding set of critical proteins. Huiru Zheng, Haiying Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Situation Awareness Inferred From Posture Transition and Location: Derived From Smartphone and Smart home SensorsabstractSituation awareness may be inferred from user context such as body posture transition and location data. Smartphones and smart homes incorporate sensors that can record this information without significant inconvenience to the user. Algorithms were developed to classify activity postures to infer current situations; and to measure user's physical location, in order to provide context that assists such interpretation. Location was detected using a subarea-mapping algorithm; activity classification was performed using a hierarchical algorithm with backward reasoning; and falls were detected using fused multiple contexts (current posture, posture transition, location, and heart rate) based on two models: “certain fall” and “possible fall.” The approaches were evaluated on nine volunteers using a smartphone, which provided accelerometer and orientation data, and a radio frequency identification network deployed at an indoor environment. Experimental results illustrated falls detection sensitivity of 94.7% and specificity of 85.7%. By providing appropriate context the robustness of situation recognition algorithms can be enhanced. Shumei Zhang, Paul J. McCullagh, Huiru Zheng, Chris D. Nugent |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2016 | A network analysis of methane and feed conversion genes in the rumen microbial communityabstractMetagenomics involves the genetic analysis of microbial DNA extracted from communities in an environment sample. Advent and falling costs of next-generation sequencing technologies has accelerated metagenomics research providing an improved understanding of microbial communities. In this study we investigate if the traits methane production and feed conversion rates in the rumen microbial community overlap with top genes ranked by topological metrics in a co-abundance network. A co-abundance network was constructed from abundance values of 1570 microbial genes in rumen samples of 8 cattle identified in a metagenomics study at the Beef and Sheep Research Centre of Scotland's Rural College. We used 4 different topological measures: Degree Centrality, Betweenness Centrality, Bonacich Power Centrality and PageRank to the network. Using permutation testing, we discovered, methane production trait genes significantly overlapped with top ranked genes obtained using the metrics PageRank and Bonacich Power Centrality. Feed conversion trait genes overlapped with top ranked genes using Bonacich Power Centrality and Betweenness. Furthermore, we observed the top ranked genes from PageRank and Bonacich Power Centrality significantly overlapped with genes involved in the KEGG methane metabolism pathway and ranked highly key methanogenesis genes such as mcrA and fmdB. Identified functional clusters containing most methane and feed conversion genes were also analyzed in terms of overlap with top ranked genes from topological metrics. Fiona Browne, Haiying Wang 0001, Huiru Zheng, Rainer Roehe, Richard J. Dewhurst, Paul Walsh |
BIBM | 3 |
| 2016 | The role of high performance, grid and cloud computing in high-throughput sequencingabstractWe have reached the era of full genome sequencing using high throughput sequencing technologies pouring out gigabases of reads in a day. To fully benefit from such a profusion of data high performance tools and systems are needed to extract the information lying within the sequences. This paper provides an overview of the evolution of high-throughput sequencing and the tools, infrastructure and data management developing in this space to support a key area in personalized medicine. The paper concludes by providing an outlook in the future of such technologies and their applications and how they might shape clinical governance. Gaye Lightbody, Fiona Browne, Huiru Zheng, Valeriia Haberland, Jaine K. Blayney |
BIBM | 3 |
| 2016 | Analysis of rumen microbial community in cattle through the integration of metagenomic and network-based approachesabstractA better understanding of the composition of rumen microbial communities and the association between host genetic and microbial activities has important applications and implication in bioscience. Being capable of revealing the full extent of microbial gene diversity, metagenomics-based approaches hold great promises in this endeavor. This study investigates the rumen microbial community in cattle through the integration of metagenomic and network-based approaches. Based on the relative abundance of 1570 microbial genes identified in a metagenomics analysis, the co-abundance network was constructed and functional modules of microbial genes were identified. One of the main contributions of this study is to develop a random matrix theory-based approach to automatically determine the correlation threshold used to construct the co-abundance network. It has been shown that the network exhibits a highly modular structure with each of the three main modules well separated. The involvement of KEGG pathways in each module was analysed. A close look at the abundance profiles highlights that Module B is strongly associated with methane emissions while Module C is highly enriched with microbial genes associated with feed conversion efficiency. Haiying Wang 0001, Huiru Zheng, Fiona Browne, Rainer Roehe, Richard J. Dewhurst, Felix Engel 0002, Matthias L. Hemmje, Paul Walsh |
BIBM | 2 |
| 2016 | Modelling enteric methane emissions from milking dairy cows with Bayesian networksabstractAs one of the potent greenhouse gases, methane emission from ruminants has been intensively studied over the past decades. Various regression-based models have been applied to examine factors affecting enteric methane emission. Based on Bayesian networks, this paper proposes an alternative network-based approach to model the relationship among factors affecting enteric methane emissions from milking cows. It was evaluated on the dataset consisting of 934 milking dairy cows collected at Agri-Food and Biosciences Institute, Northern Ireland. The preliminary results demonstrated that the proposed model has a great potential to capture the complex relationship among factors and establish causal influence among predictors. To the best of our knowledge, this is the first study to use Bayesian networks to model causal influence among factors associated with enteric methane emission from milking cows. Huiru Zheng, Haiying Wang 0001, Tianhai Yan |
BIBM | 1 |
| 2015 | Assessment of gait patterns of chronic low back pain patients: A smart mobile phone based approachabstractChronic low back pain is a common and costly condition and has been shown to affect gait. This paper describes the use of gait analysis as measured by a smart phone in a group of chronic low back pain subjects. Reliability of features extracted from the smart phone sensors was investigated using a mutual information based minimum redundancy and maximum relevance feature selection method to identify a key feature set related to lower back pain. This analysis was carried out using a KStar classification model. Results indicate the feasibility of reducing gait features to 6 key components while still achieving very promising classification accuracy (92.50%). The results also demonstrated that it is feasible to use a smart mobile phone in gait tele-monitoring and tele-assessment suggesting potential as both a prognostic and potential treatment outcome. In addition, we show that predicting context such as age and gender using smart mobile phones is achievable, which has potential to provide personalised services and context-related monitoring and intervention. Herman Chan, Huiru Zheng, Haiying Wang 0001, Dave Newell |
BIBM | 2 |
| 2015 | Combining AR filter and sparse Wavelet representation for P300 spellerabstractA variety of experimental paradigms have been proposed in the field of Brain-Computer Interface(BCI). Among them, the P300 speller allows participators to input characters to a computer directly from their own brains. Estimating available features of P300 from raw electroencephalogram(EEG) is a key step of implementing P300 speller. In this paper, a novel combination of Autoregressive model and sparse Wavelet representation is proposed to estimate the P300 features in raw EEG acquired from the P300 speller experiments. Instead of superposition, the P300 features are estimated from raw EEG of single trial in this way. By introducing this method to process signals for BCI, the number of repeated trials may be reduced so that the information transfer rate of P300 speller could be remarkably improved. The proposed approach was tested in off-line data. The results show that the number of repeated trials for a wanted character could be reduced to 4 in general when the feature estimation method is used together with the linear discriminant functions. Huiru Zheng |
BIBM | 2 |
| 2015 | Integrating omics data for identifying disease subtypes: A multiplex network-based approachabstractIt is widely acknowledged that technologies centred on the integration of omics data could play an important role in capturing heterogeneity of phenotypes and identifying disease subtypes. This paper proposed a multiplex network-based approach for integrative analysis of heterogeneous omics data. It represents a useful alternative network-based solution to the problem and a significant step forward to the methods in which each type of data is treated independently. It has been tested on the identification of the subtypes of glioblastoma multiforme. Results obtained have shown that it can achieve comparable performance in comparison to state-of-the-art techniques. The proposed methodology has several advantages. It provides a flexible platform to integrate different types of patient data, potentially from multiple sources, allowing discovering complex disease patterns with multiple facets. Haiying Wang 0001, Huiru Zheng |
BIBM | 2 |
| 2015 | Thermal sensor based multi-occupancy motion tracking and visualisation in smart environmentsabstractA smart environment is a physical space where smart devices/technologies are applied to continuously sense the occupant's daily living activities or health condition. Recent years have witnessed the application of smart environments to support independent living for elderly people or for people with chronic conditions. Nevertheless, most of the projects have been exclusively designed to support instances of single occupancy. In an effort to move beyond the scenario of a single occupant within a smart environment, this paper investigates the feasibility of using thermal sensors to detect the presence of multi-occupants and to track and visualise their motions. Results indicate that the use of thermals sensor can detect multi-occupancy, including moving subjects in addition to static subjects. The system developed demonstrated the ability to track up to 5 occupants, although mistrackings were found to occur when the individuals were apart after they stayed closely together. Huiru Zheng, Haiying Wang 0001, Jonathan Synnott, Chris D. Nugent, Paul Jeffers 0001 |
BIBM | 2 |
| 2015 | Design of a smart insole for ambulatory assessment of gaitabstractIn this paper, we present the design and development of a smart insole that may be used to assess long term chronic conditions that affect the elderly population such as Stroke, Dementia, Parkinson's disease, Cancer, Cardiac Disease and Diabetes. This smart insole offers the potential for evidence base rehabilitation. The ICT solution detect the plantar foot pressure in a free living context through the integration of piezo sensors, microcontroller and Bluetooth technology to empirically measure the pressure at important pressure points. The insole consists of 32 piezo sensors, 01 tri-axial accelerometers, temperature sensor and force sensor to automatically switch ON/OFF the insole. The accelerometers provide context for orientation. The design comprises two flexible PCBs encased in a padded layer, in order to protect the sensors and provide comfort to wearer. Y. S. Ashad Mustufa, John Barton, Brendan O'Flynn, Richard J. Davies, Paul J. McCullagh, Huiru Zheng |
BSN | 6 |
| 2015 | Home-Based Self-Management of Dementia: Closing the Loop
Timothy Patterson, Ian Cleland, Phillip J. Hartin, Chris D. Nugent, Norman D. Black, Mark P. Donnelly, Paul J. McCullagh, Huiru Zheng, Suzanne McDonough |
ICOST | 8 |
| 2015 | A semi-automated food voting classification system: Combining user interaction and Support Vector MachinesabstractObesity is prevalent worldwide including UK and Ireland, affecting all demographics. Obesity can have a detrimental affect on an individual's health, which can lead to chronic conditions. Different digital interventions have enabled users to photograph food items to be identified using different feature extraction methods. In this research, we proposed a system that allows users to draw a polygon around a food item for segmentation. After segmented, the region is then classified using an automated voting system. Different features will then be extracted from the specified area. Support Vector Machines will be issued for each feature type. This system is a proof-of-concept and is designed to research the effectiveness of employing multiple feature detection algorithms to classify food images. To classify food regions a Bag-of-features (BoFs) approach will be used for each. Speeded Up Robust Features point detection and descriptors was used along with colour spatial features, and also MSER region detection with SURF. Each of these methods will have their own BoF to train an SVM. The aim of this research was to create a voting classification system that utilises each feature detection algorithm to ultimately identify the segmented food region through plurality (or majority) vote. Testing showed that the system achieved 75% accuracy when combining each feature SVM to create a voting system. The system outperforms two of the feature classifiers (SURF and MSER with SURF). LAB colour classifier slightly outperformed the voting mechanism within the developed system. In regards to future work, further development and testing would be completed through increasing the variety of food items used in the training phase and a larger test dataset would also be used. Patrick McAllister, Huiru Zheng, Raymond R. Bond, Anne Moorhead |
ISTAS | 2 |
| 2015 | Identification of Protein Complexes from Tandem Affinity Purification/Mass Spectrometry Data via Biased Random WalkabstractSystematic identification of protein complexes from protein-protein interaction networks (PPIs) is an important application of data mining in life science. Over the past decades, various new clustering techniques have been developed based on modelling PPIs as binary relations. Non-binary information of co-complex relations (prey/bait) in PPIs data derived from tandem affinity purification/mass spectrometry (TAP-MS) experiments has been unfairly disregarded. In this paper, we propose a Biased Random Walk based algorithm for detecting protein complexes from TAP-MS data, resulting in the random walk with restarting baits (RWRB). RWRB is developed based on Random walk with restart. The main contribution of RWRB is the incorporation of co-complex relations in TAP-MS PPI networks into the clustering process, by implementing a new restarting strategy during the process of random walk. Through experimentation on un-weighted and weighted TAP-MS data sets, we validated biological significance of our results by mapping them to manually curated complexes. Results showed that, by incorporating non-binary, co-membership information, significant improvement has been achieved in terms of both statistical measurements and biological relevance. Better accuracy demonstrates that the proposed method outperformed several state-of-the-art clustering algorithms for the detection of protein complexes in TAP-MS data. Bingjing Cai, Haiying Wang 0001, Huiru Zheng, Hui Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | COPD lifestyle support through self-management (CALS)abstractA dramatic shift in population ageing is taking place, which is subsequently leading to a rise in the prevalence of chronic conditions. COPD is one such condition that places a heavy burned on national Health Services. A healthier lifestyle can greatly reduce the speed of conditional decline in those with COPD and therefore provide a better quality of life. This paper investigates how technology can be utilized for lifestyle support through self-management providing participants with the ability to monitor information on their lifestyle, behavior and physiological parameters. This information is analyzed and presented to the participant providing a comprehensive picture of their health and wellbeing. With this information at hand, informed lifestyle decisions and behavioral changes can be made and lead to a better quality of life. Mark Beattie, Huiru Zheng, Chris D. Nugent, Paul J. McCullagh |
BIBM | 2 |
| 2014 | An integrative network-driven pipeline for the prioritization of Alzheimer's disease genesabstractLarge-scale, high-throughput technologies and genome-wide studies have been pivotal in the identification of disease-gene candidates from patient cohorts. Output from these studies often result in gene candidate lists which are large in size. Therefore, there is a pressing need for computational tools to integrate heterogeneous data and prioritize disease-gene candidates for further experimental investigation. To address this need, we propose a computational pipeline for the prioritization of disease-gene candidates. Our pipeline integrates diverse heterogeneous data including: gene-expression, protein-protein interaction network, ontology-based similarity and betweenness measures. Furthermore, we incorporate tissue-specific gene expression data into the evaluation section of our approach. The pipeline was applied to prioritize Alzheimer's Disease (AD) genes, whereby a list of 31 prioritized genes was generated. This approach correctly identified key AD susceptible genes: INPP5D and PSEN1. Biological process enrichment analysis revealed the prioritized genes are modulated in AD pathogenesis including: regulation of neurogenesis and generation of neurons. KEGG pathway analysis identified significant hub involvement in the Neurotrophin signaling and Huntington Disease pathways. Furthermore, our evaluation demonstrated a relatively high predictive performance (AUC: 0.73) when classifying AD and normal gene expression profiles from individuals using leave-one-out cross validation. This work provides a foundation for future investigation of diverse heterogeneous data integration for disease-gene prioritization. Fiona Browne, Haiying Wang 0001, Huiru Zheng |
BIBM | 3 |
| 2014 | Design and evaluation of a tool for reminiscence of life-logged dataabstractIn this paper we present a design, development and evaluation of a tool to facilitate the reminiscence of life-logged data intended for persons with dementia. Using off-the-shelf technologies such as a smartphone it is possible to effectively record and log a person's daily activates such as places visited and persons interacted with. We developed a reminiscence tool that visualizes life-logging data. The system was evaluated with six healthy participants aged between 24-46 years of age. They recorded approximately 2.5 hours of life-logging data that was later loaded into the developed tool. During the subsequent evaluation, the participant's gaze was tracked using eye-tracking hardware. Results show that the tool was easy to use and navigate through, facilitating the objective identification of events that may be useful for reminiscence. William Burns, Paul J. McCullagh, Chris D. Nugent, Huiru Zheng |
BIBM | 4 |
| 2014 | Drug resistance gene identification algorithm for next-generation sequencing dataabstractIn the 21stcentury, antibiotic resistance has become a crucial and growing phenomenon in contemporary medicine. Multidrug resistant leads that antibiotics cannot be used to treat infections. In this paper, we propose a novel and efficient method to identify drug resistance genes from raw reads produced by next generation sequencing technology for metagenomes. The experimental results show that the proposed method is able to identify the resistance genes of Acinetobacter baumannii, TYTH-1. Guan-Jie Hua, Chuan Yi Tang, Che-Lun Hung, Huiru Zheng |
BIBM | 4 |
| 2014 | A multi-model reverse-engineering algorithm for large gene regulation networksabstractModeling and simulation of gene-regulatory networks (GRNs) has become an important aspect of modern systems biology investigations into mechanisms underlying gene regulation. An important and unsolved problem in this area is the automated inference (reverse-engineering) of dynamic, mechanistic GRN models from time-course gene expression data. The conventional one-stage model inference algorithm determines the values of all model parameters simultaneously. Recently, two-stage algorithms have been proposed to improve the accuracy of the inferred models and the efficiency of the reverse-engineering process. The main objective of this study is to compare the performance of the conventional one-stage and the modern two-stage algorithm, with emphasis on the computational complexity. We explored data generated from artificial and real GRN systems under different experimental conditions and regulatory structure constraints. Our results suggest that the 2-stage approach outperforms the one-stage methods by far in terms of model inference speed without a loss of accuracy. Alexandru E. Mizeranschi, Huiru Zheng, Werner Dubitzky |
BIBM | 3 |
| 2014 | Towards a generic platform for the self-management of chronic conditionsabstractSelf-management is an approach to healthcare which aims to empower individuals to manage their own health conditions. This is of particular importance given the shifting demographics, increased prevalence of chronic conditions and financial austerity facing many countries. In this paper we present our current work on the development of a flexible, generic self-management platform which can be readily extended for specific chronic conditions. A Unified Modelling Language (UML) system class diagram is presented with the class design based upon functional requirements derived from literature and engagement with key stakeholders. We subsequently extend the UML class diagram to demonstrate how the design may be adapted for specific conditions. Timothy Patterson, Ian Cleland, Chris D. Nugent, Norman D. Black, Paul J. McCullagh, Huiru Zheng, Mark P. Donnelly, Suzanne McDonough |
BIBM | 6 |
| 2014 | Minimum dominating sets in cell cycle specific protein interaction networksabstractRecently, scientists start to examine the dynamics of biological networks from a control theory perspective. Based on the determination of minimum dominating sets (MDSets), this paper investigated the properties associated with MDSet proteins in the context of the analysis of protein interaction networks specific to the yeast cell cycle. Statistically significant differences between MDSet and non-MDSet proteins were observed in terms of topological features, Gene Ontology-driven semantic similarities, and the number of protein domains associated with each protein. However, unlike previous studies, MDSet proteins were found to be enriched with essential genes. Furthermore, we constructed and analyzed a PPI network specific to the human cell cycle and highlighted that the distinction between MDSet and non-MDSet proteins is far more complex than that observed in yeast. The system used to determine a minimum dominating set in a protein interaction network was implemented as a user-friendly Java-based plugin for Cytoscape. Haiying Wang 0001, Huiru Zheng, Fiona Browne |
BIBM | 2 |
| 2014 | A low power and high accuracy MEMS sensor based activity recognition algorithmabstractWearable sensors and smart phones have been used in human activity recognitions and can achieve relative high accuracy however the power consumption is also high. In this paper, we propose an activity recognition approach that can achieve high accuracy with low power consumption. Two strategies have been applied to reduce the power consumption. The first strategy is using the hierarchical support vector machine classification algorithm to reduce the computational complexity. The second strategy is to reduce the sensor data sampling rates. Data collected from sensors in low sampling rate were processed using a wider time window for the feature extraction. The experiment results show that the average recognition accuracy of human activities (sitting, standing, walking, and running) in 1 Hz sampling rate can reach 98.50%. It indicates that the proposed approach can effectively extend the battery lifetime while maintaining high prediction accuracy in activity recognition. Shaolin Weng, Luping Xiang, Weiwei Tang, Lingxiang Zheng, Huiru Zheng |
BIBM | 7 |
| 2014 | Human-machine-environment cyber-physical system and hierarchical task planning to support independent livingabstractIn this paper we propose a novel concept of human-machine-environment cyber-physical system (HME-CPS) in smart homes to support independent living of elderly and physically disabled people. Within this system, a hierarchical task planning approach is developed on purpose of combining qualitative reasoning and quantitative calculation via bidirectional data exchanges between C++ and Prolog. A group of experiments are conducted with respect to a housework task. Their results show C++ and Prolog are connected by an software interface design that enables different formed data to be exchanged among C++ and Prolog modules of the HME-CPS, and the complete process of hierarchical task planning is accomplished in both modes of autonomous planning of the intelligent agent and cooperated planning by human-machine collaboration (HMC). Moreover, methodologies and techniques of the proposed HME-CPS are extendable and adaptive to real applications with robustness and scalability. Guixin Wu, Qijie Zhao, Dawei Tu, Huiru Zheng |
BIBM | 5 |
| 2014 | Design and Evaluation of a Smartphone Based Wearable Life-Logging and Social Interaction SystemabstractIn this paper we outline the design, development and evaluation of a smartphone based life-logging and social interaction reminder system intended for use by persons with dementia. By using a smartphone, the wearer's daily activities can be recorded in picture format, along with meta data providing activity levels and location data. In addition to this data, social interactions can also be logged and subsequently identified, using Quick Response (QR) codes. The intervention was evaluated on six healthy participants aged between 24 - 46 years of age who wore the system for 2.5 hours. The qualitative feedback received was that the technology was easy to use and was responsive and accurate at identifying, recording and displaying social interaction data. William Burns, Chris D. Nugent, Paul J. McCullagh, Huiru Zheng |
CBMS | 4 |
| 2014 | Night optimised care technology for users needing assisted lifestylesabstractThere is growing interest in the development of ambient assisted living services to increase the quality of life of the increasing proportion of the older population. We report on the Night Optimised Care Technology for UseRs Needing Assisted Lifestyles project, which provides specialised night time support to people at early stages of dementia. This article explains the technical infrastructure, the intelligent software behind the decision-making driving the system, the software development process followed, the interfaces used to interact with the user, and the findings and lessons of our user-centred approach. Juan Carlos Augusto, Maurice D. Mulvenna, Huiru Zheng, Haiying Wang 0001, Suzanne Martin, Paul J. McCullagh, Jonathan G. Wallace |
Behav. Inf. Technol. | 3 |
| 2014 | Organized Modularity in the Interactome: Evidence from the Analysis of Dynamic Organization in the Cell CycleabstractThe organization of global protein interaction networks (PINs) has been extensively studied and heatedly debated. We revisited this issue in the context of the analysis of dynamic organization of a PIN in the yeast cell cycle. Statistically significant bimodality was observed when analyzing the distribution of the differences in expression peak between periodically expressed partners. A close look at their behavior revealed that date and party hubs derived from this analysis have some distinct features. There are no significant differences between them in terms of protein essentiality, expression correlation and semantic similarity derived from gene ontology (GO) biological process hierarchy. However, date hubs exhibit significantly greater values than party hubs in terms of semantic similarity derived from both GO molecular function and cellular component hierarchies. Relating to three-dimensional structures, we found that both single- and multi-interface proteins could become date hubs coordinating multiple functions performed at different times while party hubs are mainly multi-interface proteins. Furthermore, we constructed and analyzed a PPI network specific to the human cell cycle and highlighted that the dynamic organization in human interactome is far more complex than the dichotomy of hubs observed in the yeast cell cycle. Haiying Wang 0001, Huiru Zheng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | Detection of protein complexes in protein interaction networks is improved through network-driven functional homogeneity analysisabstractThe detection of biologically meaningful clusters in protein interaction networks is crucial in systems biology. Among its applications, it can enable the identification of protein complexes. Notwithstanding significant advances, the detection of meaningful clusters faces important challenges, including the need to aid researchers in the prioritization of hundreds or even thousands of clusters. To address this need, we developed a method for the prioritization of network clusters based on the analysis of their functional homogeneity, Horn. Based on Horn scores, clustering results can be statistically ranked and attention directed toward clusters that are more likely to be biologically meaningful. We tested it on a global human protein-protein interaction network and four network clustering algorithms. Our method substantially reduced the space of potentially spurious clusters. Furthermore, we evaluated its protein complex detection capability on an independent reference dataset of protein complexes. Irrespectively of clustering approach, our approach improved protein complex identification capacity. Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
BIBM | 2 |
| 2013 | Correlating adverse drug reactions with biological pathways in humansabstractIt has been well recognized that adverse drug reactions (ADRs) are a significant cause of morbidity and mortality. There is a growing interest in investigating biological pathways involved in cellular response to drugs. Based on examining the co-occurrence of drugs in pathway activity and ADR profiles, in this paper, we propose a new method to explore the relationship between biological pathways and ADRs at a large scale. Using sparse canonical correlation analysis of 495 drugs with two profiles for 173 pathways and 1385 ADRs, a total of 80 correlated sets of pathways and ADRs were extracted. To evaluate the performance of our method, extracted correlated components were used to retrieve known ADR profiles from drug pathway profiles using a 5-fold cross validation. A relatively high prediction performance (AUC: 0.881) was achieved. This work provides a foundation for future investigation of ADRs in the context of biological pathways under different conditions. Huiru Zheng, Haiying Wang 0001, Hua Xu 0001, Zhongming Zhao, Francisco Azuaje |
BIBM | 1 |
| 2013 | Reverse engineering of gene regulation models from multi-condition experimentsabstractReverse-engineering of quantitative, dynamic gene-regulatory network (GRN) models from time-series gene expression data is becoming important as such data are increasingly generated for research and other purposes. A key problem in the reverse-engineering process is the under-determined nature of these data. Because of this, the reverse-engineered GRN models often lack robustness and perform poorly when used to simulate system responses to new conditions. In this study, we present a novel method capable of inferring robust GRN models from multi-condition GRN experiments. This study uses two important computational intelligence methods: artificial neural networks and particle swarm optimization. Noel Kennedy, Alexandru E. Mizeranschi, Huiru Zheng, Werner Dubitzky |
CIBCB | 4 |
| 2013 | Assessing Gait Patterns of Healthy Adults Climbing Stairs Employing Machine Learning TechniquesabstractSo far, stair climbing has not been studied as extensively as gait has, although the significance of the prevention of falling on stairs has been well recognized. Based on acceleration data taken from 25 healthy subjects climbing up and down a set of 13 stairs with an accelerometer placed on the lumbo-sacral joint, this paper aims to assess gait patterns of younger and older adults climbing stairs using a machine learning approach. A total of 14 gait features were extracted and analyzed. The performance of six representative classification models: Multilayer Perceptron (MLP), KStar, Support Vector Machine (SVM), Naïve Bayesian (NB), C4.5 Decision Trees, and Random Forests were evaluated in terms of their ability to discriminate between younger and older adults climbing up- and downstairs. MLP was found to provide the highest accuracy for classification. Accuracy of 95.7% was found for classifying a subject walking either up or down the stairs and an accuracy of 80.6% for classifying whether the subject was younger or older. An evaluation of individual features showed poor performance of classification for younger and older subjects climbing up- and downstairs, and in most cases failed to distinguish between the two classes. To access which set of features derived from a triaxial accelerometer can better describe the performance differences between younger and older adults climbing up- and downstairs, two feature selection algorithms, sequential feature selection and correlation-based feature selection, were implemented. Results show that 10 features derived from correlation-based feature selection were able to produce a 96.8% accuracy for classification between subjects climbing up and down. A subset of seven features achieved a performance of 84.9% accuracy for classification between younger and older subjects. Herman Chan, Mingjing Yang 0001, Haiying Wang 0001, Huiru Zheng, Sally I. McClean, Roy Sterritt, Ruth E. Mayagoitia |
Int. J. Intell. Syst. | 4 |
| 2012 | Incorporating semantic similarity into clustering process for identifying protein complexes from Affinity Purification/Mass Spectrometry dataabstractThis paper presents a framework for incorporating semantic similarities in the detection of protein complexes from Affinity Purification/Mass Spectrometry (AP-MS) data. AP-MS data is modeled as a bipartite network, where one set of nodes consist of bait proteins and the other set are prey proteins. Pair-wise similarities of bait proteins are computed by combining similarities based on topological features and functional semantic similarities. A hierarchical clustering algorithm is then applied to obtain `seed clusters' consisting of bait proteins. Starting from these `seed' clusters, an expansion process is developed to recruit prey proteins which are significantly associated with bait proteins, to produce final sets of identified protein complexes. In the application to real AP-MS datasets, we validate biological significance of predicted protein complexes by using curated protein complexes. Six statistical metrics have been applied. Results show that by integrating semantic similarities into the clustering process, the accuracy of identifying complexes has been greatly improved. Meanwhile, clustering results obtained by the proposed framework are better than those from several existent clustering methods. Bingjing Cai, Haiying Wang 0001, Huiru Zheng, Hui Wang 0001 |
BIBM | 3 |
| 2012 | Drug-target network in myocardial infarction: A structural analysisabstractThe identification of drug-target interactions is a crucial step in the drug-discovery process. It has been suggested that drug-target interactions are driven by drug-domain interactions. Based on the integration of two recently published datasets, i.e., Drug-target interactions in myocardial infarction (My-DTome) and drug-domain interaction network, this paper reports the association between drugs and protein domains in the context of myocardial infarction (MI). A MI drug-domain interaction network, My-DDome, was constructed. The functional similarity between domains based on their Gene Ontology (GO) annotations was estimated. The association between domains and therapeutic effects was investigated. Lists of GO annotations and Anatomical Therapeutic Chemical classification (ATC) codes highly enriched in My-DDome were identified. We show that drugs acting on blood and blood forming organs (ATC code B) and sensory organs (ATC code S) are significantly enriched in My-DDome (p <; 0.000001). Top enriched GO terms include GO:0003824 (catalytic activity), GO:0008152 (metabolic process) and GO:0030170 (pyridoxal phosphate binding). By incorporating protein domain information into My-DTome, more detailed insights into the interplay between drugs, their known targets and seemingly unrelated proteins are provided. Haiying Wang 0001, Huiru Zheng, Francisco Azuaje, Xing-Ming Zhao |
BIBM | 2 |
| 2011 | A Subarea Mapping Approach for Indoor Localization
Shumei Zhang, Paul J. McCullagh, Chris D. Nugent, Huiru Zheng, Norman D. Black |
ICOST | 4 |
| 2011 | An improved random walk based clustering algorithm for community detection in complex networksabstractIn recent years, there is an increasing interest in the research community in finding community structure in complex networks. The networks are usually represented as graphs, and the task is usually cast as a graph clustering problem. Traditional clustering algorithms and graph partitioning algorithms have been applied to this problem. New graph clustering algorithms have also been proposed. Random walk based clustering, in which the similarities between pairs of nodes in a graph are usually estimated using random walk with restart (RWR) algorithm, is one of the most popular graph clustering methods. Most of these clustering algorithms only find disjoint partitions in networks; however, communities in many real-world networks often overlap to some degree. In this paper, we propose an efficient clustering method based on random walks for discovering communities in graphs. The proposed method makes use of network topology and edge weights, and is able to discover overlapping communities. We analyze the effect of parameters in the proposed method on clustering results. We evaluate the proposed method on real world social networks that are well documented in the literature, using both topological-based and knowledge-based evaluation methods. We compare the proposed method to other clustering methods including recently published Repeated Random Walks, and find that the proposed method achieves better precision and accuracy values in terms of six statistical measurements including both data-driven and knowledge-driven evaluation metrics. Bingjing Cai, Haiying Wang 0001, Huiru Zheng, Hui Wang 0001 |
SMC | 3 |
| 2011 | Systems-based biological concordance and predictive reproducibility of gene set discovery methods in cardiovascular disease
Francisco Azuaje, Huiru Zheng, Anyela Camargo, Haiying Wang 0001 |
J. Biomed. Informatics | 2 |
| 2011 | Feature Selection and Classification in Supporting Report-Based Self-Management for People with Chronic PainabstractChronic pain is a common long-term condition that affects a person's physical and emotional functioning. Currently, the integrated biopsychosocial approach is the mainstay treatment for people with chronic pain. Self-reporting (the use of questionnaires) is one of the most common methods to evaluate treatment outcome. The questionnaires can consist of more than 300 questions, which is tedious for people to complete at home. This paper presents a machine learning approach to analyze self-reporting data collected from the integrated biopsychosocial treatment, in order to identify an optimal set of features for supporting self-management. In addition, a classification model is proposed to differentiate the treatment stages. Four different feature selection methods were applied to rank the questions. In addition, four supervised learning classifiers were used to investigate the relationships between the numbers of questions and classification performance. There were no significant differences between the feature ranking methods for each classifier in overall classification accuracy or AUC ( p > 0.05); however, there were significant differences between the classifiers for each ranking method ( p < 0.001). The results showed the multilayer perceptron classifier had the best classification performance on an optimized subset of questions, which consisted of ten questions. Its overall classification accuracy and AUC were 100% and 1, respectively. Huiru Zheng, Chris D. Nugent, Paul J. McCullagh, Norman D. Black, K. E. Vowles, L. McCracken |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2010 | seGOsa: Software environment for gene ontology-driven similarity assessmentabstractIn recent years there has been a growing trend towards the adoption of ontologies to support comprehensive, large-scale functional genomics research. This paper introduces seGOsa, a user-friendly cross-platform system to support large-scale assessment of Gene Ontology (GO)-driven similarity among gene products. Using information-theoretic approaches, the system exploits both topological features of the GO (i.e., between-term relationships in the hierarchy) and statistical features of the model organism databases annotated to the GO (i.e., term frequency) to assess functional similarity among gene products. Based on the assumption that the more information two terms share in common, the more similar they are, three GO-driven similarity measures (Resnik's, Lin's and Jiang's metrics) have been implemented to measure between-term similarity within each of the GO hierarchies. Meanwhile, seGOsa offers two approaches (simple and highest average similarity) to assessing the similarity between gene products based on the aggregation of between-term similarities. The program is freely available for non-profit use on request from the authors. Huiru Zheng, Francisco Azuaje, Haiying Wang 0001 |
BIBM | 1 |
| 2010 | Activity Monitoring Using a Smart Phone's Accelerometer with Hierarchical ClassificationabstractThis paper presents details of a convenient and unobtrusive system for monitoring daily activities. A smart phone equipped with an embedded 3D-accelerometer was worn on the belt for the purposes of data recording. Once collected the data was processed to identify 6 activities offline (walking, posture transition, gentle motion, standing, sitting and lying). The processing technique adopted a novel hierarchical classification. In the first instance, rule-based reasoning is used to discriminate between motion and motionless activities. Following this the classification process utilizes two multiclass SVM (support vector machines) classifiers to classify the motion and motionless activities, respectively. The classifiers were trained on data from one subject and tested on 10 subjects. The experiments demonstrate that the hierarchical method can reduce misclassification between motion and motionless activities. The average accuracy was improved compared with using a single classifier by using this classification method (82.8% vs. 63.8%), and is important for providing appropriate feedback in free living applications. Shumei Zhang, Paul J. McCullagh, Chris D. Nugent, Huiru Zheng |
Intelligent Environments | 4 |
| 2010 | Ontology- and graph-based similarity assessment in biological networksabstractSUMMARY: A standard systems-based approach to biomarker and drug target discovery consists of placing putative biomarkers in the context of a network of biological interactions, followed by different 'guilt-by-association' analyses. The latter is typically done based on network structural features. Here, an alternative analysis approach in which the networks are analyzed on a 'semantic similarity' space is reported. Such information is extracted from ontology-based functional annotations. We present SimTrek, a Cytoscape plugin for ontology-based similarity assessment in biological networks. AVAILABILITY: http://rosalind.infj.ulst.ac.uk/SimTrek.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
Bioinform. | 2 |
| 2010 | Integration of Gene Ontology-based similarities for supporting analysis of protein-protein interaction networks
Haiying Wang 0001, Huiru Zheng, Fiona Browne, David H. Glass, Francisco Azuaje |
Pattern Recognit. Lett. | 2 |
| 2009 | Assessing the impact of network depth on the analysis of PPI networks: A case studyabstractRecent years have seen a growing interest in the incorporation of protein-protein interaction (PPI) networks to support functional genomic research. Often a default depth is assumed by network inference software. This case study considers the impact of network depth on the analysis of PPI networks using seven proteins known to be relevant to heart failure as inputs into the analysis. This paper analyses how the characteristics of a PPI network vary according to the level examined, suggesting that the investigation of network topology is an essential first step in PPI analysis. The classification of nodes, in terms of degree and betweenness centrality, within the network is also considered. The effect of network depth is also proved to be significant in the identification of potentially essential proteins with large connectivity and/or high betweenness centrality values. Jaine K. Blayney, Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
CIBCB | 3 |
| 2009 | Home Based Self-management of Chronic Diseases
William Burns, Chris D. Nugent, Paul J. McCullagh, Huiru Zheng, Norman D. Black, Peter C. Wright, Gail A. Mountain |
ICOST | 4 |
| 2008 | Signature genes in human heart failure based on gene expression analysis: Can we identify a unique set?abstractDilated Cardiomyopathy is one of leading courses of heart failure. Recent advances in microarray technology have promised significant advantages in understanding the molecular mechanisms underlying dilated cardiomyopathy and heart failure. Several microarray studies have successfully yielded a set of signature genes associated with heart failure. However, it has been found that the overlap of these heart failure associated genes derived from different experiments is very small. Based on the analysis of two publicly available microarray datasets associated with heart failure with three types of machine learning and statistical prediction models, this paper explores this phenomenon. We found that there is no unique set of genes associated with heart failure. Many sets of genes can achieve very high prediction accuracy. In order to identify biomarkers in human heart failure, it may not be sufficient to just focus a certain number of top genes. Such main candidates should be chosen from the much longer list of genes. Haiying Wang 0001, Huiru Zheng |
BIBE | 2 |
| 2008 | Reassessing the limit of data integration for the prediction of protein-protein interactions in Saccharomyces cerevisiaeabstractThis paper investigates the integration of functional genomic data for the prediction of protein-protein interactions (PPI) in Saccharomyces cerevisiae. A previous benchmark study observed a marginal increase in predictive power when integrating diverse features. Classification performance was evaluated using the Receiver Operating Characteristic (ROC) curve. In this study we propose the implementation of a likelihood ratio based Bayesian classifier to reassess the limits of genomic integration. The classifier combines seven genomic features ranging from co-expression to essentiality. Due to the imbalance of the dataset in this study, ROC curves may present an overly optimistic view of the classification performance. We use the true positive/false positive (TP/FP) rate and sensitivity as comparative predictive measures to the ROC curve. Predicted interactions are verified using a Gold Standard constructed from the Munich Database of Interacting Proteins Complex Catalogue. Using the measures TP/FP and sensitivity, a clear increase in classification performance was observed with the integration of features. This framework could be extended to the analysis of PPI in more complex organisms such as Drosophila melanogaster and Homo sapiens. Fiona Browne, Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
CIBCB | 3 |
| 2008 | An Improved Support Vector Machine for the Classification of Imbalanced Biological Datasets
Haiying Wang 0001, Huiru Zheng |
ICIC (1) | 2 |
| 2008 | Poisson-Based Self-Organizing Neural Networks for Pattern Discovery
Haiying Wang 0001, Huiru Zheng |
ICIC (1) | 2 |
| 2008 | Improving Pattern Discovery and Visualization of SAGE Data Through Poisson-Based Self-Adaptive Neural NetworksabstractSerial analysis of gene expression (SAGE) allows a detailed, simultaneous analysis of thousands of genes without the need for prior, complete gene sequence information. However, due to its inherent complexity and the lack of complete structural and function knowledge, mining vast collections of SAGE data to extract useful knowledge poses great challenges to traditional analytical techniques. Moreover, SAGE data are characterized by a specific statistical model that has not been incorporated into traditional data analysis techniques. The analysis of SAGE data requires advanced, intelligent computational techniques, which consider the underlying biology and the statistical nature of SAGE data. By addressing the statistical properties demonstrated by SAGE data, this paper presents a new self-adaptive neural network, Poisson-based growing self-organizing map (PGSOM), which implements novel weight adaptation and neuron growing strategies. An empirical study of key dynamic mechanisms of PGSOM is presented. It was tested on three datasets, including synthetic and experimental SAGE data. The results indicate that, in comparison to traditional techniques, the PGSOM offers significant advantages in the context of pattern discovery and visualization in SAGE data. The pattern discovery and visualization platform discussed in this paper can be applied to other problem domains where the data are better approximated by a Poisson distribution. Huiru Zheng, Haiying Wang 0001, Francisco Azuaje |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2008 | Integration of Genomic Data for Inferring Protein Complexes from Global Protein-Protein Interaction NetworksabstractProtein-protein interactions (PPIs) play crucial roles in virtually every aspect of cellular function within an organism. One important objective of modern biology is the extraction of functional modules, such as protein complexes from global protein interaction networks. This paper describes how seven genomic features and four experimental interaction data sets were combined using a Bayesian-networks-based data integration approach to infer PPI networks in yeast. Greater coverage and higher accuracy were achieved than in previous high-throughput studies of PPI networks in yeast. A Markov clustering algorithm was then used to extract protein complexes from the inferred protein interaction networks. The quality of the computed complexes was evaluated using the hand-curated complexes from the Munich Information Center for Protein Sequences database and gene-ontology-driven semantic similarity. The results indicated that, by integrating multiple genomic information sources, a better clustering result was obtained in terms of both statistical measures and biological relevance. Huiru Zheng, Haiying Wang 0001, David H. Glass |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | Supervised Statistical and Machine Learning Approaches to Inferring Pairwise and Module-Based Protein Interaction NetworksabstractThis paper evaluates three classification techniques: Naive Bayesian (NB), multilayer perceptron (MLP) and K-nearest neighbour (KNN) that integrate diverse, large-scale functional data to infer pairwise (PW) and module-based (MB) interaction networks in Saccharomyces cerevisiae. Existing multi-source functional data from S. cerevisiae were merged and transformed to construct MB datasets. The results indicate that selection of a classifier depends upon the specific PPI classification problem. Feature integration and encoding methods proposed significantly impact the predictive performance of the classifiers. Generation of PPI maps for S. cerevisiae and beyond will be improved with new, high-quality, large-scale datasets with increased interactome coverage and the integration of classification methods. Fiona Browne, Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
BIBE | 3 |
| 2007 | Cluster Analysis of Regulatory Sequences with a Log Likelihood Ratio Statistics-based Similarity MeasureabstractUpstream regions in the DNA sequence are characterized by the presence of short regulatory motifs, which function as target binding sites for transcription factors. Finding two genes with common motifs in their regulatory regions may aid users in identifying co-regulated genes or inferring regulatory modules. By modelling pattern occurrences in the regulatory regions with Poisson statistics, this paper presents a log likelihood ratio statistics-based distance measure to calculate pair-wise similarities between sequences. To perform cluster analysis of regulatory sequences, this paper introduces two clustering algorithms on the basis of the incorporation of the log likelihood ratio statistics-based distance into hierarchical clustering and Self-Organizing Map. The proposed approach has been tested on a synthetic dataset and a real biological example. The results indicate that, in comparison to traditional distance functions, the log likelihood ratio statistics-based similarity measure offers considerable improvements in the process of regulatory sequence-based gene classification. Huiru Zheng, Haiying Wang 0001, Jinglu Hu |
BIBE | 1 |
| 2007 | homeML - An Open Standard for the Exchange of Data Within Smart Environments
Chris D. Nugent, Dewar D. Finlay, Richard J. Davies, Haiying Wang 0001, Huiru Zheng, Josef Hallberg, Kåre Synnes, Maurice D. Mulvenna |
ICOST | 5 |
| 2007 | Self-Adaptive Neural Networks Based on a Poisson Approach for Knowledge Discovery
Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
IJCAI | 2 |
| 2007 | A Hierarchical Learning System Incorporating with Supervised, Unsupervised and Reinforcement Learning
Jinglu Hu, Takafumi Sasakawa, Kotaro Hirasawa, Huiru Zheng |
ISNN (1) | 4 |
| 2007 | Poisson-Based Self-Organizing Feature Maps and Hierarchical Clustering for Serial Analysis of Gene Expression DataabstractSerial analysis of gene expression (SAGE) is a powerful technique for global gene expression profiling, allowing simultaneous analysis of thousands of transcripts without prior structural and functional knowledge. Pattern discovery and visualization have become fundamental approaches to analyzing such large-scale gene expression data. From the pattern discovery perspective, clustering techniques have received great attention. However, due to the statistical nature of SAGE data (i.e., underlying distribution), traditional clustering techniques may not be suitable for SAGE data analysis. Based on the adaptation and improvement of Self-Organizing Maps and hierarchical clustering techniques, this paper presents two new clustering algorithms, namely, PoissonS and PoissonHC, for SAGE data analysis. Tested on synthetic and experimental SAGE data, these algorithms demonstrate several advantages over traditional pattern discovery techniques. The results indicate that, by incorporating statistical properties of SAGE data, PoissonS and PoissonHC, as well as a hybrid approach (neuro-hierarchical approach) based on the combination of PoissonS and PoissonHC, offer significant improvements in pattern discovery and visualization for SAGE data. Moreover, a user-friendly platform, which may improve and accelerate SAGE data mining, was implemented. The system is freely available on request from the authors for nonprofit use. Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2006 | Computational Approaches to Supporting Large-Scale Analysis of Photoreceptor-Enriched Gene ExpressionabstractRetinal photoreceptor cells are responsible for light detection and phototransduction. The understanding of molecular mechanisms regulating photoreceptor gene expression during retinal development may have important implications in clinical neuroscience. Using self-adaptive neural networks and pattern validation statistical tools, this paper explores large-scale analysis of photoreceptor gene expression. Based on the analysis of data generated by serial analysis of gene expression (SAGE) in the developing mouse retina, significant expression patterns for the in silico detection of photoreceptor-enriched genes were revealed. This study demonstrates how machine learning and statistical techniques may be effectively combined to detect key complex relationships encoded in SAGE data. Such approaches may support inexpensive functional predictions prior to the application of experimental methodologies. Haiying Wang 0001, Huiru Zheng, Francisco Azuaje |
CBMS | 2 |
| 2006 | Machine learning approaches to supporting the identification of photoreceptor-enriched genes based on expression dataabstractBACKGROUND: Retinal photoreceptors are highly specialised cells, which detect light and are central to mammalian vision. Many retinal diseases occur as a result of inherited dysfunction of the rod and cone photoreceptor cells. Development and maintenance of photoreceptors requires appropriate regulation of the many genes specifically or highly expressed in these cells. Over the last decades, different experimental approaches have been developed to identify photoreceptor enriched genes. Recent progress in RNA analysis technology has generated large amounts of gene expression data relevant to retinal development. This paper assesses a machine learning methodology for supporting the identification of photoreceptor enriched genes based on expression data. RESULTS: Based on the analysis of publicly-available gene expression data from the developing mouse retina generated by serial analysis of gene expression (SAGE), this paper presents a predictive methodology comprising several in silico models for detecting key complex features and relationships encoded in the data, which may be useful to distinguish genes in terms of their functional roles. In order to understand temporal patterns of photoreceptor gene expression during retinal development, a two-way cluster analysis was firstly performed. By clustering SAGE libraries, a hierarchical tree reflecting relationships between developmental stages was obtained. By clustering SAGE tags, a more comprehensive expression profile for photoreceptor cells was revealed. To demonstrate the usefulness of machine learning-based models in predicting functional associations from the SAGE data, three supervised classification models were compared. The results indicated that a relatively simple instance-based model (KStar model) performed significantly better than relatively more complex algorithms, e.g. neural networks. To deal with the problem of functional class imbalance occurring in the dataset, two data re-sampling techniques were studied. A random over-sampling method supported the implementation of the most powerful prediction models. The KStar model was also able to achieve higher predictive sensitivities and specificities using random over-sampling techniques. CONCLUSION: The approaches assessed in this paper represent an efficient and relatively inexpensive in silico methodology for supporting large-scale analysis of photoreceptor gene expression by SAGE. They may be applied as complementary methodologies to support functional predictions before implementing more comprehensive, experimental prediction and validation methods. They may also be combined with other large-scale, data-driven methods to facilitate the inference of transcriptional regulatory networks in the developing retina. Furthermore, the methodology assessed may be applied to other data domains. Haiying Wang 0001, Huiru Zheng, David Simpson, Francisco Azuaje |
BMC Bioinform. | 2 |
| 2005 | Case-Based Tissue Classification for Monitoring Leg Ulcer HealingabstractThe ability to automatically monitor the wound healing process would reduce the workload of professionals, provide standardization, reduce costs, and improve the quality of care for patients. Here we propose an automatic monitoring system for leg ulcers based on case-based reasoning. We focus on the first stage of the monitoring process in this work, that of tissue classification and examine a number of different feature extraction techniques based on texture and Red, Green, and Blue histograms. Results clearly show a case-based approach to be ideal for this type of task. Mykola Galushka, Huiru Zheng, David Patterson 0002, Lilian Bradley |
CBMS | 2 |
| 2005 | Web-Based Monitoring System for Home-Based Rehabilitation with Stroke PatientsabstractResearch on developing low cost home-based rehabilitation systems aim to provide support for the rehabilitation of post stroke patients in the home environment to promote/aid functional recovery and ultimately enhance their quality of life (QoL). A Web-based system has been proposed for monitoring the home-based rehabilitation and providing both therapeutic instruction and support information. The system will support specific rehabilitation interventions, provide a three-dimensional (3D) visual output and measure the effectiveness of the resulting actions undertaken by the participant. Information regarding process can be reviewed and accessed by the patient, their carers and health professionals. Huiru Zheng, Richard J. Davies, Norman D. Black |
CBMS | 1 |