Rui Yin 0002

dblp:07/2864-2 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-1403-0396ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2025 RNALoc-LM: RNA subcellular localization prediction using pre-trained RNA language model
abstract
MOTIVATION: Accurately predicting RNA subcellular localization is crucial for understanding the cellular functions and regulatory mechanisms of RNAs. Although many computational methods have been developed to predict the subcellular localization of lncRNAs, miRNAs, and circRNAs, very few of them are designed to simultaneously predict the subcellular localization of multiple types of RNAs. In addition, the emergence of pre-trained RNA language model has shown remarkable performance in various bioinformatics tasks, such as structure prediction and functional annotation. Despite these advancements, there remains a significant gap in applying pre-trained RNA language models specifically for predicting RNA subcellular localization. RESULTS: In this study, we proposed RNALoc-LM, the first interpretable deep-learning framework that leverages a pre-trained RNA language model for predicting RNA subcellular localization. RNALoc-LM uses a pre-trained RNA language model to encode RNA sequences, then captures local patterns and long-range dependencies through TextCNN and BiLSTM modules. A multi-head attention mechanism is used to focus on important regions within the RNA sequences. The results demonstrate that RNALoc-LM significantly outperforms both deep-learning baselines and existing state-of-the-art predictors. Additionally, motif analysis highlights RNALoc-LM's potential for discovering important motifs, while an ablation study confirms the effectiveness of the RNA sequence embeddings generated by the pre-trained RNA language model. AVAILABILITY AND IMPLEMENTATION: The RNALoc-LM web server is available at http://csuligroup.com:8000/RNALoc-LM. The source code can be obtained from https://github.com/CSUBioGroup/RNALoc-LM.
Min Zeng 0004, Chengqian Lu, Rui Yin 0002, Fei Guo 0001, Min Li 0007
Bioinform.5
2025 VarPPUD: Pinpointing diagnostic variants from sets of prioritized, strong candidate variants
abstract
Rare and ultra-rare genetic conditions are estimated to impact nearly 1 in 17 people worldwide, yet accurately pinpointing the diagnostic variants underlying each of these conditions remains a formidable challenge. Because comprehensive, in vivo functional assessment of all possible genetic variants is infeasible, clinicians instead consider in silico variant pathogenicity predictions to distinguish plausibly disease-causing from benign variants across the genome. However, in the most difficult undiagnosed cases, such as those accepted to the Undiagnosed Diseases Network (UDN), existing pathogenicity predictions cannot reliably discern true etiological variant(s) from other deleterious candidate variants that were prioritized through case- or family-level analyses. Pinpointing the disease-causing variant from a small pool of plausible candidates remains a largely manual effort requiring extensive clinical workups, functional and experimental assays, and eventual identification of genotype- and phenotype-matched individuals. Here, we introduce VarPPUD, a tool trained on prioritized variants from UDN cases, that leverages gene-, amino acid-, and nucleotide-level features to discern pathogenic (disease causative) variants from other damaging or deleterious variants that are unlikely to be confirmed as relevant to the disease. VarPPUD achieves a cross-validated accuracy of 79.3% and precision of 77.5% on a held-out subset of uniquely challenging UDN cases, respectively representing an average 18.6% and 23.4% improvement over nine existing state-of-the-art pathogenicity prediction tools on this task. We validate VarPPUD's ability to discriminate likely from unlikely pathogenic variants using both synthetic data generated via a GAN-based framework and a temporally held-out set of UDN patients evaluated between 2022 and 2024. The model was trained exclusively on data available through 2021 and applied without retraining to the post-2021 cohort, demonstrating strong generalizability to newly accrued cases. Finally, we show how VarPPUD can be probed to evaluate each input feature's importance and contribution toward prediction-an essential step toward understanding the distinct characteristics of newly-uncovered disease-causing variants.
Rui Yin 0002, Alba Gutiérrez-Sacristán, Shilpa Nadimpalli Kobren, Paul Avillach
PLoS Comput. Biol.1
2024 BertSNR: an interpretable deep learning framework for single-nucleotide resolution identification of transcription factor binding sites based on DNA language model
abstract
MOTIVATION: Transcription factors are pivotal in the regulation of gene expression, and accurate identification of transcription factor binding sites (TFBSs) at high resolution is crucial for understanding the mechanisms underlying gene regulation. The task of identifying TFBSs from DNA sequences is a significant challenge in the field of computational biology today. To address this challenge, a variety of computational approaches have been developed. However, these methods face limitations in their ability to achieve high-resolution identification and often lack interpretability. RESULTS: We propose BertSNR, an interpretable deep learning framework for identifying TFBSs at single-nucleotide resolution. BertSNR integrates sequence-level and token-level information by multi-task learning based on pre-trained DNA language models. Benchmarking comparisons show that our BertSNR outperforms the existing state-of-the-art methods in TFBS predictions. Importantly, we enhanced the interpretability of the model through attentional weight visualization and motif analysis, and discovered the subtle relationship between attention weight and motif. Moreover, BertSNR effectively identifies TFBSs in promoter regions, facilitating the study of intricate gene regulation. AVAILABILITY AND IMPLEMENTATION: The BertSNR source code can be found at https://github.com/lhy0322/BertSNR.
Hanyu Luo, Min Zeng 0004, Rui Yin 0002, Pingjian Ding, Lingyun Luo, Min Li 0007
Bioinform.4
2024 GenoM7GNet: An Efficient N7-Methylguanosine Site Prediction Approach Based on a Nucleotide Language Model
abstract
N-methylguanosine (m7G), one of the mainstream post-transcriptional RNA modifications, occupies an exceedingly significant place in medical treatments. However, classic approaches for identifying m7G sites are costly both in time and equipment. Meanwhile, the existing machine learning methods extract limited hidden information from RNA sequences, thus making it difficult to improve the accuracy. Therefore, we put forward to a deep learning network, called "GenoM7GNet," for m7G site identification. This model utilizes a Bidirectional Encoder Representation from Transformers (BERT) and is pretrained on nucleotide sequences data to capture hidden patterns from RNA sequences for m7G site prediction. Moreover, through detailed comparative experiments with various deep learning models, we discovered that the one-dimensional convolutional neural network (CNN) exhibits outstanding performance in sequence feature learning and classification. The proposed GenoM7GNet model achieved 0.953in accuracy, 0.932in sensitivity, 0.976in specificity, 0.907in Matthews Correlation Coefficient and 0.984in Area Under the receiver operating characteristic Curve on performance evaluation. Extensive experimental results further prove that our GenoM7GNet model markedly surpasses other state-of-the-art models in predicting m7G sites, exhibiting high computing performance.
Chuang Li 0004, Heshi Wang, Yanhua Wen, Rui Yin 0002, Xiangxiang Zeng, Keqin Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 PESI: Paratope-Epitope Set Interaction for SARS-CoV-2 Neutralization Prediction
abstract
Prediction of neutralization antibodies is important for the development of effective vaccines and antibody-based therapeutics. Traditional methods rely on features based on first principles derived from the binding interface. However, they are burdened by arduous data preprocessing from a limited quantity of protein structures. In comparison, deep learning allows automatic substructure characterization and representation without hand-crafted feature engineering. In particular, large language models (LLMs) based method predicts neutralization using Fv sequences of antibody and antigen. Despite LLM’s success, incorporating full-length Fv sequences suffers from: 1) inaccurate sequence-level labels in existing datasets, 2) inefficient modeling due to noisy non-contributing motifs, and 3) ignorance of non-bonded interactions that play a key role in facilitating epitope-paratope pairing. In this paper, we propose a novel approach that incorporates only the paratope and epitope for antibody-antigen neutralization prediction while adopting a novel set modeling that regards the paratope and epitope as bags of residues. Specifically, we hand-crafted a dataset containing neutralizing paratope-epitope pairs where epitopes are potentially generalizable to future unseen variants of SARS-CoV-2. Training on such a dataset enables deep learning models to predict neutralizing antibodies for prospective mutated variants of SARS-CoV-2, meanwhile addressing the problem of inaccurate sequence-level labels. A higher modeling efficiency is also achieved by disregarding non-contributing motifs. Furthermore, we also propose paratope-epitope set interaction (PESI), a set modeling model inspired by first principles that learns intra-inter non-covalent interactions through a global attention mechanism. To validate PESI, we perform a 10-fold cross-validation on our dataset. Experimental results show that PESI achieves a more balanced overall performance and a significant improvement on MCC as compared to existing architectures.
Zhang Wan, Zhuoyi Lin, Shamima Rashid, Shaun Yue-Hao Ng, Rui Yin 0002, J. Senthilnath 0001, Chee Keong Kwoh 0001
BIBM5
2023 HSELDA: Heterogeneous Sub-Graph Learning for lncRNA-Disease Associations Prediction
abstract
Numerous studies have shown that long non-coding RNAs (lncRNAs) can be the key factors in causing a variety of human diseases, such as cancers. Predicting the potential associations between lncRNAs and diseases can be highly beneficial for pathogenesis investigation and understanding the functionalities of lncRNAs in disease development. Existing computational methods for lncRNAs-disease association (LDA) prediction focuses on integrating heterogeneous information from given data including associations of various biological elements. However, many of them do not fully utilize the topological information obtained from graph-structured frameworks, such as heterogeneous graphs consisting of different biological variables including lncRNA, disease, miRNA, and others. We proposed a novel method, named HSELDA, that utilized a relational graph convolutional neural network based on a link prediction technique via subgraphs extraction to predict the potential lncRNAs-disease associations. In this method, we first constructed a heterogeneous graph consisting of three types of nodes: lncRNAs, miRNAs, and disease, with their known associations between each node pair as edges. Then, we extracted a list of local subgraphs around each target link, where the subgraphs preserve rich heuristic information related to link existence with the neighbor nodes directly connected to the target links. We employed a relational graph convolutional network to identify the most promising lncRNAs-diseases associations based on local subgraphs. The experimental results showed that HSELDA was superior to several state-of-the-art methods. Of note, our proposed model shows significant superiority over other compared baseline approaches in AUPRC.
Tiancheng Zhou, Zachary Glanz, Jiang Bian 0001, Rui Yin 0002, Hongyang Gao
BIBM5
2023 GraphLncLoc: long non-coding RNA subcellular localization prediction using graph convolutional networks based on sequence to graph transformation
abstract
The subcellular localization of long non-coding RNAs (lncRNAs) is crucial for understanding lncRNA functions. Most of existing lncRNA subcellular localization prediction methods use k-mer frequency features to encode lncRNA sequences. However, k-mer frequency features lose sequence order information and fail to capture sequence patterns and motifs of different lengths. In this paper, we proposed GraphLncLoc, a graph convolutional network-based deep learning model, for predicting lncRNA subcellular localization. Unlike previous studies encoding lncRNA sequences by using k-mer frequency features, GraphLncLoc transforms lncRNA sequences into de Bruijn graphs, which transforms the sequence classification problem into a graph classification problem. To extract the high-level features from the de Bruijn graph, GraphLncLoc employs graph convolutional networks to learn latent representations. Then, the high-level feature vectors derived from de Bruijn graph are fed into a fully connected layer to perform the prediction task. Extensive experiments show that GraphLncLoc achieves better performance than traditional machine learning models and existing predictors. In addition, our analyses show that transforming sequences into graphs has more distinguishable features and is more robust than k-mer frequency features. The case study shows that GraphLncLoc can uncover important motifs for nucleus subcellular localization. GraphLncLoc web server is available at http://csuligroup.com:8000/GraphLncLoc/.
Min Li 0007, Baoying Zhao, Rui Yin 0002, Chengqian Lu, Fei Guo 0001, Min Zeng 0004
Briefings Bioinform.3
2023 A review of enzyme design in catalytic stability by artificial intelligence
abstract
The design of enzyme catalytic stability is of great significance in medicine and industry. However, traditional methods are time-consuming and costly. Hence, a growing number of complementary computational tools have been developed, e.g. ESMFold, AlphaFold2, Rosetta, RosettaFold, FireProt, ProteinMPNN. They are proposed for algorithm-driven and data-driven enzyme design through artificial intelligence (AI) algorithms including natural language processing, machine learning, deep learning, variational autoencoder/generative adversarial network, message passing neural network (MPNN). In addition, the challenges of design of enzyme catalytic stability include insufficient structured data, large sequence search space, inaccurate quantitative prediction, low efficiency in experimental validation and a cumbersome design process. The first principle of the enzyme catalytic stability design is to treat amino acids as the basic element. By designing the sequence of an enzyme, the flexibility and stability of the structure are adjusted, thus controlling the catalytic stability of the enzyme in a specific industrial environment or in an organism. Common indicators of design goals include the change in denaturation energy (ΔΔG), melting temperature (ΔTm), optimal temperature (Topt), optimal pH (pHopt), etc. In this review, we summarized and evaluated the enzyme design in catalytic stability by AI in terms of mechanism, strategy, data, labeling, coding, prediction, testing, unit, integration and prospect.
Yongfan Ming, Wenkang Wang, Rui Yin 0002, Min Zeng 0004, Shizhe Tang, Min Li 0007
Briefings Bioinform.3
2023 LncLocFormer: a Transformer-based deep learning model for multi-label lncRNA subcellular localization prediction by using localization-specific attention mechanism
abstract
MOTIVATION: There is mounting evidence that the subcellular localization of lncRNAs can provide valuable insights into their biological functions. In the real world of transcriptomes, lncRNAs are usually localized in multiple subcellular localizations. Furthermore, lncRNAs have specific localization patterns for different subcellular localizations. Although several computational methods have been developed to predict the subcellular localization of lncRNAs, few of them are designed for lncRNAs that have multiple subcellular localizations, and none of them take motif specificity into consideration. RESULTS: In this study, we proposed a novel deep learning model, called LncLocFormer, which uses only lncRNA sequences to predict multi-label lncRNA subcellular localization. LncLocFormer utilizes eight Transformer blocks to model long-range dependencies within the lncRNA sequence and shares information across the lncRNA sequence. To exploit the relationship between different subcellular localizations and find distinct localization patterns for different subcellular localizations, LncLocFormer employs a localization-specific attention mechanism. The results demonstrate that LncLocFormer outperforms existing state-of-the-art predictors on the hold-out test set. Furthermore, we conducted a motif analysis and found LncLocFormer can capture known motifs. Ablation studies confirmed the contribution of the localization-specific attention mechanism in improving the prediction performance. AVAILABILITY AND IMPLEMENTATION: The LncLocFormer web server is available at http://csuligroup.com:9000/LncLocFormer. The source code can be obtained from https://github.com/CSUBioGroup/LncLocFormer.
Min Zeng 0004, Yifan Wu 0008, Rui Yin 0002, Chengqian Lu, Junwen Duan, Min Li 0007
Bioinform.4
2023 ViPal: A framework for virulence prediction of influenza viruses with prior viral knowledge using genomic sequences
Rui Yin 0002, Zihan Luo 0001, Pei Zhuang, Min Zeng 0004, Min Li 0007, Zhuoyi Lin, Chee Keong Kwoh 0001
J. Biomed. Informatics1
2023 COMET: Convolutional Dimension Interaction for Collaborative Filtering
abstract
Representation learning-based recommendation models play a dominant role among recommendation techniques. However, most of the existing methods assume both historical interactions and embedding dimensions are independent of each other, and thus regrettably ignore the high-order interaction information among historical interactions and embedding dimensions. In this article, we propose a novel representation learning-based model called COMET (COnvolutional diMEnsion inTeraction), which simultaneously models the high-order interaction patterns among historical interactions and embedding dimensions. To be specific, COMET stacks the embeddings of historical interactions horizontally at first, which results in two “embedding maps”. In this way, internal interactions and dimensional interactions can be exploited by convolutional neural networks (CNN) with kernels of different sizes simultaneously. A fully connected multi-layer perceptron (MLP) is then applied to obtain two interaction vectors. Lastly, the representations of users and items are enriched by the learnt interaction vectors, which can further be used to produce the final prediction. Extensive experiments and ablation studies on various public implicit feedback datasets clearly demonstrate the effectiveness and rationality of our proposed method.
Zhuoyi Lin, Lei Feng 0006, Xingzhi Guo, Yu Zhang 0084, Rui Yin 0002, Chee Keong Kwoh 0001
ACM Trans. Intell. Syst. Technol.5
2022 A framework for predicting variable-length epitopes of human-adapted viruses using machine learning methods
abstract
The coronavirus disease 2019 pandemic has alerted people of the threat caused by viruses. Vaccine is the most effective way to prevent the disease from spreading. The interaction between antibodies and antigens will clear the infectious organisms from the host. Identifying B-cell epitopes is critical in vaccine design, development of disease diagnostics and antibody production. However, traditional experimental methods to determine epitopes are time-consuming and expensive, and the predictive performance using the existing in silico methods is not satisfactory. This paper develops a general framework to predict variable-length linear B-cell epitopes specific for human-adapted viruses with machine learning approaches based on Protvec representation of peptides and physicochemical properties of amino acids. QR decomposition is incorporated during the embedding process that enables our models to handle variable-length sequences. Experimental results on large immune epitope datasets validate that our proposed model's performance is superior to the state-of-the-art methods in terms of AUROC (0.827) and AUPR (0.831) on the testing set. Moreover, sequence analysis also provides the results of the viral category for the corresponding predicted epitopes with high precision. Therefore, this framework is shown to reliably identify linear B-cell epitopes of human-adapted viruses given protein sequences and could provide assistance for potential future pandemics and epidemics.
Rui Yin 0002, Xianghe Zhu, Min Zeng 0004, Min Li 0007, Chee Keong Kwoh 0001
Briefings Bioinform.1
2022 IAV-CNN: A 2D Convolutional Neural Network Model to Predict Antigenic Variants of Influenza A Virus
abstract
The rapid evolution of influenza viruses constantly leads to the emergence of novel influenza strains that are capable of escaping from population immunity. The timely determination of antigenic variants is critical to vaccine design. Empirical experimental methods like hemagglutination inhibition (HI) assays are time-consuming and labor-intensive, requiring live viruses. Recently, many computational models have been developed to predict the antigenic variants without considerations of explicitly modeling the interdependencies between the channels of feature maps. Moreover, the influenza sequences consisting of similar distribution of residues will have high degrees of similarity and will affect the prediction outcome. Consequently, it is challenging but vital to determine the importance of different residue sites and enhance the predictive performance of influenza antigenicity. We have proposed a 2D convolutional neural network (CNN) model to infer influenza antigenic variants (IAV-CNN). Specifically, we apply a new distributed representation of amino acids, named ProtVec that can be applied to a variety of downstream proteomic machine learning tasks. After splittings and embeddings of influenza strains, a 2D squeeze-and-excitation CNN architecture is constructed that enables networks to focus on informative residue features by fusing both spatial and channel-wise information with local receptive fields at each layer. Experimental results on three influenza datasets show IAV-CNN achieves state-of-the-art performance combining the new distributed representation with our proposed architecture. It outperforms both traditional machine algorithms with the same feature representations and the majority of existing models in the independent test data. Therefore we believe that our model can be served as a reliable and robust tool for the prediction of antigenic variants.
Rui Yin 0002, Nyi Nyi Thwin, Pei Zhuang, Zhuoyi Lin, Chee Keong Kwoh 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 VirPreNet: a weighted ensemble convolutional neural network for the virulence prediction of influenza A virus using all eight segments
abstract
MOTIVATION: Influenza viruses are persistently threatening public health, causing annual epidemics and sporadic pandemics. The evolution of influenza viruses remains to be the main obstacle in the effectiveness of antiviral treatments due to rapid mutations. Previous work has been investigated to reveal the determinants of virulence of the influenza A virus. To further facilitate flu surveillance, explicit detection of influenza virulence is crucial to protect public health from potential future pandemics. RESULTS: In this article, we propose a weighted ensemble convolutional neural network (CNN) for the virulence prediction of influenza A viruses named VirPreNet that uses all eight segments. Firstly, mouse lethal dose 50 is exerted to label the virulence of infections into two classes, namely avirulent and virulent. A numerical representation of amino acids named ProtVec is applied to the eight-segments in a distributed manner to encode the biological sequences. After splittings and embeddings of influenza strains, the ensemble CNN is constructed as the base model on the influenza dataset of each segment, which serves as the VirPreNet's main part. Followed by a linear layer, the initial predictive outcomes are integrated and assigned with different weights for the final prediction. The experimental results on the collected influenza dataset indicate that VirPreNet achieves state-of-the-art performance combining ProtVec with our proposed architecture. It outperforms baseline methods on the independent testing data. Moreover, our proposed model reveals the importance of PB2 and HA segments on the virulence prediction. We believe that our model may provide new insights into the investigation of influenza virulence. AVAILABILITY AND IMPLEMENTATION: Codes and data to generate the VirPreNet are publicly available at https://github.com/Rayin-saber/VirPreNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rui Yin 0002, Zihan Luo 0001, Pei Zhuang, Zhuoyi Lin, Chee Keong Kwoh 0001
Bioinform.1
2021 GLIMG: Global and local item graphs for top-N recommender systems
Zhuoyi Lin, Lei Feng 0006, Rui Yin 0002, Chee Keong Kwoh 0001
Inf. Sci.3
2020 Tempel: time-series mutation prediction of influenza A viruses via attention-based recurrent neural networks
abstract
MOTIVATION: Influenza viruses are persistently threatening public health, causing annual epidemics and sporadic pandemics. The evolution of influenza viruses remains to be the main obstacle in the effectiveness of antiviral treatments due to rapid mutations. The goal of this work is to predict whether mutations are likely to occur in the next flu season using historical glycoprotein hemagglutinin sequence data. One of the major challenges is to model the temporality and dimensionality of sequential influenza strains and to interpret the prediction results. RESULTS: In this article, we propose an efficient and robust time-series mutation prediction model (Tempel) for the mutation prediction of influenza A viruses. We first construct the sequential training samples with splittings and embeddings. By employing recurrent neural networks with attention mechanisms, Tempel is capable of considering the historical residue information. Attention mechanisms are being increasingly used to improve the performance of mutation prediction by selectively focusing on the parts of the residues. A framework is established based on Tempel that enables us to predict the mutations at any specific residue site. Experimental results on three influenza datasets show that Tempel can significantly enhance the predictive performance compared with widely used approaches and provide novel insights into the dynamics of viral mutation and evolution. AVAILABILITY AND IMPLEMENTATION: The datasets, source code and supplementary documents are available at: https://drive.google.com/drive/folders/15WULR5__6k47iRotRPl3H7ghi3RpeNXH. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rui Yin 0002, Emil Luusua, Jan Dabrowski, Yu Zhang 0084, Chee Keong Kwoh 0001
Bioinform.1
2019 Ensemble Prediction of Synergistic Drug Combinations Incorporating Biological, Chemical, Pharmacological, and Network Knowledge
abstract
Combinatorial therapy may reduce drug side effects and improve drug efficacy, making combination therapy a promising strategy to treat complex diseases. However, in the existing computational methods, the natural properties and network knowledge of drugs have not been adequately and simultaneously considered, making it difficult to identify effective drug combinations. Computational methods that incorporate multiple sources of information (biological, chemical, pharmacological, and network knowledge) offer more opportunities to screen synergistic drug combinations. Therefore, we developed a novel Ensemble Prediction framework of Synergistic Drug Combinations (EPSDC) to accurately and efficiently predict drug combinations by integrating information from multiple-sources. EPSDC constructs feature vector of drug pair by concatenating different types of drug similarities, and then uses these groups in a feature-based base predictor. Next, transductive learning is applied on heterogeneous drug-target networks to achieve a network-based score for the drug pair. Finally, two types of ensemble rules are introduced to combine the feature-based score and the network-based score, and then potential drug combinations are prioritized. To demonstrate the effect of the ensemble rule, comprehensive experiments were conducted to compare single models and ensemble models. The experimental results indicated that our method outperformed the state-of-the-art method in five-fold cross validation and de novo prediction tests on the two benchmark datasets. We further analyzed the effect of maximum length of the meta-path and the impacts of different types of features. Moreover, the practical usefulness of our method was confirmed in the predicted novel drug combinations. The source code of EPSDC is available at https://github.com/KDDing/EPSDC.
Pingjian Ding, Rui Yin 0002, Jiawei Luo 0001, Chee Keong Kwoh 0001
IEEE J. Biomed. Health Informatics2