Chee Keong Kwoh 0001

dblp:32/228 · also Chee-Keong Kwoh 0001, Kwoh Chee Keong 0001 · DBLP profile ↗
← Back
103ranked-venue papers
4as first author
39since 2021 · last 2026
0000-0002-8547-6387ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 70 · 25 since 2021Artificial intelligence and machine learning · 20 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 FOCUS: Frequency-Optimized Conditioning of diffUSion models for mitigating catastrophic forgetting during test-time adaptation
Gabriel Tjio, Jie Zhang 0002, Xulei Yang, Nhat Chung, Xiaofeng Cao 0002, Ivor W. Tsang, Chee Keong Kwoh 0001, Qing Guo 0005
Mach. Vis. Appl.8
2026 Target-Specific Adaptation and Consistent Degradation Alignment for Cross-Domain Remaining Useful Life Prediction
abstract
Accurate prediction of the Remaining Useful Life (RUL) in machinery can significantly diminish maintenance costs, enhance equipment up-time, and mitigate adverse outcomes. Data-driven RUL prediction techniques have demonstrated commendable performance. However, their efficacy often relies on the assumption that training and testing data are drawn from the same distribution or domain, which does not hold in real industrial settings. To mitigate this domain discrepancy issue, prior adversarial domain adaptation methods focused on deriving domain-invariant features. Nevertheless, they overlook target-specific information and inconsistency characteristics pertinent to the degradation stages, resulting in suboptimal performance. To tackle these issues, we propose a novel domain adaptation approach for cross-domain RUL prediction named TACDA. Specifically, we propose a target domain reconstruction strategy within the adversarial adaptation process, thereby retaining target-specific information while learning domain-invariant features. Furthermore, we develop a novel clustering and pairing strategy for consistent alignment between similar degradation stages. Through extensive experiments, our results demonstrate the remarkable performance of our proposed TACDA method, surpassing state-of-the-art approaches with regard to two different evaluation metrics. Our code is available at https://github.com/keyplay/TACDA.
Yubo Hou, Mohamed Ragab 0002, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Zhenghua Chen
IEEE Trans Autom. Sci. Eng.4
2026 Image-Enhanced Multi-Modal Contrastive Transformer for Subcellular Spatial Transcriptomics
abstract
Recent advances in spatial molecular imaging technologies have enabled gene expression profiling alongside high-resolution imaging, providing unprecedented opportunities to resolve molecular heterogeneity at subcellular resolution. However, these technologies fail to fully capture cellular characteristics due to the limited number of genes they can detect, which hinder downstream analysis. Spatial imaging data provide high-resolution and fine-grained morphology information, developing computational methods that effectively integrate image features with transcriptomic profiles is crucial for enabling comprehensive subcellular data analysis. In this study, we present SIMMT, an image-enhanced multi-modal contrastive transformer framework for identifying spatial domains and enhancing subcellular data. In the framework, we design a dual transformer architecture to learn multi-modal representations for cells by modeling transcriptomics and morphological images respectively. To fully capture modality interactions within spatial contexts, we introduce a contrastive learning module that enhances cell representation by aligning tissue morphology and gene expression at the cell level. We tested SIMMT on subcellular spatial transcriptomics datasets from human lung cancer tissue, mouse brain tissue, human colorectal cancer tissue, and human ovarian cancer tissue. The results demonstrated that SIMMT consistently outperformed state-of-the-art methods in spatial clustering and gene expression pattern analysis. Our method also effectively demonstrated its ability to identify tumor spatial heterogeneity and uncover potential gene biomarkers in the human bronchiolar adenoma (BA) dataset.
Wanwan Shi, Ying Liu 0027, Qiu Xiao, Yuting Bai, Xinling Zeng, Chee Keong Kwoh 0001, Jiawei Luo 0001
IEEE J. Biomed. Health Informatics7
2025 DeepPhosPPI: a deep learning framework with attention-CNN and transformer for predicting phosphorylation effects on protein-protein interactions
abstract
Protein phosphorylation regulates protein function and cellular signaling pathways, and is strongly associated with diseases, including neurodegenerative disorders and cancer. Phosphorylation plays a critical role in regulating protein activity and cellular signaling by modulating protein-protein interactions (PPIs). It alters binding affinities and interaction networks, thereby influencing biological processes and maintaining cellular homeostasis. Experimental validation of these effects is labor-intensive and expensive, highlighting the need for efficient computational approaches. We propose DeepPhosPPI, the first sequence-based deep learning framework for phosphorylation effects on PPIs prediction, which employs the pre-trained protein language model for feature embedding, with ProtBERT and ESM-2 as alternative backbone encoders. By combining attention-based convolutional neural network and Transformer models, DeepPhosPPI accurately predicts phosphorylation effects. The experimental results show that DeepPhosPPI consistently outperforms state-of-the-art methods in multiple tasks, including functional sites identification and regulatory effect classification.
Yinyin Gong, Rui Li 0019, Yan Liu 0032, Jilong Wang 0002, Danny Ziyi Chen, Chee Keong Kwoh 0001
Briefings Bioinform.6
2025 STCGAN: a novel cycle-consistent generative adversarial network for spatial transcriptomics cellular deconvolution
abstract
MOTIVATION: Spatial transcriptomics (ST) technologies have revolutionized our ability to map gene expression patterns within native tissue context, providing unprecedented insights into tissue architecture and cellular heterogeneity. However, accurately deconvolving cell-type compositions from ST spots remains challenging due to the sparse and averaged nature of ST data, which is essential for accurately depicting tissue architecture. While numerous computational methods have been developed for cell-type deconvolution and spatial distribution reconstruction, most fail to capture tissue complexity at the single-cell level, thereby limiting their applicability in practical scenarios. RESULTS: To this end, we propose a novel cycle-consistent generative adversarial network named STCGAN for cellular deconvolution in spatial transcriptomic. STCGAN first employs a cycle-consistent generative adversarial network (CGAN) to pre-train on ST data, ensuring that both the mapping from ST data to latent space and its reverse mapping are consistent, capturing complex spatial gene expression patterns and learning robust latent representations. Based on the learned representation, STCGAN then optimizes a trainable cell-to-spot mapping matrix to integrate scRNA-seq data with ST data, accurately estimating cellular composition within each capture spot and effectively reconstructing the spatial distribution of cells across the tissue. To further enhance deconvolution accuracy, we incorporate spatial-aware regularization that ensures accurate cellular distribution reconstruction within the spatial context. Benchmarking against seven state-of-the-art methods on five simulated and real datasets from various tissues, STCGAN consistently delivers superior cell-type deconvolution performance. AVAILABILITY: The code of STCGAN can be downloaded from https://github.com/cs-wangbo/STCGAN and all the mentioned datasets are available on Zenodo at https://zenodo.org/doi/10.5281/zenodo.10799113.
Yahui Long, Yuting Bai, Jiawei Luo 0001, Chee Keong Kwoh 0001
Briefings Bioinform.5
2024 SMMGCL: a novel multi-level graph contrastive learning framework for integrating spatial multi-omics data
abstract
Recent advances in spatial omics technologies have allowed various omics data to be obtained from a single tissue section. To fully explore the relationships among these different types of omics data, it is urgent to develop more effective methods for spatial multi-omics data integration. In this work, we propose a novel Multi-level Graph Contrastive Learning framework, named SMMGCL, to simultaneously mine complementary information at both spot and graph levels for integrating Spatial Multi-omics data. Specifically, to adaptively fuse multi-omics modalities, we first design a multi-modality autoencoder that integrates spatial locations with spot omic expressions to extract modality-specific embeddings. These embeddings are then fused into a consensus representation using an attention mechanism to capture spot-level cross-omics representations. Next, to explore the complex inter-omic structural information, we connect corresponding spots across different omics adjacency graphs into a heterogeneous graph. We then employ a graph convolutional network (GCN) to extract spatial correlations across the omics, learning a graph-level cross-omics global representation. Finally, SMMGCL aligns feature similarity matrixes between spot-level and graph-level representations with their pseudo-label similarity matrix, ensuring multi-level clustering consistency and leading to more accurate spatial multi-omics integration. Experimental results on simulated and real datasets from across tissues show that SMMGCL consistently outperforms other state-of-the-art methods in spatial multi-omics integration performance. The code for SMMGCL is available for download from the GitHub repository at https://github.com/cs-wangbo/SMMGCL.
Wei Liu 0296, Jiawei Luo 0001, Xiangtao Chen, Chee Keong Kwoh 0001
BIBM5
2024 RmsdXNA: RMSD prediction of nucleic acid-ligand docking poses using machine-learning method
abstract
Small molecule drugs can be used to target nucleic acids (NA) to regulate biological processes. Computational modeling methods, such as molecular docking or scoring functions, are commonly employed to facilitate drug design. However, the accuracy of the scoring function in predicting the closest-to-native docking pose is often suboptimal. To overcome this problem, a machine learning model, RmsdXNA, was developed to predict the root-mean-square-deviation (RMSD) of ligand docking poses in NA complexes. The versatility of RmsdXNA has been demonstrated by its successful application to various complexes involving different types of NA receptors and ligands, including metal complexes and short peptides. The predicted RMSD by RmsdXNA was strongly correlated with the actual RMSD of the docked poses. RmsdXNA also outperformed the rDock scoring function in ranking and identifying closest-to-native docking poses across different structural groups and on the testing dataset. Using experimental validated results conducted on polyadenylated nuclear element for nuclear expression triplex, RmsdXNA demonstrated better screening power for the RNA-small molecule complex compared to rDock. Molecular dynamics simulations were subsequently employed to validate the binding of top-scoring ligand candidates selected by RmsdXNA and rDock on MALAT1. The results showed that RmsdXNA has a higher success rate in identifying promising ligands that can bind well to the receptor. The development of an accurate docking score for a NA-ligand complex can aid in drug discovery and development advancements. The code to use RmsdXNA is available at the GitHub repository https://github.com/laiheng001/RmsdXNA.
Lai Heng Tan, Chee Keong Kwoh 0001, Yuguang Mu
Briefings Bioinform.2
2024 Systematic benchmarking of deep-learning methods for tertiary RNA structure prediction
abstract
The 3D structure of RNA critically influences its functionality, and understanding this structure is vital for deciphering RNA biology. Experimental methods for determining RNA structures are labour-intensive, expensive, and time-consuming. Computational approaches have emerged as valuable tools, leveraging physics-based-principles and machine learning to predict RNA structures rapidly. Despite advancements, the accuracy of computational methods remains modest, especially when compared to protein structure prediction. Deep learning methods, while successful in protein structure prediction, have shown some promise for RNA structure prediction as well, but face unique challenges. This study systematically benchmarks state-of-the-art deep learning methods for RNA structure prediction across diverse datasets. Our aim is to identify factors influencing performance variation, such as RNA family diversity, sequence length, RNA type, multiple sequence alignment (MSA) quality, and deep learning model architecture. We show that generally ML-based methods perform much better than non-ML methods on most RNA targets, although the performance difference isn't substantial when working with unseen novel or synthetic RNAs. The quality of the MSA and secondary structure prediction both play an important role and most methods aren't able to predict non-Watson-Crick pairs in the RNAs. Overall among the automated 3D RNA structure prediction methods, DeepFoldRNA has the best prediction results followed by DRFold as the second best method. Finally, we also suggest possible mitigations to improve the quality of the prediction for future method development.
Akash Bahai, Chee Keong Kwoh 0001, Yuguang Mu
PLoS Comput. Biol.2
2024 Self-Supervised Autoregressive Domain Adaptation for Time Series Data
abstract
Unsupervised domain adaptation (UDA) has successfully addressed the domain shift problem for visual applications. Yet, these approaches may have limited performance for time series data due to the following reasons. First, they mainly rely on the large-scale dataset (i.e., ImageNet) for source pretraining, which is not applicable for time series data. Second, they ignore the temporal dimension on the feature space of the source and target domains during the domain alignment step. Finally, most of the prior UDA methods can only align the global features without considering the fine-grained class distribution of the target domain. To address these limitations, we propose a SeLf-supervised AutoRegressive Domain Adaptation (SLARDA) framework. In particular, we first design a self-supervised (SL) learning module that uses forecasting as an auxiliary task to improve the transferability of source features. Second, we propose a novel autoregressive domain adaptation technique that incorporates temporal dependence of both source and target features during domain alignment. Finally, we develop an ensemble teacher model to align class-wise distribution in the target domain via a confident pseudo labeling approach. Extensive experiments have been conducted on three real-world time series applications with 30 cross-domain scenarios. The results demonstrate that our proposed SLARDA method significantly outperforms the state-of-the-art approaches for time series domain adaptation. Our source code is available at: https://github.com/mohamedr002/SLARDA.
Mohamed Ragab 0002, Emadeldeen Eldele, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 PESI: Paratope-Epitope Set Interaction for SARS-CoV-2 Neutralization Prediction
abstract
Prediction of neutralization antibodies is important for the development of effective vaccines and antibody-based therapeutics. Traditional methods rely on features based on first principles derived from the binding interface. However, they are burdened by arduous data preprocessing from a limited quantity of protein structures. In comparison, deep learning allows automatic substructure characterization and representation without hand-crafted feature engineering. In particular, large language models (LLMs) based method predicts neutralization using Fv sequences of antibody and antigen. Despite LLM’s success, incorporating full-length Fv sequences suffers from: 1) inaccurate sequence-level labels in existing datasets, 2) inefficient modeling due to noisy non-contributing motifs, and 3) ignorance of non-bonded interactions that play a key role in facilitating epitope-paratope pairing. In this paper, we propose a novel approach that incorporates only the paratope and epitope for antibody-antigen neutralization prediction while adopting a novel set modeling that regards the paratope and epitope as bags of residues. Specifically, we hand-crafted a dataset containing neutralizing paratope-epitope pairs where epitopes are potentially generalizable to future unseen variants of SARS-CoV-2. Training on such a dataset enables deep learning models to predict neutralizing antibodies for prospective mutated variants of SARS-CoV-2, meanwhile addressing the problem of inaccurate sequence-level labels. A higher modeling efficiency is also achieved by disregarding non-contributing motifs. Furthermore, we also propose paratope-epitope set interaction (PESI), a set modeling model inspired by first principles that learns intra-inter non-covalent interactions through a global attention mechanism. To validate PESI, we perform a 10-fold cross-validation on our dataset. Experimental results show that PESI achieves a more balanced overall performance and a significant improvement on MCC as compared to existing architectures.
Zhang Wan, Zhuoyi Lin, Shamima Rashid, Shaun Yue-Hao Ng, Rui Yin 0002, J. Senthilnath 0001, Chee Keong Kwoh 0001
BIBM7
2023 Directed collaboration patterns in funded teams: A perspective of knowledge flow
Bentao Zou, Yuefen Wang, Chee Keong Kwoh 0001, Yonghua Cen
Inf. Process. Manag.3
2023 ViPal: A framework for virulence prediction of influenza viruses with prior viral knowledge using genomic sequences
Rui Yin 0002, Zihan Luo 0001, Pei Zhuang, Min Zeng 0004, Min Li 0007, Zhuoyi Lin, Chee Keong Kwoh 0001
J. Biomed. Informatics7
2023 Self-Supervised Contrastive Representation Learning for Semi-Supervised Time-Series Classification
abstract
Learning time-series representations when only unlabeled data or few labeled samples are available can be a challenging task. Recently, contrastive self-supervised learning has shown great improvement in extracting useful representations from unlabeled data via contrasting different augmented views of data. In this work, we propose a novel Time-Series representation learning framework via Temporal and Contextual Contrasting (TS-TCC) that learns representations from unlabeled data with contrastive learning. Specifically, we propose time-series-specific weak and strong augmentations and use their views to learn robust temporal relations in the proposed temporal contrasting module, besides learning discriminative representations by our proposed contextual contrasting module. Additionally, we conduct a systematic study of time-series data augmentation selection, which is a key part of contrastive learning. We also extend TS-TCC to the semi-supervised learning settings and propose a Class-Aware TS-TCC (CA-TCC) that benefits from the available few labeled data to further improve representations learned by TS-TCC. Specifically, we leverage the robust pseudo labels produced by TS-TCC to realize a class-aware contrastive loss. Extensive experiments show that the linear evaluation of the features learned by our proposed framework performs comparably with the fully supervised training. Additionally, our framework shows high efficiency in few labeled data and transfer learning scenarios.
Emadeldeen Eldele, Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Cuntai Guan
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Empirical Study of Protein Feature Representation on Deep Belief Networks Trained With Small Data for Secondary Structure Prediction
abstract
Protein secondary structure (SS) prediction is a classic problem of computational biology and is widely used in structural characterization and to infer homology. While most SS predictors have been trained on thousands of sequences, a previous approach had developed a compact model of training proteins that used aC-Alpha, C-BetaSide Chain (CABS)-algorithm derived energy based feature representation. Here, the previous approach is extended to Deep Belief Networks (DBN). Deep learning methods are notorious for requiring large datasets and there is a wide consensus that training deep models from scratch on small datasets, works poorly. By contrast, we demonstrate a simple DBN architecture containing a single hidden layer, trained only on the CB513 dataset. Testing on an independent set of G Switch proteins improved the Q$_{3}$score of the previous compact model by almost 3%. The findings are further confirmed by comparison to several deep learning models which are trained on thousands of proteins. Finally, the DBN performance is also compared withPositionSpecificScoringMatrix (PSSM)-profile based feature representation. The importance of (i) structural information in protein feature representation and (ii) complementary small dataset learning approaches for detection of structural fold switching are demonstrated.
Shamima Rashid, Suresh Sundaram 0002, Chee Keong Kwoh 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 COMET: Convolutional Dimension Interaction for Collaborative Filtering
abstract
Representation learning-based recommendation models play a dominant role among recommendation techniques. However, most of the existing methods assume both historical interactions and embedding dimensions are independent of each other, and thus regrettably ignore the high-order interaction information among historical interactions and embedding dimensions. In this article, we propose a novel representation learning-based model called COMET (COnvolutional diMEnsion inTeraction), which simultaneously models the high-order interaction patterns among historical interactions and embedding dimensions. To be specific, COMET stacks the embeddings of historical interactions horizontally at first, which results in two “embedding maps”. In this way, internal interactions and dimensional interactions can be exploited by convolutional neural networks (CNN) with kernels of different sizes simultaneously. A fully connected multi-layer perceptron (MLP) is then applied to obtain two interaction vectors. Lastly, the representations of users and items are enriched by the learnt interaction vectors, which can further be used to produce the final prediction. Extensive experiments and ablation studies on various public implicit feedback datasets clearly demonstrate the effectiveness and rationality of our proposed method.
Zhuoyi Lin, Lei Feng 0006, Xingzhi Guo, Yu Zhang 0084, Rui Yin 0002, Chee Keong Kwoh 0001
ACM Trans. Intell. Syst. Technol.6
2023 ADATIME: A Benchmarking Suite for Domain Adaptation on Time Series Data
abstract
Unsupervised domain adaptation methods aim at generalizing well on unlabeled test data that may have a different (shifted) distribution from the training data. Such methods are typically developed on image data, and their application to time series data is less explored. Existing works on time series domain adaptation suffer from inconsistencies in evaluation schemes, datasets, and backbone neural network architectures. Moreover, labeled target data are often used for model selection, which violates the fundamental assumption of unsupervised domain adaptation. To address these issues, we develop a benchmarking evaluation suite ( AdaTime ) to systematically and fairly evaluate different domain adaptation methods on time series data. Specifically, we standardize the backbone neural network architectures and benchmarking datasets, while also exploring more realistic model selection approaches that can work with no labeled data or just a few labeled samples. Our evaluation includes adapting state-of-the-art visual domain adaptation methods to time series data as well as the recent methods specifically developed for time series data. We conduct extensive experiments to evaluate 11 state-of-the-art methods on five representative datasets spanning 50 cross-domain scenarios. Our results suggest that with careful selection of hyper-parameters, visual domain adaptation methods are competitive with methods proposed for time series domain adaptation. In addition, we find that hyper-parameters could be selected based on realistic model selection approaches. Our work unveils practical insights for applying domain adaptation methods on time series data and builds a solid foundation for future works in the field. The code is available at github.com/emadeldeen24/AdaTime .
Mohamed Ragab 0002, Emadeldeen Eldele, Wee Ling Tan, Chuan-Sheng Foo, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001
ACM Trans. Knowl. Discov. Data7
2023 Attention Over Self-Attention: Intention-Aware Re-Ranking With Dynamic Transformer Encoders for Recommendation
abstract
Re-ranking models refine item recommendation lists generated by the prior global ranking model, which have demonstrated their effectiveness in improving the recommendation quality. However, most existing re-ranking solutions only learn from implicit feedback with a shared prediction model, which regrettably ignore inter-item relationships under diverse user intentions. In this paper, we propose a novel Intention-aware Re-ranking Model with Dynamic TransformerEncoder (RAISE), aiming to perform user-specific prediction for each individual user based on her intentions. Specifically, we first propose to mine latent user intentions from text reviews with an intention discovering module (IDM). By differentiating the importance of review information with a co-attention network, the latent user intention can be explicitly modeled for each user-item pair. We then introduce a dynamic transformer encoder (DTE) to capture user-specific inter-item relationships among item candidates by seamlessly accommodating the learned latent user intentions via IDM. As such, one can not only achieve more personalized recommendations but also obtain corresponding explanations by constructing RAISE upon existing recommendation engines. Empirical study on four public datasets shows the superiority of our proposed RAISE, with up to 13.95%, 9.60%, and 13.03% relative improvements evaluated by Precision@5, MAP@5, and NDCG@5 respectively.
Zhuoyi Lin, Sheng Zang, Zhu Sun 0001, J. Senthilnath 0001, Chee Keong Kwoh 0001
IEEE Trans. Knowl. Data Eng.7
2022 A heterogeneous graph cross-omics attention model for single-cell representation learning
abstract
Single-cell multi-omics sequencing technologies allow simultaneous measurement of transcriptome and epigenome profiles in the same cell, providing unprecedented opportunities to dissect cell heterogeneity. Despite great efforts, conjoint analysis of single-cell multi-omics data still suffers from sparsity, high dimensionality and binary. In this study, we present a heterogeneous graph cross-omics attention model (scHGA), a computational tool based on a heterogeneous graph neural network combining two attention mechanisms to jointly analyze single-cell multi-omics data based on different protocols data, including SNARE-seq, scMT-seq and sci-CAR. To avoid the cell heterogeneity of single-omics data, scHGA automatically learns a cell association graph to capture neighbor information. The latent representation of aggregated cells generated by hierarchical attention can fuse knowledge across different omics to dissect cellular heterogeneity, providing a better scheme to characterize the features of cells. scHGA is an effective exploration of graph neural networks in single-cell multi-omics analysis, providing new insights into the understanding of single-cell sequencing data.
Yue Liu 0041, Shulin Wang, Wei Zhang 0089, Xiangxiang Zeng, Chee Keong Kwoh 0001
BIBM6
2022 Rotating Machinery Fault Diagnosis Based on Multi-sensor Information Fusion Using Graph Attention Network
abstract
Multi-sensor information acquisition system can reflect the operation status of machinery more comprehensively and reliably, but also demands higher requirements on data analysis algorithms. Unlike previous deep learning models, the emerging Graph Neural Network (GNN) has a remarkable performance in mining graph structure and patterns, effectively integrating multiple node relationships and features. This paper presents a fault diagnosis algorithm based on multi-sensor information fusion using the modified Graph Attention Network-GATv2. Firstly, the dependencies between multi-sensor signals are explicitly extracted by the Grow-Shrink (GS) algorithm, where the topology of the constructed graph can characterize different failure states of the equipment. During the aggregation process, the attention mechanism in the GATv2 assigns higher weights to informative nodes for the effective fusion of multi-sensor information. Experiments show that the proposed diagnosis framework can yield more expressive multi-sensor representations, and the diagnostic accuracy is improved significantly compared to the single-sensor graph.
Chenyang Li 0005, Chee Keong Kwoh 0001, Xiaoli Li 0001, Lingfei Mo, Ruqiang Yan 0001
ICARCV2
2022 Jupytope: computational extraction of structural properties of viral epitopes
abstract
Epitope residues located on viral surface proteins are of immense interest in immunology and related applications such as vaccine development, disease diagnosis and drug design. Most tools rely on sequence-based statistical comparisons, such as information entropy of residue positions in aligned columns to infer location and properties of epitope sites. To facilitate cross-structural comparisons of epitopes on viral surface proteins, a python-based extraction tool implemented with Jupyter notebook is presented (Jupytope). Given a viral antigen structure of interest, a list of known epitope sites and a reference structure, the corresponding epitope structural properties can quickly be obtained. The tool integrates biopython modules for commonly used software such as NACCESS, DSSP as well as residue depth and outputs a list of structure-derived properties such as dihedral angles, solvent accessibility, residue depth and secondary structure that can be saved in several convenient data formats. To ensure correct spatial alignment, Jupytope takes a list of given epitope sites and their corresponding reference structure and aligns them before extracting the desired properties. Examples are demonstrated for epitopes of Influenza and severe acute respiratory syndrome coronavirus 2 (SARS-CoV2) viral strains. The extracted properties assist detection of two Influenza subtypes and show potential in distinguishing between four major clades of SARS-CoV2, as compared with randomized labels. The tool will facilitate analytical and predictive works on viral epitopes through the extracted structural information. Jupytope and extracted datasets are available at https://github.com/shamimarashid/Jupytope.
Shamima Rashid, Teng Ann Ng, Chee Keong Kwoh 0001
Briefings Bioinform.3
2022 Graph representation learning in bioinformatics: trends, methods and applications
abstract
Graph is a natural data structure for describing complex systems, which contains a set of objects and relationships. Ubiquitous real-life biomedical problems can be modeled as graph analytics tasks. Machine learning, especially deep learning, succeeds in vast bioinformatics scenarios with data represented in Euclidean domain. However, rich relational information between biological elements is retained in the non-Euclidean biomedical graphs, which is not learning friendly to classic machine learning methods. Graph representation learning aims to embed graph into a low-dimensional space while preserving graph topology and node properties. It bridges biomedical graphs and modern machine learning methods and has recently raised widespread interest in both machine learning and bioinformatics communities. In this work, we summarize the advances of graph representation learning and its representative applications in bioinformatics. To provide a comprehensive and structured analysis and perspective, we first categorize and analyze both graph embedding methods (homogeneous graph embedding, heterogeneous graph embedding, attribute graph embedding) and graph neural networks. Furthermore, we summarize their representative applications from molecular level to genomics, pharmaceutical and healthcare systems level. Moreover, we provide open resource platforms and libraries for implementing these graph representation learning methods and discuss the challenges and opportunities of graph representation learning in bioinformatics. This work provides a comprehensive survey of emerging graph representation learning algorithms and their applications in bioinformatics. It is anticipated that it could bring valuable insights for researchers to contribute their knowledge to graph representation learning and future-oriented bioinformatics studies.
Zhu-Hong You, De-Shuang Huang, Chee Keong Kwoh 0001
Briefings Bioinform.4
2022 A framework for predicting variable-length epitopes of human-adapted viruses using machine learning methods
abstract
The coronavirus disease 2019 pandemic has alerted people of the threat caused by viruses. Vaccine is the most effective way to prevent the disease from spreading. The interaction between antibodies and antigens will clear the infectious organisms from the host. Identifying B-cell epitopes is critical in vaccine design, development of disease diagnostics and antibody production. However, traditional experimental methods to determine epitopes are time-consuming and expensive, and the predictive performance using the existing in silico methods is not satisfactory. This paper develops a general framework to predict variable-length linear B-cell epitopes specific for human-adapted viruses with machine learning approaches based on Protvec representation of peptides and physicochemical properties of amino acids. QR decomposition is incorporated during the embedding process that enables our models to handle variable-length sequences. Experimental results on large immune epitope datasets validate that our proposed model's performance is superior to the state-of-the-art methods in terms of AUROC (0.827) and AUPR (0.831) on the testing set. Moreover, sequence analysis also provides the results of the viral category for the corresponding predicted epitopes with high precision. Therefore, this framework is shown to reliably identify linear B-cell epitopes of human-adapted viruses given protein sequences and could provide assistance for potential future pandemics and epidemics.
Rui Yin 0002, Xianghe Zhu, Min Zeng 0004, Min Li 0007, Chee Keong Kwoh 0001
Briefings Bioinform.6
2022 Pre-training graph neural networks for link prediction in biomedical networks
abstract
MOTIVATION: Graphs or networks are widely utilized to model the interactions between different entities (e.g. proteins, drugs, etc.) for biomedical applications. Predicting potential interactions/links in biomedical networks is important for understanding the pathological mechanisms of various complex human diseases, as well as screening compound targets for drug discovery. Graph neural networks (GNNs) have been utilized for link prediction in various biomedical networks, which rely on the node features extracted from different data sources, e.g. sequence, structure and network data. However, it is challenging to effectively integrate these data sources and automatically extract features for different link prediction tasks. RESULTS: In this article, we propose a novel Pre-Training Graph Neural Networks-based framework named PT-GNN to integrate different data sources for link prediction in biomedical networks. First, we design expressive deep learning methods [e.g. convolutional neural network and graph convolutional network (GCN)] to learn features for individual nodes from sequence and structure data. Second, we further propose a GCN-based encoder to effectively refine the node features by modelling the dependencies among nodes in the network. Third, the node features are pre-trained based on graph reconstruction tasks. The pre-trained features can be used for model initialization in downstream tasks. Extensive experiments have been conducted on two critical link prediction tasks, i.e. synthetic lethality (SL) prediction and drug-target interaction (DTI) prediction. Experimental results demonstrate PT-GNN outperforms the state-of-the-art methods for SL prediction and DTI prediction. In addition, the pre-trained features benefit improving the performance and reduce the training time of existing models. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at: https://github.com/longyahui/PT-GNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yahui Long, Min Wu 0008, Yong Liu 0020, Yuan Fang 0001, Chee Keong Kwoh 0001, Jinmiao Chen, Jiawei Luo 0001, Xiaoli Li 0001
Bioinform.5
2022 An Efficient Multiresolution Clustering for Motif Discovery in Complex Networks
abstract
Motif discovery and network clustering in complex networks have received a lot of attention in recent years, also they are still challenging tasks in bioinformatics, big data analytics and data mining applications. Motif discovery in big data networks has a lot of important applications in different domains such as engineering, bioinformatics, cheminformatics, genomics, sociology and ecology for revealing hidden frequent structures, functional building blocks, or knowledge discovery. In this paper, a motif localization method based on a novel clustering algorithm in complex networks is presented. In our method, for each complex network, a novel structure so-called Augmented Multiresolution Network (AMN) is generated, then it is adaptively partitioned into several clusters and their corresponding subnets. Then top ranked subnets are chosen to discover network motifs. We show that the proposed method provides an efficient solution for clustering and motif discovery; It speeds up current motif discovery algorithms by pruning non-promising regions of complex networks. Experimental results show our algorithm efficiently deals with complex networks representing large datasets with high-dimensionality such as big scientific data. Our method also provides motivations for future studies in big data and complex networks.
Mahdi Pursalim, Chee Keong Kwoh 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 IAV-CNN: A 2D Convolutional Neural Network Model to Predict Antigenic Variants of Influenza A Virus
abstract
The rapid evolution of influenza viruses constantly leads to the emergence of novel influenza strains that are capable of escaping from population immunity. The timely determination of antigenic variants is critical to vaccine design. Empirical experimental methods like hemagglutination inhibition (HI) assays are time-consuming and labor-intensive, requiring live viruses. Recently, many computational models have been developed to predict the antigenic variants without considerations of explicitly modeling the interdependencies between the channels of feature maps. Moreover, the influenza sequences consisting of similar distribution of residues will have high degrees of similarity and will affect the prediction outcome. Consequently, it is challenging but vital to determine the importance of different residue sites and enhance the predictive performance of influenza antigenicity. We have proposed a 2D convolutional neural network (CNN) model to infer influenza antigenic variants (IAV-CNN). Specifically, we apply a new distributed representation of amino acids, named ProtVec that can be applied to a variety of downstream proteomic machine learning tasks. After splittings and embeddings of influenza strains, a 2D squeeze-and-excitation CNN architecture is constructed that enables networks to focus on informative residue features by fusing both spatial and channel-wise information with local receptive fields at each layer. Experimental results on three influenza datasets show IAV-CNN achieves state-of-the-art performance combining the new distributed representation with our proposed architecture. It outperforms both traditional machine algorithms with the same feature representations and the majority of existing models in the independent test data. Therefore we believe that our model can be served as a reliable and robust tool for the prediction of antigenic variants.
Rui Yin 0002, Nyi Nyi Thwin, Pei Zhuang, Zhuoyi Lin, Chee Keong Kwoh 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Toward Multidiversified Ensemble Clustering of High-Dimensional Data: From Subspaces to Metrics and Beyond
abstract
The rapid emergence of high-dimensional data in various areas has brought new challenges to current ensemble clustering research. To deal with the curse of dimensionality, recently considerable efforts in ensemble clustering have been made by means of different subspace-based techniques. However, besides the emphasis on subspaces, rather limited attention has been paid to the potential diversity in similarity/dissimilarity metrics. It remains a surprisingly open problem in ensemble clustering how to create and aggregate a large population of diversified metrics, and furthermore, how to jointly investigate the multilevel diversity in the large populations of metrics, subspaces, and clusters in a unified framework. To tackle this problem, this article proposes a novel multidiversified ensemble clustering approach. In particular, we create a large number of diversified metrics by randomizing a scaled exponential similarity kernel, which are then coupled with random subspaces to form a large set of metric-subspace pairs. Based on the similarity matrices derived from these metric-subspace pairs, an ensemble of diversified base clusterings can be thereby constructed. Furthermore, an entropy-based criterion is utilized to explore the cluster wise diversity in ensembles, based on which three specific ensemble clustering algorithms are presented by incorporating three types of consensus functions. Extensive experiments are conducted on 30 high-dimensional datasets, including 18 cancer gene expression datasets and 12 image/speech datasets, which demonstrate the superiority of our algorithms over the state of the art. The source code is available at https://github.com/huangdonghere/MDEC.
Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai, Chee Keong Kwoh 0001
IEEE Trans. Cybern.4
2021 Time-Series Representation Learning via Temporal and Contextual Contrasting
abstract
Learning decent representations from unlabeled time-series data with temporal dynamics is a very challenging task. In this paper, we propose an unsupervised Time-Series representation learning framework via Temporal and Contextual Contrasting (TS-TCC), to learn time-series representation from unlabeled data. First, the raw time-series data are transformed into two different yet correlated views by using weak and strong augmentations. Second, we propose a novel temporal contrasting module to learn robust temporal representations by designing a tough cross-view prediction task. Last, to further learn discriminative representations, we propose a contextual contrasting module built upon the contexts from the temporal contrasting module. It attempts to maximize the similarity among different contexts of the same sample while minimizing similarity among contexts of different samples. Experiments have been carried out on three real-world time-series datasets. The results manifest that training a linear classifier on top of the features learned by our proposed TS-TCC performs comparably with the supervised training. Additionally, our proposed TS-TCC shows high efficiency in few-labeled data and transfer learning scenarios. The code is publicly available at https://github.com/emadeldeen24/TS-TCC.
Emadeldeen Eldele, Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Cuntai Guan
IJCAI5
2021 Chromatin loop anchors predict transcript and exon usage
abstract
Epigenomics and transcriptomics data from high-throughput sequencing techniques such as RNA-seq and ChIP-seq have been successfully applied in predicting gene transcript expression. However, the locations of chromatin loops in the genome identified by techniques such as Chromatin Interaction Analysis with Paired End Tag sequencing (ChIA-PET) have never been used for prediction tasks. Here, we developed machine learning models to investigate if ChIA-PET could contribute to transcript and exon usage prediction. In doing so, we used a large set of transcription factors as well as ChIA-PET data. We developed different Gradient Boosting Trees models according to the different tasks with the integrated datasets from three cell lines, including GM12878, HeLaS3 and K562. We validated the models via 10-fold cross validation, chromosome-split validation and cross-cell validation. Our results show that both transcript and splicing-derived exon usage can be effectively predicted with at least 0.7512 and 0.7459 of accuracy, respectively, on all cell lines from all kinds of validations. Examining the predictive features, we found that RNA Polymerase II ChIA-PET was one of the most important features in both transcript and exon usage prediction, suggesting that chromatin loop anchors are predictive of both transcript and exon usage.
Yu Zhang 0084, Xavier Roca, Chee Keong Kwoh 0001, Melissa Jane Fullwood
Briefings Bioinform.4
2021 Recent advances in network-based methods for disease gene prediction
abstract
Disease-gene association through genome-wide association study (GWAS) is an arduous task for researchers. Investigating single nucleotide polymorphisms that correlate with specific diseases needs statistical analysis of associations. Considering the huge number of possible mutations, in addition to its high cost, another important drawback of GWAS analysis is the large number of false positives. Thus, researchers search for more evidence to cross-check their results through different sources. To provide the researchers with alternative and complementary low-cost disease-gene association evidence, computational approaches come into play. Since molecular networks are able to capture complex interplay among molecules in diseases, they become one of the most extensively used data for disease-gene association prediction. In this survey, we aim to provide a comprehensive and up-to-date review of network-based methods for disease gene prediction. We also conduct an empirical analysis on 14 state-of-the-art methods. To summarize, we first elucidate the task definition for disease gene prediction. Secondly, we categorize existing network-based efforts into network diffusion methods, traditional machine learning methods with handcrafted graph features and graph representation learning methods. Thirdly, an empirical analysis is conducted to evaluate the performance of the selected methods across seven diseases. We also provide distinguishing findings about the discussed methods based on our empirical analysis. Finally, we highlight potential research directions for future studies on disease gene prediction.
Sezin Kircali Ata, Min Wu 0008, Yuan Fang 0001, Le Ou-Yang, Chee Keong Kwoh 0001, Xiaoli Li 0001
Briefings Bioinform.5
2021 DeepCPP: a deep neural network based on nucleotide bias information and minimum distribution similarity feature selection for RNA coding potential prediction
abstract
The development of deep sequencing technologies has led to the discovery of novel transcripts. Many in silico methods have been developed to assess the coding potential of these transcripts to further investigate their functions. Existing methods perform well on distinguishing majority long noncoding RNAs (lncRNAs) and coding RNAs (mRNAs) but poorly on RNAs with small open reading frames (sORFs). Here, we present DeepCPP (deep neural network for coding potential prediction), a deep learning method for RNA coding potential prediction. Extensive evaluations on four previous datasets and six new datasets constructed in different species show that DeepCPP outperforms other state-of-the-art methods, especially on sORF type data, which overcomes the bottleneck of sORF mRNA identification by improving more than 4.31, 37.24 and 5.89% on its accuracy for newly discovered human, vertebrate and insect data, respectively. Additionally, we also revealed that discontinuous k-mer, and our newly proposed nucleotide bias and minimal distribution similarity feature selection method play crucial roles in this classification problem. Taken together, DeepCPP is an effective method for RNA coding potential prediction.
Yu Zhang 0084, Cangzhi Jia, Melissa Jane Fullwood, Chee Keong Kwoh 0001
Briefings Bioinform.4
2021 Predicting the interaction biomolecule types for lncRNA: an ensemble deep learning approach
abstract
Long noncoding RNAs (lncRNAs) play significant roles in various physiological and pathological processes via their interactions with biomolecules like DNA, RNA and protein. The existing in silico methods used for predicting the functions of lncRNA mainly rely on calculating the similarity of lncRNA or investigating whether an lncRNA can interact with a specific biomolecule or disease. In this work, we explored the functions of lncRNA from a different perspective: we presented a tool for predicting the interaction biomolecule type for a given lncRNA. For this purpose, we first investigated the main molecular mechanisms of the interactions of lncRNA-RNA, lncRNA-protein and lncRNA-DNA. Then, we developed an ensemble deep learning model: lncIBTP (lncRNA Interaction Biomolecule Type Prediction). This model predicted the interactions between lncRNA and different types of biomolecules. On the 5-fold cross-validation, the lncIBTP achieves average values of 0.7042 in accuracy, 0.7903 and 0.6421 in macro-average area under receiver operating characteristic curve and precision-recall curve, respectively, which illustrates the model effectiveness. Besides, based on the analysis of the collected published data and prediction results, we hypothesized that the characteristics of lncRNAs that interacted with DNA may be different from those that interacted with only RNA.
Yu Zhang 0084, Cangzhi Jia, Chee Keong Kwoh 0001
Briefings Bioinform.3
2021 Graph contextualized attention network for predicting synthetic lethality in human cancers
abstract
MOTIVATION: Synthetic Lethality (SL) plays an increasingly critical role in the targeted anticancer therapeutics. In addition, identifying SL interactions can create opportunities to selectively kill cancer cells without harming normal cells. Given the high cost of wet-lab experiments, in silico prediction of SL interactions as an alternative can be a rapid and cost-effective way to guide the experimental screening of candidate SL pairs. Several matrix factorization-based methods have recently been proposed for human SL prediction. However, they are limited in capturing the dependencies of neighbors. In addition, it is also highly challenging to make accurate predictions for new genes without any known SL partners. RESULTS: In this work, we propose a novel graph contextualized attention network named GCATSL to learn gene representations for SL prediction. First, we leverage different data sources to construct multiple feature graphs for genes, which serve as the feature inputs for our GCATSL method. Second, for each feature graph, we design node-level attention mechanism to effectively capture the importance of local and global neighbors and learn local and global representations for the nodes, respectively. We further exploit multi-layer perceptron (MLP) to aggregate the original features with the local and global representations and then derive the feature-specific representations. Third, to derive the final representations, we design feature-level attention to integrate feature-specific representations by taking the importance of different feature graphs into account. Extensive experimental results on three datasets under different settings demonstrated that our GCATSL model outperforms 14 state-of-the-art methods consistently. In addition, case studies further validated the effectiveness of our proposed model in identifying novel SL pairs. AVAILABILITYAND IMPLEMENTATION: Python codes and dataset are freely available on GitHub (https://github.com/longyahui/GCATSL) and Zenodo (https://zenodo.org/record/4522679) under the MIT license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yahui Long, Min Wu 0008, Yong Liu 0020, Jie Zheng 0002, Chee Keong Kwoh 0001, Jiawei Luo 0001, Xiaoli Li 0001
Bioinform.5
2021 VirPreNet: a weighted ensemble convolutional neural network for the virulence prediction of influenza A virus using all eight segments
abstract
MOTIVATION: Influenza viruses are persistently threatening public health, causing annual epidemics and sporadic pandemics. The evolution of influenza viruses remains to be the main obstacle in the effectiveness of antiviral treatments due to rapid mutations. Previous work has been investigated to reveal the determinants of virulence of the influenza A virus. To further facilitate flu surveillance, explicit detection of influenza virulence is crucial to protect public health from potential future pandemics. RESULTS: In this article, we propose a weighted ensemble convolutional neural network (CNN) for the virulence prediction of influenza A viruses named VirPreNet that uses all eight segments. Firstly, mouse lethal dose 50 is exerted to label the virulence of infections into two classes, namely avirulent and virulent. A numerical representation of amino acids named ProtVec is applied to the eight-segments in a distributed manner to encode the biological sequences. After splittings and embeddings of influenza strains, the ensemble CNN is constructed as the base model on the influenza dataset of each segment, which serves as the VirPreNet's main part. Followed by a linear layer, the initial predictive outcomes are integrated and assigned with different weights for the final prediction. The experimental results on the collected influenza dataset indicate that VirPreNet achieves state-of-the-art performance combining ProtVec with our proposed architecture. It outperforms baseline methods on the independent testing data. Moreover, our proposed model reveals the importance of PB2 and HA segments on the virulence prediction. We believe that our model may provide new insights into the investigation of influenza virulence. AVAILABILITY AND IMPLEMENTATION: Codes and data to generate the VirPreNet are publicly available at https://github.com/Rayin-saber/VirPreNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rui Yin 0002, Zihan Luo 0001, Pei Zhuang, Zhuoyi Lin, Chee Keong Kwoh 0001
Bioinform.5
2021 Class similarity network for coding and long non-coding RNA classification
abstract
BACKGROUND: Long non-coding RNAs (lncRNAs) play significant roles in varieties of physiological and pathological processes.The premise of the lncRNA functional study is that the lncRNAs are identified correctly. Recently, deep learning method like convolutional neural network (CNN) has been successfully applied to identify the lncRNAs. However, the traditional CNN considers little relationships among samples via an indirect way. RESULTS: Inspired by the Siamese Neural Network (SNN), here we propose a novel network named Class Similarity Network in coding RNA and lncRNA classification. Class Similarity Network considers more relationships among input samples in a direct way. It focuses on exploring the potential relationships between input samples and samples from both the same class and the different classes. To achieve this, Class Similarity Network trains the parameters specific to each class to obtain the high-level features and represents the general similarity to each class in a node. The comparison results on the validation dataset under the same conditions illustrate the superiority of our Class Similarity Network to the baseline CNN. Besides, our method performs effectively and achieves state-of-the-art performances on two test datasets. CONCLUSIONS: We construct Class Similarity Network in coding RNA and lncRNA classification, which is shown to work effectively on two different datasets by achieving accuracy, precision, and F1-score as 98.43%, 0.9247, 0.9374, and 97.54%, 0.9990, 0.9860, respectively.
Yu Zhang 0084, Yahui Long, Chee Keong Kwoh 0001
BMC Bioinform.3
2021 Attention-based sequence to sequence model for machine remaining useful life prediction
Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Ruqiang Yan 0001, Xiaoli Li 0001
Neurocomputing4
2021 GLIMG: Global and local item graphs for top-N recommender systems
Zhuoyi Lin, Lei Feng 0006, Rui Yin 0002, Chee Keong Kwoh 0001
Inf. Sci.5
2021 Contrastive Adversarial Domain Adaptation for Machine Remaining Useful Life Prediction
abstract
Enabling precise forecasting of the remaining useful life (RUL) for machines can reduce maintenance cost, increase availability, and prevent catastrophic consequences. Data-driven RUL prediction methods have already achieved acclaimed performance. However, they usually assume that the training and testing data are collected from the same condition (same distribution or domain), which is generally not valid in real industry. Conventional approaches to address domain shift problems attempt to derive domain-invariant features, but fail to consider target-specific information, leading to limited performance. To tackle this issue, in this article, we propose a contrastive adversarial domain adaptation (CADA) method for cross-domain RUL prediction. The proposed CADA approach is built upon an adversarial domain adaptation architecture with a contrastive loss, such that it is able to take target-specific information into consideration when learning domain-invariant features. To validate the superiority of the proposed approach, comprehensive experiments have been conducted to predict the RULs of aeroengines across 12 cross-domain scenarios. The experimental results show that the proposed method significantly outperforms state-of-the-arts with over 21% and 38% improvements in terms of two different evaluation metrics.
Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chuan-Sheng Foo, Chee Keong Kwoh 0001, Ruqiang Yan 0001, Xiaoli Li 0001
IEEE Trans. Ind. Informatics5
2021 Multi-View Collaborative Network Embedding
abstract
Real-world networks often exist with multiple views, where each view describes one type of interaction among a common set of nodes. For example, on a video-sharing network, while two user nodes are linked, if they have common favorite videos in one view, then they can also be linked in another view if they share common subscribers. Unlike traditional single-view networks, multiple views maintain different semantics to complement each other. In this article, we propose M ulti-view coll A borative N etwork E mbedding (MANE), a multi-view network embedding approach to learn low-dimensional representations. Similar to existing studies, MANE hinges on diversity and collaboration—while diversity enables views to maintain their individual semantics, collaboration enables views to work together. However, we also discover a novel form of second-order collaboration that has not been explored previously, and further unify it into our framework to attain superior node representations. Furthermore, as each view often has varying importance w.r.t. different nodes, we propose MANE , an attention -based extension of MANE, to model node-wise view importance. Finally, we conduct comprehensive experiments on three public, real-world multi-view networks, and the results demonstrate that our models consistently outperform state-of-the-art approaches.
Sezin Kircali Ata, Yuan Fang 0001, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001
ACM Trans. Knowl. Discov. Data5
2021 Enhanced Ensemble Clustering via Fast Propagation of Cluster-Wise Similarities
abstract
Ensemble clustering has been a popular research topic in data mining and machine learning. Despite its significant progress in recent years, there are still two challenging issues in the current ensemble clustering research. First, most of the existing algorithms tend to investigate the ensemble information at the object-level, yet often lack the ability to explore the rich information at higher levels of granularity. Second, they mostly focus on the direct connections (e.g., direct intersection or pair-wise co-occurrence) in the multiple base clusterings, but generally neglect the multiscale indirect relationship hidden in them. To address these two issues, this paper presents a novel ensemble clustering approach based on fast propagation of cluster-wise similarities via random walks. We first construct a cluster similarity graph with the base clusters treated as graph nodes and the cluster-wise Jaccard coefficient exploited to compute the initial edge weights. Upon the constructed graph, a transition probability matrix is defined, based on which the random walk process is conducted to propagate the graph structural information. Specifically, by investigating the propagating trajectories starting from different nodes, a new cluster-wise similarity matrix can be derived by considering the trajectory relationship. Then, the newly obtained cluster-wise similarity matrix is mapped from the cluster-level to the object-level to achieve an enhanced co-association matrix, which is able to simultaneously capture the object-wise co-occurrence relationship as well as the multiscale cluster-wise relationship in ensembles. Finally, two novel consensus functions are proposed to obtain the consensus clustering result. Extensive experiments on a variety of real-world datasets have demonstrated the effectiveness and efficiency of our approach.
Dong Huang 0001, Chang-Dong Wang 0001, Hongxing Peng, Jian-Huang Lai, Chee Keong Kwoh 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2020 Predicting Drugs for COVID-19/SARS-CoV-2 via Heterogeneous Graph Attention Networks
abstract
Coronavirus Disease-19 (COVID-19) has led to global epidemics with high morbidity and mortality. However, there are currently no proven effective drugs targeting COVID19. Identifying drug-virus associations can not only provide insights into the understanding of drug-virus interaction mechanism, but also guide and facilitate the screening of compound candidates for antiviral drug discovery. In this work, we propose a novel framework of Heterogeneous Graph Attention Networks for Drug-Virus Association predictions, named HGATDVA. First, we fully incorporate multiple sources of biomedical data to construct abundant features for drugs and viruses. Second, we construct two drug-virus heterogeneous graphs. For each graph, we design a self-enhanced graph attention network (SGAT) to explicitly model the dependency between a node and its local neighbors and derive the graph-specific representations for nodes. Third, we further develop a neural network architecture with tri-aggregator to aggregate the graph-specific representations to generate the final node representations. Experiments on two datasets were conducted to demonstrate the effectiveness of our proposed method in identifying candidate drugs for viruses.
Yahui Long, Yu Zhang 0084, Min Wu 0008, Shaoliang Peng, Chee Keong Kwoh 0001, Jiawei Luo 0001, Xiaoli Li 0001
BIBM5
2020 Spectral Clustering by Subspace Randomization and Graph Fusion for High-Dimensional Data
Xiaosha Cai, Dong Huang 0001, Chang-Dong Wang 0001, Chee Keong Kwoh 0001
PAKDD (1)4
2020 Heterogeneous information network and its application to human health and disease
abstract
The molecular components with the functional interdependencies in human cell form complicated biological network. Diseases are mostly caused by the perturbations of the composite of the interaction multi-biomolecules, rather than an abnormality of a single biomolecule. Furthermore, new biological functions and processes could be revealed by discovering novel biological entity relationships. Hence, more and more biologists focus on studying the complex biological system instead of the individual biological components. The emergence of heterogeneous information network (HIN) offers a promising way to systematically explore complicated and heterogeneous relationships between various molecules for apparently distinct phenotypes. In this review, we first present the basic definition of HIN and the biological system considered as a complex HIN. Then, we discuss the topological properties of HIN and how these can be applied to detect network motif and functional module. Afterwards, methodologies of discovering relationships between disease and biomolecule are presented. Useful insights on how HIN aids in drug development and explores human interactome are provided. Finally, we analyze the challenges and opportunities for uncovering combinatorial patterns among pharmacogenomics and cell-type detection based on single-cell genomic data.
Pingjian Ding, Wenjue Ouyang, Jiawei Luo 0001, Chee Keong Kwoh 0001
Briefings Bioinform.4
2020 Ensembling graph attention networks for human microbe-drug association prediction
abstract
MOTIVATION: Human microbes get closely involved in an extensive variety of complex human diseases and become new drug targets. In silico methods for identifying potential microbe-drug associations provide an effective complement to conventional experimental methods, which can not only benefit screening candidate compounds for drug development but also facilitate novel knowledge discovery for understanding microbe-drug interaction mechanisms. On the other hand, the recent increased availability of accumulated biomedical data for microbes and drugs provides a great opportunity for a machine learning approach to predict microbe-drug associations. We are thus highly motivated to integrate these data sources to improve prediction accuracy. In addition, it is extremely challenging to predict interactions for new drugs or new microbes, which have no existing microbe-drug associations. RESULTS: In this work, we leverage various sources of biomedical information and construct multiple networks (graphs) for microbes and drugs. Then, we develop a novel ensemble framework of graph attention networks with a hierarchical attention mechanism for microbe-drug association prediction from the constructed multiple microbe-drug graphs, denoted as EGATMDA. In particular, for each input graph, we design a graph convolutional network with node-level attention to learn embeddings for nodes (i.e. microbes and drugs). To effectively aggregate node embeddings from multiple input graphs, we implement graph-level attention to learn the importance of different input graphs. Experimental results under different cross-validation settings (e.g. the setting for predicting associations for new drugs) showed that our proposed method outperformed seven state-of-the-art methods. Case studies on predicted microbe-drug associations further demonstrated the effectiveness of our proposed EGATMDA method. AVAILABILITY: Source codes and supplementary materials are available at: https://github.com/longyahui/EGATMDA/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yahui Long, Min Wu 0008, Yong Liu 0020, Chee Keong Kwoh 0001, Jiawei Luo 0001, Xiaoli Li 0001
Bioinform.4
2020 Predicting human microbe-drug associations via graph convolutional network with conditional random field
abstract
MOTIVATION: Human microbes play critical roles in drug development and precision medicine. How to systematically understand the complex interaction mechanism between human microbes and drugs remains a challenge nowadays. Identifying microbe-drug associations can not only provide great insights into understanding the mechanism, but also boost the development of drug discovery and repurposing. Considering the high cost and risk of biological experiments, the computational approach is an alternative choice. However, at present, few computational approaches have been developed to tackle this task. RESULTS: In this work, we leveraged rich biological information to construct a heterogeneous network for drugs and microbes, including a microbe similarity network, a drug similarity network and a microbe-drug interaction network. We then proposed a novel graph convolutional network (GCN)-based framework for predicting human Microbe-Drug Associations, named GCNMDA. In the hidden layer of GCN, we further exploited the Conditional Random Field (CRF), which can ensure that similar nodes (i.e. microbes or drugs) have similar representations. To more accurately aggregate representations of neighborhoods, an attention mechanism was designed in the CRF layer. Moreover, we performed a random walk with restart-based scheme on both drug and microbe similarity networks to learn valuable features for drugs and microbes, respectively. Experimental results on three different datasets showed that our GCNMDA model consistently achieved better performance than seven state-of-the-art methods. Case studies for three microbes including SARS-CoV-2 and two antimicrobial drugs (i.e. Ciprofloxacin and Moxifloxacin) further confirmed the effectiveness of GCNMDA in identifying potential microbe-drug associations. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at: https://github.com/longyahui/GCNMDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yahui Long, Min Wu 0008, Chee Keong Kwoh 0001, Jiawei Luo 0001, Xiaoli Li 0001
Bioinform.3
2020 Tempel: time-series mutation prediction of influenza A viruses via attention-based recurrent neural networks
abstract
MOTIVATION: Influenza viruses are persistently threatening public health, causing annual epidemics and sporadic pandemics. The evolution of influenza viruses remains to be the main obstacle in the effectiveness of antiviral treatments due to rapid mutations. The goal of this work is to predict whether mutations are likely to occur in the next flu season using historical glycoprotein hemagglutinin sequence data. One of the major challenges is to model the temporality and dimensionality of sequential influenza strains and to interpret the prediction results. RESULTS: In this article, we propose an efficient and robust time-series mutation prediction model (Tempel) for the mutation prediction of influenza A viruses. We first construct the sequential training samples with splittings and embeddings. By employing recurrent neural networks with attention mechanisms, Tempel is capable of considering the historical residue information. Attention mechanisms are being increasingly used to improve the performance of mutation prediction by selectively focusing on the parts of the residues. A framework is established based on Tempel that enables us to predict the mutations at any specific residue site. Experimental results on three influenza datasets show that Tempel can significantly enhance the predictive performance compared with widely used approaches and provide novel insights into the dynamics of viral mutation and evolution. AVAILABILITY AND IMPLEMENTATION: The datasets, source code and supplementary documents are available at: https://drive.google.com/drive/folders/15WULR5__6k47iRotRPl3H7ghi3RpeNXH. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rui Yin 0002, Emil Luusua, Jan Dabrowski, Yu Zhang 0084, Chee Keong Kwoh 0001
Bioinform.5
2020 A random forest based computational model for predicting novel lncRNA-disease associations
abstract
BACKGROUND: Accumulated evidence shows that the abnormal regulation of long non-coding RNA (lncRNA) is associated with various human diseases. Accurately identifying disease-associated lncRNAs is helpful to study the mechanism of lncRNAs in diseases and explore new therapies of diseases. Many lncRNA-disease association (LDA) prediction models have been implemented by integrating multiple kinds of data resources. However, most of the existing models ignore the interference of noisy and redundancy information among these data resources. RESULTS: To improve the ability of LDA prediction models, we implemented a random forest and feature selection based LDA prediction model (RFLDA in short). First, the RFLDA integrates the experiment-supported miRNA-disease associations (MDAs) and LDAs, the disease semantic similarity (DSS), the lncRNA functional similarity (LFS) and the lncRNA-miRNA interactions (LMI) as input features. Then, the RFLDA chooses the most useful features to train prediction model by feature selection based on the random forest variable importance score that takes into account not only the effect of individual feature on prediction results but also the joint effects of multiple features on prediction results. Finally, a random forest regression model is trained to score potential lncRNA-disease associations. In terms of the area under the receiver operating characteristic curve (AUC) of 0.976 and the area under the precision-recall curve (AUPR) of 0.779 under 5-fold cross-validation, the performance of the RFLDA is better than several state-of-the-art LDA prediction models. Moreover, case studies on three cancers demonstrate that 43 of the 45 lncRNAs predicted by the RFLDA are validated by experimental data, and the other two predicted lncRNAs are supported by other LDA prediction models. CONCLUSIONS: Cross-validation and case studies indicate that the RFLDA has excellent ability to identify potential disease-associated lncRNAs.
Dengju Yao, Xiaojuan Zhan, Xiaorong Zhan, Chee Keong Kwoh 0001, Jinke Wang
BMC Bioinform.4
2020 Deep learning based DNA: RNA triplex forming potential prediction
abstract
BACKGROUND: Long non-coding RNAs (lncRNAs) can exert functions via forming triplex with DNA. The current methods in predicting the triplex formation mainly rely on mathematic statistic according to the base paring rules. However, these methods have two main limitations: (1) they identify a large number of triplex-forming lncRNAs, but the limited number of experimentally verified triplex-forming lncRNA indicates that maybe not all of them can form triplex in practice, and (2) their predictions only consider the theoretical relationship while lacking the features from the experimentally verified data. RESULTS: In this work, we develop an integrated program named TriplexFPP (Triplex Forming Potential Prediction), which is the first machine learning model in DNA:RNA triplex prediction. TriplexFPP predicts the most likely triplex-forming lncRNAs and DNA sites based on the experimentally verified data, where the high-level features are learned by the convolutional neural networks. In the fivefold cross validation, the average values of Area Under the ROC curves and PRC curves for removed redundancy triplex-forming lncRNA dataset with threshold 0.8 are 0.9649 and 0.9996, and these two values for triplex DNA sites prediction are 0.8705 and 0.9671, respectively. Besides, we also briefly summarize the cis and trans targeting of triplexes lncRNAs. CONCLUSIONS: The TriplexFPP is able to predict the most likely triplex-forming lncRNAs from all the lncRNAs with computationally defined triplex forming capacities and the potential of a DNA site to become a triplex. It may provide insights to the exploration of lncRNA functions.
Yu Zhang 0084, Yahui Long, Chee Keong Kwoh 0001
BMC Bioinform.3
2020 Ultra-Scalable Spectral Clustering and Ensemble Clustering
abstract
This paper focuses on scalability and robustness of spectral clustering for extremely large-scale datasets with limited resources. Two novel algorithms are proposed, namely, ultra-scalable spectral clustering (U-SPEC) and ultra-scalable ensemble clustering (U-SENC). In U-SPEC, a hybrid representative selection strategy and a fast approximation method for K-nearest representatives are proposed for the construction of a sparse affinity sub-matrix. By interpreting the sparse sub-matrix as a bipartite graph, the transfer cut is then utilized to efficiently partition the graph and obtain the clustering result. In U-SENC, multiple U-SPEC clusterers are further integrated into an ensemble clustering framework to enhance the robustness of U-SPEC while maintaining high efficiency. Based on the ensemble generation via multiple U-SEPC's, a new bipartite graph is constructed between objects and base clusters and then efficiently partitioned to achieve the consensus clustering result. It is noteworthy that both U-SPEC and U-SENC have nearly linear time and space complexity, and are capable of robustly and efficiently partitioning 10-million-level nonlinearly-separable datasets on a PC with 64 GB memory. Experiments on various large-scale datasets have demonstrated the scalability and robustness of our algorithms. The MATLAB code and experimental data are available at https://www.researchgate.net/publication/330760669.
Dong Huang 0001, Chang-Dong Wang 0001, Jian-Sheng Wu, Jian-Huang Lai, Chee Keong Kwoh 0001
IEEE Trans. Knowl. Data Eng.5
2019 Fast Top-N Personalized Recommendation on Item Graph
abstract
In the era of big data, traditional supply chain systems can not match the requirement of e-commerce. The analysis of customers’ demands and behaviors are necessary to exploit the potential insights and to build intelligent supply chain systems, which can be achieved by recommender systems. Graph-based recommendation models work well for top-N recommender systems due to their capability to capture the potential relationships between entities. In this paper, we propose a novel graph-based recommendation model to achieve personalized item ranking. To be specific, we design an adapted semi-supervised learning method to capture item smoothness, item fitting, and item confidence. By exploiting the structure of item graph moderately, the proposed method achieves impressive effectiveness and efficiency. In addition, extensive experimental results on real-world datasets show that our proposed method consistently outperforms the state-of-the-art counterparts on the top-N recommendation task.
Zhuoyi Lin, Lei Feng 0006, Chee Keong Kwoh 0001
IEEE BigData3
2019 Computational prediction of drug-target interactions using chemogenomic approaches: an empirical survey
abstract
Computational prediction of drug-target interactions (DTIs) has become an essential task in the drug discovery process. It narrows down the search space for interactions by suggesting potential interaction candidates for validation via wet-lab experiments that are well known to be expensive and time-consuming. In this article, we aim to provide a comprehensive overview and empirical evaluation on the computational DTI prediction techniques, to act as a guide and reference for our fellow researchers. Specifically, we first describe the data used in such computational DTI prediction efforts. We then categorize and elaborate the state-of-the-art methods for predicting DTIs. Next, an empirical comparison is performed to demonstrate the prediction performance of some representative methods under different scenarios. We also present interesting findings from our evaluation study, discussing the advantages and disadvantages of each method. Finally, we highlight potential avenues for further enhancement of DTI prediction performance as well as related research directions.
Ali Ezzat, Min Wu 0008, Xiaoli Li 0001, Chee Keong Kwoh 0001
Briefings Bioinform.4
2019 MULTiPly: a novel multi-layer predictor for discovering general and specific types of promoters
abstract
MOTIVATION: Promoters are short DNA consensus sequences that are localized proximal to the transcription start sites of genes, allowing transcription initiation of particular genes. However, the precise prediction of promoters remains a challenging task because individual promoters often differ from the consensus at one or more positions. RESULTS: In this study, we present a new multi-layer computational approach, called MULTiPly, for recognizing promoters and their specific types. MULTiPly took into account the sequences themselves, including both local information such as k-tuple nucleotide composition, dinucleotide-based auto covariance and global information of the entire samples based on bi-profile Bayes and k-nearest neighbour feature encodings. Specifically, the F-score feature selection method was applied to identify the best unique type of feature prediction results, in combination with other types of features that were subsequently added to further improve the prediction performance of MULTiPly. Benchmarking experiments on the benchmark dataset and comparisons with five state-of-the-art tools show that MULTiPly can achieve a better prediction performance on 5-fold cross-validation and jackknife tests. Moreover, the superiority of MULTiPly was also validated on a newly constructed independent test dataset. MULTiPly is expected to be used as a useful tool that will facilitate the discovery of both general and specific types of promoters in the post-genomic era. AVAILABILITY AND IMPLEMENTATION: The MULTiPly webserver and curated datasets are freely available at http://flagshipnt.erc.monash.edu/MULTiPly/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meng Zhang 0046, Fuyi Li, Tatiana T. Marquez-Lago, André Leier, Cunshuo Fan, Chee Keong Kwoh 0001, Kuo-Chen Chou, Jiangning Song, Cangzhi Jia
Bioinform.6
2019 An improved random forest-based computational model for predicting novel miRNA-disease associations
abstract
BACKGROUND: A large body of evidence shows that miRNA regulates the expression of its target genes at post-transcriptional level and the dysregulation of miRNA is related to many complex human diseases. Accurately discovering disease-related miRNAs is conductive to the exploring of the pathogenesis and treatment of diseases. However, because of the limitation of time-consuming and expensive experimental methods, predicting miRNA-disease associations by computational models has become a more economical and effective mean. RESULTS: Inspired by the work of predecessors, we proposed an improved computational model based on random forest (RF) for identifying miRNA-disease associations (IRFMDA). First, the integrated similarity of diseases and the integrated similarity of miRNAs were calculated by combining the semantic similarity and Gaussian interaction profile kernel (GIPK) similarity of diseases, the functional similarity and GIPK similarity of miRNAs, respectively. Then, the integrated similarity of diseases and the integrated similarity of miRNAs were combined to represent each miRNA-disease relationship pair. Next, the miRNA-disease relationship pairs contained in the HMDD (v2.0) database were considered positive samples, and the randomly constructed miRNA-disease relationship pairs not included in HMDD (v2.0) were considered negative samples. Next, the feature selection based on the variable importance score of RF was performed to choose more useful features to represent samples to optimize the model's ability of inferring miRNA-disease associations. Finally, a RF regression model was trained on reduced sample space to score the unknown miRNA-disease associations. The AUCs of IRFMDA under local leave-one-out cross-validation (LOOCV), global LOOCV and 5-fold cross-validation achieved 0.8728, 0.9398 and 0.9363, which were better than several excellent models for predicting miRNA-disease associations. Moreover, case studies on oesophageal cancer, lymphoma and lung cancer showed that 94 (oesophageal cancer), 98 (lymphoma) and 100 (lung cancer) of the top 100 disease-associated miRNAs predicted by IRFMDA were supported by the experimental data in the dbDEMC (v2.0) database. CONCLUSIONS: Cross-validation and case studies demonstrated that IRFMDA is an excellent miRNA-disease association prediction model, and can provide guidance and help for experimental studies on the regulatory mechanism of miRNAs in complex human diseases in the future.
Dengju Yao, Xiaojuan Zhan, Chee Keong Kwoh 0001
BMC Bioinform.3
2019 Ensemble Prediction of Synergistic Drug Combinations Incorporating Biological, Chemical, Pharmacological, and Network Knowledge
abstract
Combinatorial therapy may reduce drug side effects and improve drug efficacy, making combination therapy a promising strategy to treat complex diseases. However, in the existing computational methods, the natural properties and network knowledge of drugs have not been adequately and simultaneously considered, making it difficult to identify effective drug combinations. Computational methods that incorporate multiple sources of information (biological, chemical, pharmacological, and network knowledge) offer more opportunities to screen synergistic drug combinations. Therefore, we developed a novel Ensemble Prediction framework of Synergistic Drug Combinations (EPSDC) to accurately and efficiently predict drug combinations by integrating information from multiple-sources. EPSDC constructs feature vector of drug pair by concatenating different types of drug similarities, and then uses these groups in a feature-based base predictor. Next, transductive learning is applied on heterogeneous drug-target networks to achieve a network-based score for the drug pair. Finally, two types of ensemble rules are introduced to combine the feature-based score and the network-based score, and then potential drug combinations are prioritized. To demonstrate the effect of the ensemble rule, comprehensive experiments were conducted to compare single models and ensemble models. The experimental results indicated that our method outperformed the state-of-the-art method in five-fold cross validation and de novo prediction tests on the two benchmark datasets. We further analyzed the effect of maximum length of the meta-path and the impacts of different types of features. Moreover, the practical usefulness of our method was confirmed in the predicted novel drug combinations. The source code of EPSDC is available at https://github.com/KDDing/EPSDC.
Pingjian Ding, Rui Yin 0002, Jiawei Luo 0001, Chee Keong Kwoh 0001
IEEE J. Biomed. Health Informatics4
2017 Drug-Target Interaction Prediction with Graph Regularized Matrix Factorization
abstract
Experimental determination of drug-target interactions is expensive and time-consuming. Therefore, there is a continuous demand for more accurate predictions of interactions using computational techniques. Algorithms have been devised to infer novel interactions on a global scale where the input to these algorithms is a drug-target network (i.e., a bipartite graph where edges connect pairs of drugs and targets that are known to interact). However, these algorithms had difficulty predicting interactions involving new drugs or targets for which there are no known interactions (i.e., "orphan" nodes in the network). Since data usually lie on or near to low-dimensional non-linear manifolds, we propose two matrix factorization methods that use graph regularization in order to learn such manifolds. In addition, considering that many of the non-occurring edges in the network are actually unknown or missing cases, we developed a preprocessing step to enhance predictions in the "new drug" and "new target" cases by adding edges with intermediate interaction likelihood scores. In our cross validation experiments, our methods achieved better results than three other state-of-the-art methods in most cases. Finally, we simulated some "new drug" and "new target" cases and found that GRMF predicted the left-out interactions reasonably well.
Ali Ezzat, Peilin Zhao, Min Wu 0008, Xiaoli Li 0001, Chee Keong Kwoh 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2016 Clustering based active learning for biomedical Named Entity Recognition
abstract
The recognition and extraction of biomedical names is an essential task for the biomedical information extraction. However, the preparation of large annotated corpora hinders the training of the Named Entity Recognition (NER) systems. Active learning is reducing the needed manual annotation work in supervised learning task. In this work, we propose a novel clustering based active learning method for the biomedical NER task. We show that the underlying NER system using the proposed method outperforms those with other state of the art active learning methods, including density, Gibbs error and entropy based approaches, as well as the random selection. We compare variations of our proposed method and find the optimal design of the active learning method, which is to use the vector representation of named entities, and to select documents that are `representative' and `informative', as well as to use the Shared Nearest Neighbor (SNN) clustering approach. In particular, the optimal variant of the proposed method achieves a deficiency gain of 36.3% over the random selection.
Xu Han 0008, Chee Keong Kwoh 0001, Jung-Jae Kim 0001
IJCNN2
2016 Drug-target interaction prediction via class imbalance-aware ensemble learning
abstract
BACKGROUND: Multiple computational methods for predicting drug-target interactions have been developed to facilitate the drug discovery process. These methods use available data on known drug-target interactions to train classifiers with the purpose of predicting new undiscovered interactions. However, a key challenge regarding this data that has not yet been addressed by these methods, namely class imbalance, is potentially degrading the prediction performance. Class imbalance can be divided into two sub-problems. Firstly, the number of known interacting drug-target pairs is much smaller than that of non-interacting drug-target pairs. This imbalance ratio between interacting and non-interacting drug-target pairs is referred to as the between-class imbalance. Between-class imbalance degrades prediction performance due to the bias in prediction results towards the majority class (i.e. the non-interacting pairs), leading to more prediction errors in the minority class (i.e. the interacting pairs). Secondly, there are multiple types of drug-target interactions in the data with some types having relatively fewer members (or are less represented) than others. This variation in representation of the different interaction types leads to another kind of imbalance referred to as the within-class imbalance. In within-class imbalance, prediction results are biased towards the better represented interaction types, leading to more prediction errors in the less represented interaction types. RESULTS: We propose an ensemble learning method that incorporates techniques to address the issues of between-class imbalance and within-class imbalance. Experiments show that the proposed method improves results over 4 state-of-the-art methods. In addition, we simulated cases for new drugs and targets to see how our method would perform in predicting their interactions. New drugs and targets are those for which no prior interactions are known. Our method displayed satisfactory prediction performance and was able to predict many of the interactions successfully. CONCLUSIONS: Our proposed method has improved the prediction performance over the existing work, thus proving the importance of addressing problems pertaining to class imbalance in the data.
Ali Ezzat, Min Wu 0008, Xiaoli Li 0001, Chee Keong Kwoh 0001
BMC Bioinform.4
2016 Gene, Environment and Methylation (GEM): a tool suite to efficiently navigate large scale epigenome wide association studies and integrate genotype and interaction between genotype and environment
abstract
BACKGROUND: The interplay among genetic, environment and epigenetic variation is not fully understood. Advances in high-throughput genotyping methods, high-density DNA methylation detection and well-characterized sample collections, enable epigenetic association studies at the genomic and population levels (EWAS). The field has extended to interrogate the interaction of environmental and genetic (GxE) influences on epigenetic variation. Also, the detection of methylation quantitative trait loci (methQTLs) and their association with health status has enhanced our knowledge of epigenetic mechanisms in disease trajectory. However analysis of this type of data brings computational challenges and there are few practical solutions to enable large scale studies in standard computational environments. RESULTS: GEM is a highly efficient R tool suite for performing epigenome wide association studies (EWAS). GEM provides three major functions named GEM_Emodel, GEM_Gmodel and GEM_GxEmodel to study the interplay of Gene, Environment and Methylation (GEM). Within GEM, the pre-existing "Matrix eQTL" package is utilized and extended to study methylation quantitative trait loci (methQTL) and the interaction of genotype and environment (GxE) to determine DNA methylation variation, using matrix based iterative correlation and memory-efficient data analysis. Benchmarking presented here on a publicly available dataset, demonstrated that GEM can facilitate reliable genome-wide methQTL and GxE analysis on a standard laptop computer within minutes. CONCLUSIONS: The GEM package facilitates efficient EWAS study in large cohorts. It is written in R code and can be freely downloaded from Bioconductor at https://www.bioconductor.org/packages/GEM/ .
Joanna D. Holbrook, Neerja Karnani, Chee Keong Kwoh 0001
BMC Bioinform.4
2016 Cross-Examination for Angle-Closure Glaucoma Feature Detection
abstract
Effective feature selection plays a vital role in anterior segment imaging for determining the mechanism involved in angle-closure glaucoma (ACG) diagnosis. This research focuses on the use of redundant features for complex disease diagnosis such as ACG using anterior segment optical coherence tomography images. Both supervised [minimum redundancy maximum relevance (MRMR)] and unsupervised [Laplacian score (L-score)] feature selection algorithms have been cross-examined with different ACG mechanisms. An AdaBoost machine learning classifier is then used for classifying the five various classes of ACG mechanism such as iris roll, lens, pupil block, plateau iris, and no mechanism using both feature selection methods. The overall accuracy has shown that the usefulness of redundant features by L-score method in improved ACG diagnosis compared to minimum redundant features by MRMR method.
S. Issac Niwas, Weisi Lin, Chee Keong Kwoh 0001, C.-C. Jay Kuo, Chelvin C. Sng, Maria Cecilia Aquino, Paul T. K. Chew
IEEE J. Biomed. Health Informatics3
2015 Fast, accurate, and reliable molecular docking with QuickVina 2
abstract
MOTIVATION: The need for efficient molecular docking tools for high-throughput screening is growing alongside the rapid growth of drug-fragment databases. AutoDock Vina ('Vina') is a widely used docking tool with parallelization for speed. QuickVina ('QVina 1') then further enhanced the speed via a heuristics, requiring high exhaustiveness. With low exhaustiveness, its accuracy was compromised. We present in this article the latest version of QuickVina ('QVina 2') that inherits both the speed of QVina 1 and the reliability of the original Vina. RESULTS: We tested the efficacy of QVina 2 on the core set of PDBbind 2014. With the default exhaustiveness level of Vina (i.e. 8), a maximum of 20.49-fold and an average of 2.30-fold acceleration with a correlation coefficient of 0.967 for the first mode and 0.911 for the sum of all modes were attained over the original Vina. A tendency for higher acceleration with increased number of rotatable bonds as the design variables was observed. On the accuracy, Vina wins over QVina 2 on 30% of the data with average energy difference of only 0.58 kcal/mol. On the same dataset, GOLD produced RMSD smaller than 2 Å on 56.9% of the data while QVina 2 attained 63.1%. AVAILABILITY AND IMPLEMENTATION: The C++ source code of QVina 2 is available at (www.qvina.org). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Amr Alhossary, Stephanus Daniel Handoko, Yuguang Mu, Chee Keong Kwoh 0001
Bioinform.4
2015 Efficient and Accurate OTU Clustering with GPU-Based Sequence Alignment and Dynamic Dendrogram Cutting
abstract
De novo clustering is a popular technique to perform taxonomic profiling of a microbial community by grouping 16S rRNA amplicon reads into operational taxonomic units (OTUs). In this work, we introduce a new dendrogram-based OTU clustering pipeline called CRiSPy. The key idea used in CRiSPy to improve clustering accuracy is the application of an anomaly detection technique to obtain a dynamic distance cutoff instead of using the de facto value of 97 percent sequence similarity as in most existing OTU clustering pipelines. This technique works by detecting an abrupt change in the merging heights of a dendrogram. To produce the output dendrograms, CRiSPy employs the OTU hierarchical clustering approach that is computed on a genetic distance matrix derived from an all-against-all read comparison by pairwise sequence alignment. However, most existing dendrogram-based tools have difficulty processing datasets larger than 10,000 unique reads due to high computational complexity. We address this difficulty by developing two efficient algorithms for CRiSPy: a compute-efficient GPU-accelerated parallel algorithm for pairwise distance matrix computation and a memory-efficient hierarchical clustering algorithm. Our experiments on various datasets with distinct attributes show that CRiSPy is able to produce more accurate OTU groupings than most OTU clustering applications.
Thuy-Diem Nguyen, Bertil Schmidt, Zejun Zheng, Chee Keong Kwoh 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2014 Integrating water exclusion theory into βcontacts to predict binding free energy changes and binding hot spots
abstract
BACKGROUND: Binding free energy and binding hot spots at protein-protein interfaces are two important research areas for understanding protein interactions. Computational methods have been developed previously for accurate prediction of binding free energy change upon mutation for interfacial residues. However, a large number of interrupted and unimportant atomic contacts are used in the training phase which caused accuracy loss. RESULTS: This work proposes a new method, βACVASA, to predict the change of binding free energy after alanine mutations. βACVASA integrates accessible surface area (ASA) and our newly defined β contacts together into an atomic contact vector (ACV). A β contact between two atoms is a direct contact without being interrupted by any other atom between them. A β contact's potential contribution to protein binding is also supposed to be inversely proportional to its ASA to follow the water exclusion hypothesis of binding hot spots. Tested on a dataset of 396 alanine mutations, our method is found to be superior in classification performance to many other methods, including Robetta, FoldX, HotPOINT, an ACV method of β contacts without ASA integration, and ACVASA methods (similar to βACVASA but based on distance-cutoff contacts). Based on our data analysis and results, we can draw conclusions that: (i) our method is powerful in the prediction of binding free energy change after alanine mutation; (ii) β contacts are better than distance-cutoff contacts for modeling the well-organized protein-binding interfaces; (iii) β contacts usually are only a small fraction number of the distance-based contacts; and (iv) water exclusion is a necessary condition for a residue to become a binding hot spot. CONCLUSIONS: βACVASA is designed using the advantages of both β contacts and water exclusion. It is an excellent tool to predict binding free energy changes and binding hot spots after alanine mutation.
Qian Liu 0014, Steven C. H. Hoi, Chee Keong Kwoh 0001, Limsoon Wong, Jinyan Li 0001
BMC Bioinform.3
2014 IFACEwat: the interfacial water-implemented re-ranking algorithm to improve the discrimination of near native structures for protein rigid docking
abstract
Protein-protein docking is an in silico method to predict the formation of protein complexes. Due to limited computational resources, the protein-protein docking approach has been developed under the assumption of rigid docking, in which one of the two protein partners remains rigid during the protein associations and water contribution is ignored or implicitly presented. Despite obtaining a number of acceptable complex predictions, it seems to-date that most initial rigid docking algorithms still find it difficult or even fail to discriminate successfully the correct predictions from the other incorrect or false positive ones. To improve the rigid docking results, re-ranking is one of the effective methods that help re-locate the correct predictions in top high ranks, discriminating them from the other incorrect ones. In this paper, we propose a new re-ranking technique using a new energy-based scoring function, namely IFACEwat - a combined Interface Atomic Contact Energy (IFACE) and water effect. The IFACEwat aims to further improve the discrimination of the near-native structures of the initial rigid docking algorithm ZDOCK3.0.2. Unlike other re-ranking techniques, the IFACEwat explicitly implements interfacial water into the protein interfaces to account for the water-mediated contacts during the protein interactions. Our results showed that the IFACEwat increased both the numbers of the near-native structures and improved their ranks as compared to the initial rigid docking ZDOCK3.0.2. In fact, the IFACEwat achieved a success rate of 83.8% for Antigen/Antibody complexes, which is 10% better than ZDOCK3.0.2. As compared to another re-ranking technique ZRANK, the IFACEwat obtains success rates of 92.3% (8% better) and 90% (5% better) respectively for medium and difficult cases. When comparing with the latest published re-ranking method F 2 Dock, the IFACEwat performed equivalently well or even better for several Antigen/Antibody complexes. With the inclusion of interfacial water, the IFACEwat improves mostly results of the initial rigid docking, especially for Antigen/Antibody complexes. The improvement is achieved by explicitly taking into account the contribution of water during the protein interactions, which was ignored or not fully presented by the initial rigid docking and other re-ranking techniques. In addition, the IFACEwat maintains sufficient computational efficiency of the initial docking algorithm, yet improves the ranks as well as the number of the near native structures found. As our implementation so far targeted to improve the results of ZDOCK3.0.2, and particularly for the Antigen/Antibody complexes, it is expected in the near future that more implementations will be conducted to be applicable for other initial rigid docking algorithms.
Chinh Tran To Su, Thuy-Diem Nguyen, Jie Zheng 0002, Chee Keong Kwoh 0001
BMC Bioinform.4
2014 LDsplit: screening for cis-regulatory motifs stimulating meiotic recombination hotspots by analysis of DNA sequence polymorphisms
abstract
BACKGROUND: As a fundamental genomic element, meiotic recombination hotspot plays important roles in life sciences. Thus uncovering its regulatory mechanisms has broad impact on biomedical research. Despite the recent identification of the zinc finger protein PRDM9 and its 13-mer binding motif as major regulators for meiotic recombination hotspots, other regulators remain to be discovered. Existing methods for finding DNA sequence motifs of recombination hotspots often rely on the enrichment of co-localizations between hotspots and short DNA patterns, which ignore the cross-individual variation of recombination rates and sequence polymorphisms in the population. Our objective in this paper is to capture signals encoded in genetic variations for the discovery of recombination-associated DNA motifs. RESULTS: Recently, an algorithm called "LDsplit" has been designed to detect the association between single nucleotide polymorphisms (SNPs) and proximal meiotic recombination hotspots. The association is measured by the difference of population recombination rates at a hotspot between two alleles of a candidate SNP. Here we present an open source software tool of LDsplit, with integrative data visualization for recombination hotspots and their proximal SNPs. Applying LDsplit on SNPs inside an established 7-mer motif bound by PRDM9 we observed that SNP alleles preserving the original motif tend to have higher recombination rates than the opposite alleles that disrupt the motif. Running on SNP windows around hotspots each containing an occurrence of the 7-mer motif, LDsplit is able to guide the established motif finding algorithm of MEME to recover the 7-mer motif. In contrast, without LDsplit the 7-mer motif could not be identified. CONCLUSIONS: LDsplit is a software tool for the discovery of cis-regulatory DNA sequence motifs stimulating meiotic recombination hotspots by screening and narrowing down to hotspot associated SNPs. It is the first computational method that utilizes the genetic variation of recombination hotspots among individuals, opening a new avenue for motif finding. Tested on an established motif and simulated datasets, LDsplit shows promise to discover novel DNA motifs for meiotic recombination hotspots.
Peng Yang 0010, Min Wu 0008, Chee Keong Kwoh 0001, Teresa M. Przytycka, Jie Zheng 0002
BMC Bioinform.4
2014 Reliable and Fast Estimation of Recombination Rates by Convergence Diagnosis and Parallel Markov Chain Monte Carlo
abstract
Genetic recombination is an essential event during the process of meiosis resulting in an exchange of segments between paired chromosomes. Estimating recombination rate is crucial for understanding the process of recombination. Experimental methods are normally difficult and limited to small scale estimations. Thus statistical methods using population genetics data are important for large-scale analysis. LDhat is an extensively used statistical method using rjMCMC algorithm to predict recombination rates. Due to the complexity of rjMCMC scheme, LDhat may take a long time for large SNP data sets. In addition, rjMCMC parameters should be manually defined in the original program which directly impact results. To address these issues, we designed an improved algorithm based on LDhat implementing MCMC convergence diagnostic algorithms to automatically predict values of parameters and monitor the mixing process. Then parallel computation methods were employed to further accelerate the new program. The new algorithms have been tested on ten samples from HapMap phase 2 data set. The results were compared with previous code and showed nearly identical output. However, our new methods achieved significant acceleration proving that they are more efficient and reliable for the estimation of recombination rates. The stand-alone package is freely available for download http://www.ntu.edu.sg/home/zhengjie/software/CPLDhat.
Ritika Jain, Peng Yang 0010, Chee Keong Kwoh 0001, Jie Zheng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2013 Review of tandem repeat search tools: a systematic approach to evaluating algorithmic performance
abstract
The prevalence of tandem repeats in eukaryotic genomes and their association with a number of genetic diseases has raised considerable interest in locating these repeats. Over the last 10-15 years, numerous tools have been developed for searching tandem repeats, but differences in the search algorithms adopted and difficulties with parameter settings have confounded many users resulting in widely varying results. In this review, we have systematically separated the algorithmic aspect of the search tools from the influence of the parameter settings. We hope that this will give a better understanding of how the tools differ in algorithmic performance, their inherent constraints and how one should approach in evaluating and selecting them.
Kian Guan Lim, Chee Keong Kwoh 0001, Li Yang Hsu, Adrianto Wirawan
Briefings Bioinform.2
2013 Drug-target interaction prediction by learning from local information and neighbors
abstract
MOTIVATION: In silico methods provide efficient ways to predict possible interactions between drugs and targets. Supervised learning approach, bipartite local model (BLM), has recently been shown to be effective in prediction of drug-target interactions. However, for drug-candidate compounds or target-candidate proteins that currently have no known interactions available, its pure 'local' model is not able to be learned and hence BLM may fail to make correct prediction when involving such kind of new candidates. RESULTS: We present a simple procedure called neighbor-based interaction-profile inferring (NII) and integrate it into the existing BLM method to handle the new candidate problem. Specifically, the inferred interaction profile is treated as label information and is used for model learning of new candidates. This functionality is particularly important in practice to find targets for new drug-candidate compounds and identify targeting drugs for new target-candidate proteins. Consistent good performance of the new BLM-NII approach has been observed in the experiment for the prediction of interactions between drugs and four categories of target proteins. Especially for nuclear receptors, BLM-NII achieves the most significant improvement as this dataset contains many drugs/targets with no interactions in the cross-validation. This demonstrates the effectiveness of the NII strategy and also shows the great potential of BLM-NII for prediction of compound-protein interactions. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jian-Ping Mei, Chee Keong Kwoh 0001, Peng Yang 0010, Xiaoli Li 0001, Jie Zheng 0002
Bioinform.2
2013 Structural analysis on mutation residues and interfacial water molecules for human TIM disease understanding
abstract
BACKGROUND: Human triosephosphate isomerase (HsTIM) deficiency is a genetic disease caused often by the pathogenic mutation E104D. This mutation, located at the side of an abnormally large cluster of water in the inter-subunit interface, reduces the thermostability of the enzyme. Why and how these water molecules are directly related to the excessive thermolability of the mutant have not been investigated in structural biology. RESULTS: This work compares the structure of the E104D mutant with its wild type counterparts. It is found that the water topology in the dimer interface of HsTIM is atypical, having a "wet-core-dry-rim" distribution with 16 water molecules tightly packed in a small deep region surrounded by 22 residues including GLU104. These water molecules are co-conserved with their surrounding residues in non-archaeal TIMs (dimers) but not conserved across archaeal TIMs (tetramers), indicating their importance in preserving the overall quaternary structure. As the structural permutation induced by the mutation is not significant, we hypothesize that the excessive thermolability of the E104D mutant is attributed to the easy propagation of atoms' flexibility from the surface into the core via the large cluster of water. It is indeed found that the B factor increment in the wet region is higher than other regions, and, more importantly, the B factor increment in the wet region is maintained in the deeply buried core. Molecular dynamics simulations revealed that for the mutant structure at normal temperature, a clear increase of the root-mean-square deviation is observed for the wet region contacting with the large cluster of interfacial water. Such increase is not observed for other interfacial regions or the whole protein. This clearly suggests that, in the E104D mutant, the large water cluster is responsible for the subunit interface flexibility and overall thermolability, and it ultimately leads to the deficiency of this enzyme. CONCLUSIONS: Our study reveals that a large cluster of water buried in protein interfaces is fragile and high-maintenance, closely related to the structure, function and evolution of the whole protein.
Ying He 0001, Qian Liu 0014, Limsoon Wong, Chee Keong Kwoh 0001, Hung T. Nguyen 0001, Jinyan Li 0001
BMC Bioinform.6
2013 Structural analysis of the novel influenza A (H7N9) viral Neuraminidase interactions with current approved neuraminidase inhibitors Oseltamivir, Zanamivir, and Peramivir in the presence of mutation R289K
abstract
BACKGROUND: Since late March 2013, there has been another global health concern with a sudden wave of flu infections by a novel strain of avian influenza A (H7N9) virus in China. To-date, there have been more than 100 infections with 23 deaths. It is more worrying as this viral strain has never been detected in humans and only been found to be of low-pathogenicity. Currently, there are 3 effective neuraminidase inhibitors for this H7N9 virus strain, i.e. oseltamivir, zanamivir, and peramivir. These drugs have been used for treatment of the H7N9 influenza in China. However, how these inhibitors work and affect the binding cavity of the novel H7N9 neuraminidase in the presence of potential mutations has not been disclosed. In our study, we investigate steric effects and subsequently show the conformational restraints of the inhibitor-binding site of the non-mutated and mutated H7N9 neuraminidase structures to different drug compounds. RESULTS: Combination of molecular docking and Molecular Dynamics simulation reveal that zanamivir forms more favorable and stable complex than oseltamivir and peramivir when binding to the active site of the H7N9 neuraminidase. And it is likely that the novel influenza A (H7N9) virus adopts a higher probability to acquire resistance to peramivir than the other two inhibitors. Conformational changes induced by the mutation R289K causes loss of number of hydrogen bonds between the inhibitors and the H7N9 viral neuraminidase in 2 out of 3 complexes. In addition, our results of binding-affinity relationships of the 3 inhibitors with the viral neuraminidase proteins of previous pandemics (H1N1, H5N1) and the current novel H7N9 reflected the extent of binding effectiveness of the 3 inhibitors to the novel H7N9 neuraminidase. CONCLUSIONS: The results are novel and specific for the A/Hangzhou/1/2013(H7N9) influenza strain. Furthermore, the protocol could be useful for further drug-binding analysis and prediction of future viral mutations to which the virus evolves through adaptation and acquires resistance to the current available drugs.
Chinh Tran To Su, Xuchang Ouyang, Jie Zheng 0002, Chee Keong Kwoh 0001
BMC Bioinform.4
2013 Molecular docking analysis of 2009-H1N1 and 2004-H5N1 influenza virus HLA-B*4405-restricted HA epitope candidates: implications for TCR cross-recognition and vaccine development
abstract
BACKGROUND: The pandemic 2009-H1N1 influenza virus circulated in the human population and caused thousands deaths worldwide. Studies on pandemic influenza vaccines have shown that T cell recognition to conserved epitopes and cross-reactive T cell responses are important when new strains emerge, especially in the absence of antibody cross-reactivity. In this work, using HLA-B*4405 and DM1-TCR structure model, we systematically generated high confidence conserved 2009-H1N1 T cell epitope candidates and investigated their potential cross-reactivity against H5N1 avian flu virus. RESULTS: Molecular docking analysis of differential DM1-TCR recognition of the 2009-H1N1 epitope candidates yielded a mosaic epitope (KEKMNTEFW) and potential H5N1 HA cross-reactive epitopes that could be applied as multivalent peptide towards influenza A vaccine development. Structural models of TCR cross-recognition between 2009-H1N1 and 2004-H5N1 revealed steric and topological effects of TCR contact residue mutations on TCR binding affinity. CONCLUSIONS: The results are novel with regard to HA epitopes and useful for developing possible vaccination strategies against the rapidly changing influenza viruses. Yet, the challenge of identifying epitope candidates that result in heterologous T cell immunity under natural influenza infection conditions can only be overcome if more structural data on the TCR repertoire become available.
Chinh Tran To Su, Christian Schönbach, Chee Keong Kwoh 0001
BMC Bioinform.3
2013 Using causality modeling and Fuzzy Lattice Reasoning algorithm for predicting blood glucose
Simon Fong 0001, Sabah Mohammed, Jinan Fiaidhi, Chee Keong Kwoh 0001
Expert Syst. Appl.4
2013 Research and applications: Automatic glaucoma diagnosis through medical imaging informatics
abstract
BACKGROUND: Computer-aided diagnosis for screening utilizes computer-based analytical methodologies to process patient information. Glaucoma is the leading irreversible cause of blindness. Due to the lack of an effective and standard screening practice, more than 50% of the cases are undiagnosed, which prevents the early treatment of the disease. OBJECTIVE: To design an automatic glaucoma diagnosis architecture automatic glaucoma diagnosis through medical imaging informatics (AGLAIA-MII) that combines patient personal data, medical retinal fundus image, and patient's genome information for screening. MATERIALS AND METHODS: 2258 cases from a population study were used to evaluate the screening software. These cases were attributed with patient personal data, retinal images and quality controlled genome data. Utilizing the multiple kernel learning-based classifier, AGLAIA-MII, combined patient personal data, major image features, and important genome single nucleotide polymorphism (SNP) features. RESULTS AND DISCUSSION: Receiver operating characteristic curves were plotted to compare AGLAIA-MII's performance with classifiers using patient personal data, images, and genome SNP separately. AGLAIA-MII was able to achieve an area under curve value of 0.866, better than 0.551, 0.722 and 0.810 by the individual personal data, image and genome information components, respectively. AGLAIA-MII also demonstrated a substantial improvement over the current glaucoma screening approach based on intraocular pressure. CONCLUSIONS: AGLAIA-MII demonstrates for the first time the capability of integrating patients' personal data, medical retinal image and genome information for automatic glaucoma diagnosis and screening in a large dataset from a population study. It paves the way for a holistic approach for automatic objective glaucoma diagnosis and screening.
Jiang Liu 0001, Zhuo Zhang 0001, Damon Wing Kee Wong, Yanwu Xu 0001, Fengshou Yin, Jun Cheng 0003, Ngan Meng Tan, Chee Keong Kwoh 0001, Dong Xu 0001, Tin Aung, Tien Yin Wong
J. Am. Medical Informatics Assoc.8
2012 GPU Accelerated Molecular Docking with Parallel Genetic Algorithm
abstract
Molecular docking is a widely used tool in Computer-aided Drug Design and Discovery. Due to the complexity of simulating the chemical events when two molecules interact, highly accelerated molecular docking programs are of great interest and importance for practical use. In this paper, we present a GPU accelerated docking program implemented with CUDA. The hardware-enabled texture interpolation is employed for fast energy evaluation. Two types of parallel genetic algorithms are mapped to the CUDA computing architecture and used for the search of optimal docking result. Comparing to the CPU implementation, the GPU accelerated docking program achieved significant speedup while producing comparable results to the CPU version. The source code is made public at http://code.google.com/p/cudock/.
Xuchang Ouyang, Chee Keong Kwoh 0001
ICPADS2
2012 Positive-unlabeled learning for disease gene identification
abstract
BACKGROUND: Identifying disease genes from human genome is an important but challenging task in biomedical research. Machine learning methods can be applied to discover new disease genes based on the known ones. Existing machine learning methods typically use the known disease genes as the positive training set P and the unknown genes as the negative training set N (non-disease gene set does not exist) to build classifiers to identify new disease genes from the unknown genes. However, such kind of classifiers is actually built from a noisy negative set N as there can be unknown disease genes in N itself. As a result, the classifiers do not perform as well as they could be. RESULT: Instead of treating the unknown genes as negative examples in N, we treat them as an unlabeled set U. We design a novel positive-unlabeled (PU) learning algorithm PUDI (PU learning for disease gene identification) to build a classifier using P and U. We first partition U into four sets, namely, reliable negative set RN, likely positive set LP, likely negative set LN and weak negative set WN. The weighted support vector machines are then used to build a multi-level classifier based on the four training sets and positive training set P to identify disease genes. Our experimental results demonstrate that our proposed PUDI algorithm outperformed the existing methods significantly. CONCLUSION: The proposed PUDI algorithm is able to identify disease genes more accurately by treating the unknown data more appropriately as unlabeled set U instead of negative set N. Given that many machine learning problems in biomedical research do involve positive and unlabeled data instead of negative data, it is possible that the machine learning methods for these problems can be further improved by adopting PU learning methods, as we have done here for disease gene identification. AVAILABILITY AND IMPLEMENTATION: The executable program and data are available at http://www1.i2r.a-star.edu.sg/~xlli/PUDI/PUDI.html.
Peng Yang 0010, Xiaoli Li 0001, Jian-Ping Mei, Chee Keong Kwoh 0001, See-Kiong Ng
Bioinform.4
2012 QuickVina: Accelerating AutoDock Vina Using Gradient-Based Heuristics for Global Optimization
abstract
Predicting binding between macromolecule and small molecule is a crucial phase in the field of rational drug design. AutoDock Vina, one of the most widely used docking software released in 2009, uses an empirical scoring function to evaluate the binding affinity between the molecules and employs the iterated local search global optimizer for global optimization, achieving a significantly improved speed and better accuracy of the binding mode prediction compared its predecessor, AutoDock 4. In this paper, we propose further improvement in the local search algorithm of Vina by heuristically preventing some intermediate points from undergoing local search. Our improved version of Vina-dubbed QVina-achieved a maximum acceleration of about 25 times with the average speed-up of 8.34 times compared to the original Vina when tested on a set of 231 protein-ligand complexes while maintaining the optimal scores mostly identical. Using our heuristics, larger number of different ligands can be quickly screened against a given receptor within the same time frame.
Stephanus Daniel Handoko, Xuchang Ouyang, Chinh Tran To Su, Chee Keong Kwoh 0001, Yew-Soon Ong
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 Erratum to "QuickVina: Accelerating AutoDock Vina Using Gradient-Based Heuristics for Global Optimization"
Stephanus Daniel Handoko, Xuchang Ouyang, Chinh Tran To Su, Chee Keong Kwoh 0001, Yew-Soon Ong
IEEE ACM Trans. Comput. Biol. Bioinform.4
2011 Prediction of Trans-regulators of Recombination Hotspots in Mouse Genome
abstract
The regulatory mechanism of recombination is a fundamental problem in genomics, with wide applications in genome wide association studies, birth-defect diseases, molecular evolution, cancer research, etc. In mammalian genomes, recombination events cluster into short genomic regions called ¡§recombination hotspots¡¨. Recently, a 13-mer motif enriched in hotspots is identified as a candidate cis-regulatory element of human recombination hotspots, moreover, a zinc finger protein, PRDM9, binds to this motif and is associated with variation of recombination phenotype in human and mouse genomes, thus is a trans-acting regulator of recombination hotspots. However, this pair of cis and trans-regulators covers only a fraction of hotspots, thus other regulators of recombination hotspots remain to be discovered. In this paper, we propose an approach to predicting additional trans-regulators from DNA-binding proteins by comparing their enrichment of binding sites in hotspots. Applying this approach on newly mapped mouse hotspots genome-wide, we confirmed that PRDM9 is a major trans-regulator of hotspots. In addition, a list of top candidate trans-regulators of mouse hotspots is reported. Using GO analysis we observed that the top genes are enriched with function of his tone modification, highlighting the epigenetic regulatory mechanisms of recombination hotspots.
Min Wu 0008, Chee Keong Kwoh 0001, Teresa M. Przytycka, Jing Li 0002, Jie Zheng 0002
BIBM2
2011 Classification-assisted memetic algorithms for solving optimization problems with restricted equality constraint function mapping
abstract
The success of Memetic Algorithms (MAs) has driven many researchers to be more focused on the efficiency aspect of the algorithms such that it would be possible to effectively employ MAs to solve computationally expensive optimization problems where single evaluation of the objective and constraint functions may require minutes to hours of CPU time. One of the important design issues in MAs is the choice of the individuals upon which local search procedure should be applied. Selecting only some potential individuals lessens the demand for functional evaluations hence accelerates convergence to the global optimum. In recent years, advances have been made targeting optimization problems with single equality constraint h(x) = 0. The presence of previously evaluated candidate solutions with different signs of constraint values within some localities thus allows the estimation of the constraint boundary. An individual will undergo local search only if it is sufficiently close to the approximated boundary. Elegant as it may seem, the approach had unfortunately assumed that every constraint function maps the design variables to optimize into unbounded real values. This, however, may not always be the case in practice. In this paper, we present a strategy to efficiently solve constrained problems with a single equality constraint; the function of which maps the design variables into restricted (either strictly non-negative or strictly non-positive) real values only.
Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Yew-Soon Ong, Jonathan H. Chan
IEEE Congress on Evolutionary Computation2
2011 Structural analysis of the hot spots in the binding between H1N1 HA and the 2D1 antibody: do mutations of H1N1 from 1918 to 2009 affect much on this binding?
abstract
MOTIVATION: Worldwide and substantial mortality caused by the 2009 H1N1 influenza A has stimulated a new surge of research on H1N1 viruses. An epitope conservation has been learned in the HA1 protein that allows antibodies to cross-neutralize both 1918 and 2009 H1N1. However, few works have thoroughly studied the binding hot spots in those two antigen-antibody interfaces which are responsible for the antibody cross-neutralization. RESULTS: We apply predictive methods to identify binding hot spots at the epitope sites of the HA1 proteins and at the paratope sites of the 2D1 antibody. We find that the six mutations at the HA1's epitope from 1918 to 2009 should not harm its binding to 2D1. Instead, the change of binding free energy on the whole exhibits an increased tendency after these mutations, making the binding stronger. This is consistent with the observation that the 1918 H1N1 neutralizing antibody can cross-react with 2009 H1N1. We identified three distinguished hot spot residues, including Lys(166), common between the two epitopes. These common hot spots again can explain why 2D1 cross-reacted. We believe that these hot spot residues are mutation candidates which may help H1N1 viruses to evade the immune system. We also identified eight residues at the paratope site of 2D1, five from its heavy chain and three from its light chain, that are predicted to be energetically important in the HA1 recognition. The identification of these hot spot residues and their structural analysis are potentially useful to fight against H1N1 viruses. CONTACT: [email protected] AVAILABILITY: Z-score is available at http://155.69.2.25/liuqian/indexz.py SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qian Liu 0014, Steven C. H. Hoi, Chinh Tran To Su, Chee Keong Kwoh 0001, Limsoon Wong, Jinyan Li 0001
Bioinform.5
2011 Construction of co-complex score matrix for protein complex prediction from AP-MS data
abstract
MOTIVATION: Protein complexes are of great importance for unraveling the secrets of cellular organization and function. The AP-MS technique has provided an effective high-throughput screening to directly measure the co-complex relationship among multiple proteins, but its performance suffers from both false positives and false negatives. To computationally predict complexes from AP-MS data, most existing approaches either required the additional knowledge from known complexes (supervised learning), or had numerous parameters to tune. METHOD: In this article, we propose a novel unsupervised approach, without relying on the knowledge of existing complexes. Our method probabilistically calculates the affinity between two proteins, where the affinity score is evaluated by a co-complexed score or C2S in brief. In particular, our method measures the log-likelihood ratio of two proteins being co-complexed to being drawn randomly, and we then predict protein complexes by applying hierarchical clustering algorithm on the C2S score matrix. RESULTS: Compared with existing approaches, our approach is computationally efficient and easy to implement. It has just one parameter to set and its value has little effect on the results. It can be applied to different species as long as the AP-MS data are available. Despite its simplicity, it is competitive or superior in performance over many aspects when compared with the state-of-the-art predictions performed by supervised or unsupervised approaches.
Zhipeng Xie, Chee Keong Kwoh 0001, Xiaoli Li 0001, Min Wu 0008
Bioinform.2
2010 Outcomes of gene association analysis of cancer microarray data are impacted by pre-processing algorithms
abstract
Gene association analysis of cancer microarray data provides a wealth of information on gene expression patterns and cancer pathways to enhance the identification of potential biomarkers for cancer diagnosis, prognosis, and prediction of therapeutic responsiveness. However, achieving these biological/clinical objectives relies heavily on the functional capabilities and accuracy of the various analytical tools to mine these cancer microarray gene expression profiles. Many preprocessing algorithms exist for analyzing Affymetrix microarray gene expression data. Previous studies have evaluated these algorithms on their capabilities in accurately determining gene expression using a variety of spike-in as well as experimental data sets. However, variations in detecting differentially expressed genes between these different pre-processing algorithms on a single cancer dataset have not been done in a systems-level evaluation. In this study, we assessed the comparability and the level of variation between PLIER, GCRMA, RMA and MAS5 for their capability to detect differentially expressed genes.
N. Baskaran, Chee Keong Kwoh 0001, Kam M. Hui
BIBM2
2010 A possible mutation that enables H1N1 influenza a virus to escape antibody recognition
abstract
The H1N1 influenza A 2009 pandemic caused a global concern as it has killed more than 18,000 people worldwide so far. Studies that have found cross-neutralizing antibodies between the 1918 and 2009 pandemic flu elicit a basis of pre-existing immunity against the 2009 H1N1 virus in old population. The cross-reactivity occurs due to conserved antigenic epitopes shared between the two pandemic viruses. However, evolutionary mutation can enable the virus to elude human immunity system, making these antibodies probably no longer effective. In our study, we found that a possible mutation in B-cell epitope (the sequence PNHDSNKG) could be the chance for the virus to escape the 1918 antibody recognition. Hence, this finding can be helpful for further vaccine designs against the H1N1 2009 influenza A virus.
Chinh Tran To Su, Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Christian Schönbach, Xiaoli Li 0001
BIBM3
2010 Multi-threaded vectorized distance matrix computation on the CELL/BE and x86/SSE2 architectures
abstract
SUMMARY: Multiple sequence alignment is an important tool in bioinformatics. Although efficient heuristic algorithms exist for this problem, the exponential growth of biological data demands an even higher throughput. The recent emergence of multi-core technologies has made it possible to achieve a highly improved execution time for many bioinformatics applications. In this article, we introduce an implementation that accelerates the distance matrix computation on x86 and Cell Broadband Engine, a homogeneous and heterogeneous multi-core system, respectively. By taking advantage of multiple processors as well as Single Instruction Multiple Data vectorization, we were able to achieve speed-ups of two orders of magnitude compared to the publicly available implementation utilized in ClustalW. AVAILABILITY AND IMPLEMENTATION: Source codes in C are publicly available at https://sourceforge.net/projects/distmatcomp/ CONTACT: [email protected]
Adrianto Wirawan, Chee Keong Kwoh 0001, Bertil Schmidt
Bioinform.2
2010 Integrating diverse biological and computational sources for reliable protein-protein interactions
abstract
BACKGROUND: Protein-protein interactions (PPIs) play important roles in various cellular processes. However, the low quality of current PPI data detected from high-throughput screening techniques has diminished the potential usefulness of the data. We need to develop a method to address the high data noise and incompleteness of PPI data, namely, to filter out inaccurate protein interactions (false positives) and predict putative protein interactions (false negatives). RESULTS: In this paper, we proposed a novel two-step method to integrate diverse biological and computational sources of supporting evidence for reliable PPIs. The first step, interaction binning or InterBIN, groups PPIs together to more accurately estimate the likelihood (Bin-Confidence score) that the protein pairs interact for each biological or computational evidence source. The second step, interaction classification or InterCLASS, integrates the collected Bin-Confidence scores to build classifiers and identify reliable interactions. CONCLUSIONS: We performed comprehensive experiments on two benchmark yeast PPI datasets. The experimental results showed that our proposed method can effectively eliminate false positives in detected PPIs and identify false negatives by predicting novel yet reliable PPIs. Our proposed method also performed significantly better than merely using each of individual evidence sources, illustrating the importance of integrating various biological and computational sources of data and evidence.
Min Wu 0008, Xiaoli Li 0001, Hon Nian Chua, Chee Keong Kwoh 0001, See-Kiong Ng
BMC Bioinform.4
2010 Feasibility Structure Modeling: An Effective Chaperone for Constrained Memetic Algorithms
abstract
An important issue in designing memetic algorithms (MAs) is the choice of solutions in the population for local refinements, which becomes particularly crucial when solving computationally expensive problems. With single evaluation of the objective/constraint functions necessitating tremendous computational power and time, it is highly desirable to be able to focus search efforts on the regions where the global optimum is potentially located so as not to waste too many function evaluations. For constrained optimization, the global optimum must either be located at the trough of some feasible basin or some particular point along the feasibility boundary. Presented in this paper is an instance of optinformatics where a new concept of modeling the feasibility structure of inequality-constrained optimization problems-dubbed the feasibility structure modeling-is proposed to perform geometrical predictions of the locations of candidate solutions in the solution space: deep inside any infeasible region, nearby any feasibility boundary, or deep inside any feasible region. This knowledge may be unknown prior to executing an MA but it can be mined as the search for the global optimum progresses. As more solutions are generated and subsequently stored in the database, the feasibility structure can thus be approximated more accurately. As an integral part, a new paradigm of incorporating the classification-rather than the regression-into the framework of MAs is introduced, allowing the MAs to estimate the feasibility boundary such that effective assessments of whether or not the candidate solutions should experience local refinements can be made. This eventually helps preventing the unnecessary refinements and consequently reducing the number of function evaluations required to reach the global optimum.
Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Yew-Soon Ong
IEEE Trans. Evol. Comput.2
2009 A GA-SVM Feature Selection Model Based on High Performance Computing Techniques
abstract
Supervised learning is well-known and widely applied in many domains including bioinformatics, cheminformatics and financial forecasting. However, the interference from irrelevant features may lead to the poor accuracy of classifiers. As a popular feature selection model, GA-SVM is desirable in many of those cases to filter out irrelevant features and improve the learning performance subsequently. However, the high computational cost strongly discourages the application of GA-SVM in large-scale datasets. In this paper, an HPC-enabled GA-SVM (HGA-SVM) is proposed by integrating data parallelization, multithreading and heuristic techniques with the ultimate goal of robustness and low computational cost. Our proposed model is comprised of four improvement strategies: 1) GA parallelization, 2) SVM parallelization, 3) neighbor search and 4) evaluation caching. All the four strategies improve various aspects of the feature selection model and contribute collectively towards higher computational throughput.
Tianyou Zhang, Xiuju Fu, Rick Siow Mong Goh, Chee Keong Kwoh 0001, Gary Kee Khoon Lee
SMC4
2009 A core-attachment based method to detect protein complexes in PPI networks
abstract
BACKGROUND: How to detect protein complexes is an important and challenging task in post genomic era. As the increasing amount of protein-protein interaction (PPI) data are available, we are able to identify protein complexes from PPI networks. However, most of current studies detect protein complexes based solely on the observation that dense regions in PPI networks may correspond to protein complexes, but fail to consider the inherent organization within protein complexes. RESULTS: To provide insights into the organization of protein complexes, this paper presents a novel core-attachment based method (COACH) which detects protein complexes in two stages. It first detects protein-complex cores as the "hearts" of protein complexes and then includes attachments into these cores to form biologically meaningful structures. We evaluate and analyze our predicted protein complexes from two aspects. First, we perform a comprehensive comparison between our proposed method and existing techniques by comparing the predicted complexes against benchmark complexes. Second, we also validate the core-attachment structures using various biological evidence and knowledge. CONCLUSION: Our proposed COACH method has been applied on two different yeast PPI networks and the experimental results show that COACH performs significantly better than the state-of-the-art techniques. In addition, the identified complexes with core-attachment structures are demonstrated to match very well with existing biological knowledge and thus provide more insights for future biological study.
Min Wu 0008, Xiaoli Li 0001, Chee Keong Kwoh 0001, See-Kiong Ng
BMC Bioinform.3
2009 Brief Overview of Bioinformatics Activities in Singapore
abstract
10.1371/journal.pcbi.1000508
Frank Eisenhaber, Chee Keong Kwoh 0001, See-Kiong Ng, Wing-Kin Sung, Limsoon Wong
PLoS Comput. Biol.2
2008 A study on constrained MA using GA and SQP: Analytical vs. finite-difference gradients
abstract
Many deterministic algorithms in the context of constrained optimization require the first-order derivatives, or the gradient vectors, of the objective and constraint functions to determine the next feasible direction along which the search should progress. Although the second-order derivatives, or the Hessian matrices, are also required by some methods such as the sequential quadratic programming (SQP), their values can be approximated based on the first-order information, making the gradients central to the deterministic algorithms for solving constrained optimization problems. In this paper, two ways of obtaining the gradients are compared under the framework of the simple memetic algorithm (MA) employing genetic algorithm (GA) and SQP. Despite the simplicity and straightforwardness of the finite-difference gradients, faster convergence rate can be achieved when the analytical gradients can be made available. The savings on the number of function evaluations as well as the amount of time taken to solve some benchmark problems are presented along with some discussions.
Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Yew-Soon Ong, Meng-Hiot Lim
IEEE Congress on Evolutionary Computation2
2008 Using classification for constrained memetic algorithm: A new paradigm
abstract
Regression has been successfully combined with the memetic algorithm (MA) for constructing surrogate models. It is essentially an attempt to approximate the objective or constraint landscape of a constrained optimization problem. Classification, on the other hand, has probably never been thought of being of any assistance to the MA. In fact, it can be used to approximate the feasibility boundary by means of some decision functions. The search effort can thus be focussed on the nearby region, recalling that many constrained optimization problems have their optimal solutions situated on the boundaries. This simply means that only potential individuals will undergo local refinements, reducing the number of function evaluations and accelerating the identification of the global optimum. Presented in this paper is a new approach that combines the support vector machine (SVM) with the MA to achieve this purpose.
Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Yew-Soon Ong
SMC2
2008 CBESW: Sequence Alignment on the Playstation 3
abstract
BACKGROUND: The exponential growth of available biological data has caused bioinformatics to be rapidly moving towards a data-intensive, computational science. As a result, the computational power needed by bioinformatics applications is growing exponentially as well. The recent emergence of accelerator technologies has made it possible to achieve an excellent improvement in execution time for many bioinformatics applications, compared to current general-purpose platforms. In this paper, we demonstrate how the PlayStation 3, powered by the Cell Broadband Engine, can be used as a computational platform to accelerate the Smith-Waterman algorithm. RESULTS: For large datasets, our implementation on the PlayStation 3 provides a significant improvement in running time compared to other implementations such as SSEARCH, Striped Smith-Waterman and CUDA. Our implementation achieves a peak performance of up to 3,646 MCUPS. CONCLUSION: The results from our experiments demonstrate that the PlayStation 3 console can be used as an efficient low cost computational platform for high performance sequence alignment applications.
Adrianto Wirawan, Chee Keong Kwoh 0001, Nim Tri Hieu, Bertil Schmidt
BMC Bioinform.2
2008 Hotspot Hunter: a computational system for large-scale screening and selection of candidate immunological hotspots in pathogen proteomes
abstract
BACKGROUND: T-cell epitopes that promiscuously bind to multiple alleles of a human leukocyte antigen (HLA) supertype are prime targets for development of vaccines and immunotherapies because they are relevant to a large proportion of the human population. The presence of clusters of promiscuous T-cell epitopes, immunological hotspots, has been observed in several antigens. These clusters may be exploited to facilitate the development of epitope-based vaccines by selecting a small number of hotspots that can elicit all of the required T-cell activation functions. Given the large size of pathogen proteomes, including of variant strains, computational tools are necessary for automated screening and selection of immunological hotspots. RESULTS: Hotspot Hunter is a web-based computational system for large-scale screening and selection of candidate immunological hotspots in pathogen proteomes through analysis of antigenic diversity. It allows screening and selection of hotspots specific to four common HLA supertypes, namely HLA class I A2, A3, B7 and class II DR. The system uses Artificial Neural Network and Support Vector Machine methods as predictive engines. Soft computing principles were employed to integrate the prediction results produced by both methods for robust prediction performance. Experimental validation of the predictions showed that Hotspot Hunter can successfully identify majority of the real hotspots. Users can predict hotspots from a single protein sequence, or from a set of aligned protein sequences representing pathogen proteome. The latter feature provides a global view of the localizations of the hotspots in the proteome set, enabling analysis of antigenic diversity and shift of hotspots across protein variants. The system also allows the integration of prediction results of the four supertypes for identification of hotspots common across multiple supertypes. The target selection feature of the system shortlists candidate peptide hotspots for the formulation of an epitope-based vaccine that could be effective against multiple variants of the pathogen and applicable to a large proportion of the human population. CONCLUSION: Hotspot Hunter is publicly accessible at http://antigen.i2r.a-star.edu.sg/hh/. It is a new generation computational tool aiding in epitope-based vaccine design.
Guanglan Zhang, Asif M. Khan, Kellathur N. Srinivasan, A. T. Heiny, Kenneth X. Lee, Chee Keong Kwoh 0001, J. Thomas August, Vladimir Brusic
BMC Bioinform.6
2007 Semi-supervised Learning of the Hidden Vector State Model for Protein-Protein Interactions Extraction
abstract
A major challenge in text mining for biology and biomedicine is automatically extracting protein-protein interactions from the vast amount of biological literature since most knowledge about them still hides in biological publications. Existing approaches can be broadly categorized as rule-based or statistical-based. Rule-based approaches require heavy manual efforts. On the other hand, statistical-based approaches require large-scale, richly annotated corpora in order to reliably estimate model parameters. This is normally difficult to obtain in practical applications. The hidden vector state (HVS) model, an extension of the basic discrete Markov model, has been successfully applied to extract protein-protein interactions. In this paper, we propose a novel approach to train the HVS model on both annotated and un-annotated corpus. Sentences selection algorithm is designed to utilize the semantic parsing results of the un-annotated corpus generated by the HVS model. Experimental results show that the performance of the initial HVS model trained on a small amount of the annotated data can be improved by employing this approach
Yulan He 0001, Chee Keong Kwoh 0001
CIDM3
2007 Semi-supervised learning of the hidden vector state model for extracting protein-protein interactions
Yulan He 0001, Chee Keong Kwoh 0001
Artif. Intell. Medicine3
2006 Functional Prediction of Snake Neurotoxins
abstract
Snake neurotoxins are important experimental tool in pharmacological research. Over the years, the number of snake neurotoxin sequences identified is increasing at a very fast pace. However, only a small portion of them are experimentally characterized from more than 200,000 variants estimated to exist in nature. In this paper, we report a systematic functional analysis on snake neurotoxins using a statistical machine learning method - nearest neighbour approach for functional prediction together with a set of rules. Based on this method we built a highly accurate functional prediction tool for putative annotation for snake neurotoxins
Seng Hong Seah, Chee Keong Kwoh 0001, Vladimir Brusic, Meena Kishore Sakharkar, Geok See Ng
ICARCV2
2006 Extreme Learning Machine for Predicting HLA-Peptide Binding
Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Yew-Soon Ong, Guanglan Zhang, Vladimir Brusic
ISNN (2)2
2005 3D Posture Reconstruction and Human Animation from 2D Feature Points
abstract
Abstract An optimal approach is proposed in this paper for posture reconstruction and human animation from 2D feature points extracted from the monocular images containing human motions. Biomechanical constraints are encoded in every joint of the adopted 3D skeletal human model to make sure that each state of the joints represents a physically valid posture. Size of the human model is adjusted to be consistent with the human figure represented by feature points. Energy Function is defined to represent the residuals between the extracted 2D feature points and the corresponding features resulted from projection of the 3D human model. Local Adjustment and Global Adjustment procedures are proposed to place the joints and body segments into proper locations and orientations in 3D space to create the posture with the minimum value of Energy Function. To find the optimal solution of the ill‐posed recovery problem from 2D to 3D, Genetic Algorithm is employed in the high‐dimensional parameter space by considering all the parameters simultaneously. Smooth and continuous changes between consecutive frames are considered in development of the human animation procedure. The proposed approach produces optimal reconstruction results of any possible human postures and movements. It is different from classical kinematics and dynamics formulations, and is an attempt to bridge the gap between computer vision and computer animation in human motion study.
Ling Li 0006, Chee Keong Kwoh 0001
Comput. Graph. Forum3
2004 Reconstructing Boolean networks from noisy gene expression data
abstract
In recent years, a lot of interests have been given to simulate gene regulatory networks (GRNs), especially the architectures of them. Boolean networks (BLNs) are a good choice to obtain the architectures of GRNs when the accessible data sets are limited. Various algorithms have been introduced to reconstruct Boolean networks from gene expression profiles, which are always noisy. However, there are still few dedicated endeavors given to noise problems in learning BLNs. In this paper, we introduce a novel way of sifting noises from gene expression data. The noises cause indefinite states in the learned BLNs, but the correct BLNs could be obtained further with the incompletely specified Karnaugh maps. The experiments on both synthetic and yeast gene expression data show that the method can detect noises and reconstruct the original models in some cases.
Zheng Yun, Chee Keong Kwoh 0001
ICARCV2
2001 Convex object based volume visualization: a formal proof and example
Zou Qingsong, Chee Keong Kwoh 0001, Wan Sing Ng
Comput. Graph.2
1999 Tessellated Surface Reconstruction from 2D Contours
Chee Fatt Chan, Chee Keong Kwoh 0001, Ming Yeong Teo, Wan Sing Ng
MICCAI2
1998 Probabilistic reasoning and multiple-expert methodology for correlated objective data
Chee Keong Kwoh 0001, Duncan Fyfe Gillies
Artif. Intell. Eng.1
1997 Choice of error cost function for training unobservable nodes in Bayesian networks
abstract
In the construction of a Bayesian network from observed data, the fundamental assumption that the variables starting from the same parent are conditionally independent can be met by introduction of hidden node (C.K. Kwoh and D.F. Gillies, 1994). We show that the conditional probability matrices for the hidden node for a triplet, linking three observed nodes, can be determined by the gradient descent method. As in all operational research problems, the quality of the result depends on the ability to locate a feasible solution for the conditional probabilities. C.K. Kwoh and D.F. Gillies (1995) presented a paper in which they detailed the methodologies for estimating the initial values of unobservable variables in Bayesian networks. We present the concept of determining the best conditional matrices as an estimation problem. The discrepancies between the observed and predicted values are mapped into a monotonic function where its gradients are used for adjusting the parameters to be estimated. We present our investigation of choosing among various popular error cost functions for training the networks with hidden nodes and determined that both cross entropy and sum of squared error cost functions work equally well for our implementation.
Chee Keong Kwoh 0001, Duncan Fyfe Gillies
KES (2)1
1996 Using Hidden Nodes in Bayesian Networks
Chee Keong Kwoh 0001, Duncan Fyfe Gillies
Artif. Intell.1
1995 Estimating the initial values of unobservable variables in visual probabilistic networks
Chee Keong Kwoh 0001, Duncan Fyfe Gillies
CAIP1