VLDB 2026 Research / reviewers in the wild / expert
Ka-Chun Wong
dblp:45/7183
· DBLP profile ↗
140ranked-venue papers
15as first author
88since 2021 · last 2026
0000-0001-6062-733XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 70 · 10 first-author · 43 since 2021Artificial intelligence and machine learning · 58 · 5 first-author · 38 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Relation model-assisted multi-region evolutionary algorithm for expensive constrained optimization
Yuxi Huang 0011, Genghui Li, Laizhong Cui, Wangjun Chen, Zhicai Zhu, Qiuzhen Lin, Ka-Chun Wong |
Expert Syst. Appl. | 9 |
| 2026 | Expensive multiobjective immune algorithm using a novel differential evolution in objective space
Wu Lin, Daxin Zhu, Anhui Tan, Ka-Chun Wong, Qiuzhen Lin |
Expert Syst. Appl. | 5 |
| 2026 | Evolutionary multi-task robust architecture search for network intrusion detection
Yeming Yang, Ka-Chun Wong, Qiuzhen Lin, Jianping Luo, Jianqiang Li 0001 |
Expert Syst. Appl. | 3 |
| 2026 | A historical search-guided evolutionary framework for dynamic multiobjective optimizationabstractIn recent years, a number of dynamic multiobjective evolutionary algorithms (DMOEAs) have been proposed for tackling dynamic multiobjective optimization problems (DMOPs). Most of DMOEAs adopt learning methods to extract search experiences from past environments, trying to predict a promising initial population in new environments. However, they often ignore the use of search experiences to guide the evolutionary trajectory in new environments, which is also important to accelerate their convergence. Thus, this paper proposes a historical search-guided evolutionary (HSGE) framework for tackling DMOPs, which designs a neural network-based pattern learning (NNPL) strategy and a historical direction-guided evolutionary (HDGE) strategy. First, the NNPL strategy trains a neural network to effectively extract search experiences from historical environments. Then, based on these search experiences, the HDGE strategy is designed to steer the evolutionary direction of population, aiming to speed up its convergence in new environments. After embedding four dynamic response mechanisms into the HSGE framework, the corresponding DMOEAs have shown superior performance over the original DMOEAs on most test DMOPs. Moreover, the experimental results also validate the advantages of HSGE over two state-of-the-art optimization frameworks for tackling DMOPs. Qingling Zhu, Ping Guo 0007, Qiuzhen Lin, Ka-Chun Wong, Jianqiang Li 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Federated surrogate-assisted evolutionary framework for distributed data-driven multiobjective optimization
Qianhui Ding, Wu Lin, Yinglan Feng, Lijia Ma, Ka-Chun Wong, Qiuzhen Lin, Jianqiang Li 0001 |
Inf. Sci. | 5 |
| 2026 | METRON: Metabolic Dynamic Perception Kolmogorov-Arnold Network for Biological Age EstimationabstractBiological age is a more direct reflection of physiological status than chronological age, serving as a vital measure to evaluate health risks and aging interventions. While steroid metabolomics offers rich information for exploring aging mechanisms, the complex and nonlinear interactions within metabolic networks remain challenging in modeling. Here, we propose and describe METRON as a deep learning framework to predict biological ages from steroid metabolomics. Specifically, a Metabolite Interaction Perception Module (MIPM) is proposed to capture the interactions. Subsequently, a Group-Rational Kolmogorov-Arnold Network is also integrated to capture intricate dependencies and enhance the representation capability. We demonstrate that METRON achieves promising performance as compared to other machine learning and deep learning methods. Beyond performance, METRON offers interpretability by recovering the established markers such as Dehydroepiandrosterone (DHEA) and identifying 17-hydroxyprogesterone (17-OH-P4) as the key signature linked to hypothalamic-pituitary-adrenal axis dynamics. These results support the capacity of METRON not only to estimate biological age but also to uncover underappreciated metabolic drivers behind aging. Zhongshen Li, Jixiang Yu, Shen You, Hao Liu 0072, Luyang Cai, Yuxuan Deng, Leyi Wei, Junkai Ji, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Comput. Biol. Bioinform. | 11 |
| 2026 | UniBreak: A Unified Evolutionary Token-Level Jailbreaking Framework for Large Language ModelsabstractLarge Language Models (LLMs) demonstrate promising capabilities in natural language understanding and reasoning with enormous parameter spaces and vast amounts of training data. These attributes have facilitated their deployment into diverse application domains. However, the underlying parameters implicitly assume decision-making boundaries, resulting in a significant number of decision spaces not covered by training data. This makes them susceptible to adversarial manipulations through carefully crafted inputs. To illuminate the vulnerabilities of LLMs, we propose a unified token-level jailbreaking attack that makes victim models generate responses for potentially harmful queries. Specifically, we propose an evolutionary algorithm to evolve perturbation sets, utilizing gradient-based and crossover-based operators to enhance performance under multiple scenarios. Furthermore, we develop a repository for reusing past perturbations and conduct an analysis of token sensitivity within LLMs, facilitating zero-shot attacks with convergence. Extensive benchmark experiments validate the effectiveness of our method on three different models, achieving increases of 62.36%, 57.89%, and 64.81% in attack success rates compared to baseline methods under white-box scenarios. In addition, the evaluation experiments demonstrate our method is effective for multiple scenarios and different size of models. This research reveals vulnerabilities of LLMs, provides theoretical foundations for developing more robust defense strategies, and contributes to building more reliable AI systems. Shen You, Wei Jiang 0016, Hefei Mei, Danei Gong, Zhongshen Li, Jixiang Yu, Junkai Ji, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Evol. Comput. | 10 |
| 2026 | SOPA: Sensitivity-Oriented Poisoning Attack for Self-Supervised Graph Embedding Model via Bilevel Evolutionary OptimizationabstractDespite the popularity of graph neural networks, perturbed graph data is still a serious threat towards its inherent vulnerabilities. Adversarial examples can still easily manipulate the output of graph neural networks across various attack scenarios. Meanwhile, attacks on graph networks also appear to be crucial, as it can help model designers enhance the robustness of their models. In this study, we propose a sensitivity-oriented poisoning attack for self-supervised graph embedding models through bilevel optimization, which employs different optimization methods at each level. In addition, in order to improve attack effectiveness, we analyze graph structure to identify sensitive nodes and edges that guide attack directions, combining gradient-based and query-based methods to target both edge connections and node attributes. Besides, according to the defects of existing graph masked auto-encoders models, we design the feature sensitivity and feature variance to reduce the feature differentiability, which impairs the performance of the downstream model. Ablation studies validate our operator is effective on three citation datasets. And benchmark-based experiments support the effectiveness of our method on three different graph tasks. Specifically, our approach can achieve an average reduction of 3% in the accuracy of node classification compared to existing methods for attacking neural structures alone. For attacking both graph structures and attributes, our model has even achieved an average reduction of 4.5% for the node classification task, outperforming the existing methods. Shen You, Kai Zhou 0001, Zhongshen Li, Kay Chen Tan, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Evol. Comput. | 7 |
| 2026 | Genetic Perturbation Modeling for Human Cell Therapy With BRNETabstractCellular responses to genetic perturbations are prevalent in wide contexts from the fundamental understandings on pathology to the development of clinical therapies and the discovery of novel drug targets. Nonetheless, the substantial amount of possible perturbation combinations renders wet-lab experiments prohibitively expensive and time-consuming. To address it, the BRNET model is proposed for predicting non-linear transcriptional outcomes where multiple perturbations exist. BRNET integrates prior knowledge with advanced embeddings into a non-stacked neural structure to predict transcriptional responses to both individual and multiple genetic perturbations. For unseen scenarios, BRNET also generalizes well under the corresponding perturbations. Experimental results highlight the capabilities of BRNET, demonstrating promising performance as compared to established deep learning models. Luyang Cai, Jixang Yu, Qiuzhen Lin, Licheng Liu, Xiangtao Li, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Synergizing Anti-Cancer Drug Combinations With Dual-View Hypergraph Representation FusionabstractDrug combination therapy plays a vital role in disease treatment, including cancer, as it contributes to treatment efficacy and can alleviate the effect of drug resistance. Although clinical trials and screening may provide valuable information about synergistic drug combinations, they suffer from challenging combinatorial space. Multiple methods are proposed to address those issues. However, they still fail in making full use of global and local triplet context relationships of known synergistic combinations. To this end, a deep learning model which leverages dual view hypergraph representation fusion for synergistic drug combinations identification is proposed, namely DVHSyn. It first extracts the transcriptome features of cancer cell lines and molecular structures of drugs. Subsequently, by modeling the synergistic effect on a hypergraph, DVHSyn simultaneously learns the local and global context of the sample triplets via a hypergraph view and its expanded heterogeneous graph view. Finally, the learned representations of the above two branches are fused selectively to predict synergistic drug combinations. Experiment results demonstrate that DVHSyn surpasses six other competing methods. One case study also reflects that DVHSyn has the potential to predict novel synergistic drug combinations. Overall, our method is effective in identifying synergistic drug combinations and provides new insights for novel drug development. Jixiang Yu, Nanjun Chen, Linlin Cao, Ming Gao 0008, Daizong Liu, Fuzhou Wang, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | TRNAS: A Training-Free Robust Neural Architecture Search
Yeming Yang, Qingling Zhu, Jianping Luo, Ka-Chun Wong, Qiuzhen Lin, Jianqiang Li 0001 |
ICCV | 4 |
| 2025 | Addressing Capture Bias in Tomato Disease Classification with Deep Transfer Learning
Muhammad Toseef, Malik Jahan Khan, Saifur Rahaman, Olutomilayo Olayemi Petinrin, Xiangtao Li, Ka-Chun Wong |
ICONIP (5) | 7 |
| 2025 | DDintensity: Addressing imbalanced drug-drug interaction risk levels using pre-trained deep learning model embeddings
Weidun Xie, Xingjian Chen, Zetian Zheng, Ruoxuan Zhang, Chengbin Peng 0001, Monika Gullerova, Ka-Chun Wong |
Artif. Intell. Medicine | 11 |
| 2025 | cis-positional information in regulatory single nucleotide variation prioritizationabstractAbstract Duttke et al. have proved and generalized past observations on the positional preferences of regulatory genomics at multiple functional levels in July 2024 [1]. However, the explicit open-box distribution learning on those positional preferences are under-explored in the existing gene regulation methods including deep learning. Contributing towards such directions, we propose to develop regulatory positional distribution models. The positional distribution models can capture the spatial features of gene transcription which can improve different downstream applications such as rSNV prioritization. Experiments have been conducted to substantiate its claim in deleterious rSNV predictions of ClinVar. In particular, we have collected the ClinVar dataset (i.e., ‘variant_summary.txt.gz’ on 2024-09-18) and retrieved all deleterious SNVs (i.e., labelled as ‘Pathogenic’ and ‘Likely pathogenic’ in the ‘ClinicalSigificance’ column) around all human TSS locations from Ensembl (i.e., Ensembl Genes 112). In particular, it was surprising that, although CADD is already an ensemble approach built upon different state-of-the-arts methods [2], CADD can still be improved with statistical significance after cis-positional information has been incorporated across different situations where evolutionary conservation signals (PhastCons and PhyloP) have been integrated. Based on the results, we propose two future research directions. The first direction is to examine different statistical distributions for position-aware gene regulation modelling while the second direction is to incorporate those distributions into different downstream applications such as eQTL analysis and deleterious rSNV predictions. The outcomes will have broad implications across different downstream gene regulation modelling studies. [1] Duttke S.H. et al. ‘Position-dependent function of human sequence-specific transcription factors.’ Nature 2024;631:891–898. [2] Rentzsch, P. et al. Nucleic acids research 2019:47(D1):D886–D894. Ka-Chun Wong, Zhongyu Yao, Weidun Xie, Zhongshen Li, Tianchi Lu, Cho Ling |
Briefings Bioinform. | 1 |
| 2025 | EnrichRBP: an automated and interpretable computational platform for predicting and analysing RNA-binding protein eventsabstractMOTIVATION: Predicting RNA-binding proteins (RBPs) is central to understanding post-transcriptional regulatory mechanisms. Here, we introduce EnrichRBP, an automated and interpretable computational platform specifically designed for the comprehensive analysis of RBP interactions with RNA. RESULTS: EnrichRBP is a web service that enables researchers to develop original deep learning and machine learning architectures to explore the complex dynamics of RBPs. The platform supports 70 deep learning algorithms, covering feature representation, selection, model training, comparison, optimization, and evaluation, all integrated within an automated pipeline. EnrichRBP is adept at providing comprehensive visualizations, enhancing model interpretability, and facilitating the discovery of functionally significant sequence regions crucial for RBP interactions. In addition, EnrichRBP supports base-level functional annotation tasks, offering explanations and graphical visualizations that confirm the reliability of the predicted RNA-binding sites. Leveraging high-performance computing, EnrichRBP provides ultra-fast predictions ranging from seconds to hours, applicable to both pre-trained and custom model scenarios, thus proving its utility in real-world applications. Case studies highlight that EnrichRBP provides robust and interpretable predictions, demonstrating the power of deep learning in the functional analysis of RBP interactions. Finally, EnrichRBP aims to enhance the reproducibility of computational method analyses for RBP sequences, as well as reduce the programming and hardware requirements for biologists, thereby offering meaningful functional insights. AVAILABILITY AND IMPLEMENTATION: EnrichRBP is available at https://airbp.aibio-lab.com/. The source code is available at https://github.com/wangyb97/EnrichRBP, and detailed online documentation can be found at https://enrichrbp.readthedocs.io/en/latest/. Yujian Huang, Jian Zhang 0020, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 7 |
| 2025 | Federated Intrusion Detection System With Cost-Sensitive Learning for Internet of ThingsabstractNetwork Intrusion Detection System (NIDS) has become more important as a large number of diverse devices connect to the Internet of Things (IoT). Generally, training an effective NIDS requires a large amount of high-quality and centralized attack data. However, in real-world scenarios, it is difficult to centralize the distributed data for training NIDS in the IoT due to the privacy concerns and data format heterogeneity. To solve this problem, a novel NIDS combining federated learning and cost-sensitive learning is proposed, named FIDS-CL. Specifically, federated learning with dynamic weights aggregation tackles the problem of non-clusterable data, where multiple clients collaboratively enhance the overall performance while dynamically aggregating weights to maximize the retention of high-performing client models. Moreover, cost-sensitive learning is employed to alleviate the problem of class imbalance in NIDS by dynamically adjusting the gradient descent weights of different classes in the loss function, thereby emphasizing the importance of minority classes. Therefore, our method can effectively handle data imbalance while safeguarding data privacy of clients, which is more effective to detect network intrusions. The experiments conducted across various scenarios validate the superior detection capabilities and computational efficiency of FIDS-CL when compared to other state-of-the-art NIDSs. Qiuzhen Lin, Shaifeng Zheng, Junkai Ji, Ka-Chun Wong, Jianqiang Li 0001, Carlos A. Coello Coello |
IEEE Internet Things J. | 5 |
| 2025 | Elucidating spatiotemporal chromatin dynamics with multi-stage differential variations from Hi-CabstractHigh-throughput sequencing such as Hi-C captures spatiotemporal chromatin interactions, revealing the intricate interplays within transcriptional regulation and chromatin dynamics during cellular reprogramming and developmental processes. However, the forecast on chromatin dynamics in successive developmental stages remains challenging due to the inherent complexity of spatial and temporal patterns in Hi-C data across developmental stages. Towards such a direction, we present StarMie, a deep learning framework that integrates the spatiotemporal-aware module and the multi-stage differential variation module to predict high-throughput chromatin interactions in next developmental stages. Our comprehensive evaluation demonstrates that StarMie outperforms existing methods and sufficiently captures discriminative spatial and temporal dependencies as well as inter-stage-level variations of chromatin interactions. Moreover, the dual importance of both spatial and temporal information in Hi-C data is observed in parameter analysis. Ablation studies also confirm the essential role of each component in StarMie. Furthermore, five cross-species case studies support StarMie's cross-species generalizability and its capability to extract universal chromatin interaction patterns in different developmental stages. In-depth analysis demonstrates that StarMie uncovers conserved genomic logic in cardiac development and disease. Overall, this work paves a new approach for exploring genome reprogramming and development through predictive modeling of Hi-C dynamics. Zhongshen Li, Jixiang Yu, Shen You, Leyi Wei, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
Knowl. Based Syst. | 7 |
| 2025 | Personalized federated learning with multiple classifier aggregation
Shaifeng Zheng, Qingling Zhu, Qiuzhen Lin, Songbai Liu, Ka-Chun Wong, Jianqiang Li 0001 |
Knowl. Based Syst. | 5 |
| 2025 | Editorial: 2020 India international congress on computational intelligence
Suash Deb, Ka-Chun Wong, Thomas Hanne |
Neural Comput. Appl. | 2 |
| 2025 | Dynamic Graph Representation Learning for Spatio-Temporal Neuroimaging AnalysisabstractNeuroimaging analysis aims to reveal the information-processing mechanisms of the human brain in a noninvasive manner. In the past, graph neural networks (GNNs) have shown promise in capturing the non-Euclidean structure of brain networks. However, existing neuroimaging studies focused primarily on spatial functional connectivity, despite temporal dynamics in complex brain networks. To address this gap, we propose a spatio-temporal interactive graph representation framework (STIGR) for dynamic neuroimaging analysis that encompasses different aspects from classification and regression tasks to interpretation tasks. STIGR leverages a dynamic adaptive-neighbor graph convolution network to capture the interrelationships between spatial and temporal dynamics. To address the limited global scope in graph convolutions, a self-attention module based on Transformers is introduced to extract long-term dependencies. Contrastive learning is used to adaptively contrast similarities between adjacent scanning windows, modeling cross-temporal correlations in dynamic graphs. Extensive experiments on six public neuroimaging datasets demonstrate the competitive performance of STIGR across different platforms, achieving state-of-the-art results in classification and regression tasks. The proposed framework enables the detection of remarkable temporal association patterns between regions of interest based on sequential neuroimaging signals, offering medical professionals a versatile and interpretable tool for exploring task-specific neurological patterns. Our codes and models are available at https://github.com/77YQ77/STIGR/. Rui Liu 0038, Yao Hu 0001, Jibin Wu, Ka-Chun Wong, Zhi-an Huang, Kay Chen Tan |
IEEE Trans. Cybern. | 4 |
| 2025 | When East Meets West: Cross-Domain Drug Interaction Annotations With Large Language Models and Bidirectional Neural NetworksabstractDrug combination therapy is a promising strategy for managing complex and co-existing diseases. However, drug-drug interactions (DDIs) can result in unexpected adverse effects, making it crucial to understand such interactions to prevent adverse drug reactions and develop new therapeutic strategies. Current DDI annotation methods heavily rely on atom-level graph structural features, overlooking valuable drug contextual representations within medical literature. Additionally, these methods are typically designed for a specific task, limiting their scalability to diverse medical scenarios. To address these limitations, we propose TEmbed-DDI, a novel framework that leverages contextual representations and pre-trained large language model embeddings to enhance feature extraction for DDI annotations. Specifically, we retrieve meaningful contextual texts for each drug to enrich semantic features and adopt pre-trained large language model embeddings to capture rich features from these long-range contextual representations. TEmbed-DDI is the first framework to incorporate LLM-powered embeddings for medical interaction annotations. Furthermore, a bidirectional neural network is integrated into TEmbed-DDI for the integrative Western and traditional Chinese medicine DDI annotation tasks. Comparative results demonstrate that TEmbed-DDI achieves state-of-the-art performance, with the highest AUC scores of 0.992 and 0.95 on the Western CHCH and DEEP interaction annotation benchmarks. Even for the newly constructed Traditional Chinese Medicine (TCM) DDI annotation benchmark, TEmbed-DDI consistently exhibits outstanding generalization capability, achieving an AUC of 0.956. Moreover, case studies further validate TEmbed-DDI's capability to annotate previously unknown interactions. These findings suggest that TEmbed-DDI can serve as a valuable tool in annotating previously unknown drug combinations for real-world applications, facilitating the development of efficacious therapies. Furthermore, as the first framework combining traditional Chinese medicine into DDI annotation tasks, its adaptability highlights the potential in supporting cross-domain medical research. TEmbed-DDI's design principles can inspire the development of flexible LLM-powered frameworks for drug combination discovery in the future. Ruoxuan Zhang, Weidun Xie, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Unsupervised Gene-Cell Collective Representation Learning with Optimal TransportabstractCell type identification plays a vital role in single-cell RNA sequencing (scRNA-seq) data analysis. Although many deep embedded methods to cluster scRNA-seq data have been proposed, they still fail in elucidating the intrinsic properties of cells and genes. Here, we present a novel end-to-end deep graph clustering model for single-cell transcriptomics data based on unsupervised Gene-Cell Collective representation learning and Optimal Transport (scGCOT) which integrates both cell and gene correlations. Specifically, scGCOT learns the latent embedding of cells and genes simultaneously and reconstructs the cell graph, the gene graph, and the gene expression count matrix. A zero-inflated negative binomial (ZINB) model is estimated via the reconstructed count matrix to capture the essential properties of scRNA-seq data. By leveraging the optimal transport-based joint representation alignment, scGCOT learns the clustering process and the latent representations through a mutually supervised self optimization strategy. Extensive experiments with 14 competing methods on 15 real scRNA-seq datasets demonstrate the competitive edges of scGCOT. Jixiang Yu, Nanjun Chen, Ming Gao 0008, Xiangtao Li, Ka-Chun Wong |
AAAI | 5 |
| 2024 | scPER2P: Parameter-Efficient Single-Cell LLM for Translated Proteome Profiles
Xingjian Chen, Zetian Zheng, Weidun Xie, Fuzhou Wang, Ka-Chun Wong |
ICONIP (5) | 7 |
| 2024 | A versatile informative diffusion model for single-cell ATAC-seq data generation and analysisabstractThe rapid advancement of single-cell ATAC sequencing (scATAC-seq) technologies holds great promise for investigating the heterogeneity of epigenetic landscapes at the cellular level. The amplification process in scATAC-seq experiments often introduces noise due to dropout events, which results in extreme sparsity that hinders accurate analysis. Consequently, there is a significant demand for the generation of high-quality scATAC-seq data in silico. Furthermore, current methodologies are typically task-specific, lacking a versatile framework capable of handling multiple tasks within a single model. In this work, we propose ATAC-Diff, a versatile framework, which is based on a diffusion model conditioned on the latent auxiliary variables to adapt for various tasks. ATAC-Diff is the first diffusion model for the scATAC-seq data generation and analysis, composed of auxiliary modules encoding the latent high-level variables to enable the model to learn the semantic information to sample high-quality data. Gaussian Mixture Model (GMM) as the latent prior and auxiliary decoder, the yield variables reserve the refined genomic information beneficial for downstream analyses. Another innovation is the incorporation of mutual information between observed and hidden variables as a regularization term to prevent the model from decoupling from latent variables. Through extensive experiments, we demonstrate that ATAC-Diff achieves high performance in both generation and analysis tasks, outperforming state-of-the-art models. Zunpeng Liu, Ka-Chun Wong, Manolis Kellis |
NeurIPS | 5 |
| 2024 | TP-LMMSG: a peptide prediction graph neural network incorporating flexible amino acid property representationabstractBioactive peptide therapeutics has been a long-standing research topic. Notably, the antimicrobial peptides (AMPs) have been extensively studied for its therapeutic potential. Meanwhile, the demand for annotating other therapeutic peptides, such as antiviral peptides (AVPs) and anticancer peptides (ACPs), also witnessed an increase in recent years. However, we conceive that the structure of peptide chains and the intrinsic information between the amino acids is not fully investigated among the existing protocols. Therefore, we develop a new graph deep learning model, namely TP-LMMSG, which offers lightweight and easy-to-deploy advantages while improving the annotation performance in a generalizable manner. The results indicate that our model can accurately predict the properties of different peptides. The model surpasses the other state-of-the-art models on AMP, AVP and ACP prediction across multiple experimental validated datasets. Moreover, TP-LMMSG also addresses the challenges of time-consuming pre-processing in graph neural network frameworks. With its flexibility in integrating heterogeneous peptide features, our model can provide substantial impacts on the screening and discovery of therapeutic peptides. The source code is available at https://github.com/NanjunChen37/TP_LMMSG. Nanjun Chen, Jixiang Yu, Fuzhou Wang, Xiangtao Li, Ka-Chun Wong |
Briefings Bioinform. | 6 |
| 2024 | TransPTM: a transformer-based model for non-histone acetylation site predictionabstractProtein acetylation is one of the extensively studied post-translational modifications (PTMs) due to its significant roles across a myriad of biological processes. Although many computational tools for acetylation site identification have been developed, there is a lack of benchmark dataset and bespoke predictors for non-histone acetylation site prediction. To address these problems, we have contributed to both dataset creation and predictor benchmark in this study. First, we construct a non-histone acetylation site benchmark dataset, namely NHAC, which includes 11 subsets according to the sequence length ranging from 11 to 61 amino acids. There are totally 886 positive samples and 4707 negative samples for each sequence length. Secondly, we propose TransPTM, a transformer-based neural network model for non-histone acetylation site predication. During the data representation phase, per-residue contextualized embeddings are extracted using ProtT5 (an existing pre-trained protein language model). This is followed by the implementation of a graph neural network framework, which consists of three TransformerConv layers for feature extraction and a multilayer perceptron module for classification. The benchmark results reflect that TransPTM has the competitive performance for non-histone acetylation site prediction over three state-of-the-art tools. It improves our comprehension on the PTM mechanism and provides a theoretical basis for developing drug targets for diseases. Moreover, the created PTM datasets fills the gap in non-histone acetylation site datasets and is beneficial to the related communities. The related source code and data utilized by TransPTM are accessible at https://www.github.com/TransPTM/TransPTM. Lingkuan Meng, Xingjian Chen, Nanjun Chen, Zetian Zheng, Fuzhou Wang, Hongyan Sun, Ka-Chun Wong |
Briefings Bioinform. | 8 |
| 2024 | HE2Gene: image-to-RNA translation via multi-task learning for spatial transcriptomics dataabstractMOTIVATION: Tissue context and molecular profiling are commonly used measures in understanding normal development and disease pathology. In recent years, the development of spatial molecular profiling technologies (e.g. spatial resolved transcriptomics) has enabled the exploration of quantitative links between tissue morphology and gene expression. However, these technologies remain expensive and time-consuming, with subsequent analyses necessitating high-throughput pathological annotations. On the other hand, existing computational tools are limited to predicting only a few dozen to several hundred genes, and the majority of the methods are designed for bulk RNA-seq. RESULTS: In this context, we propose HE2Gene, the first multi-task learning-based method capable of predicting tens of thousands of spot-level gene expressions along with pathological annotations from H&E-stained images. Experimental results demonstrate that HE2Gene is comparable to state-of-the-art methods and generalizes well on an external dataset without the need for re-training. Moreover, HE2Gene preserves the annotated spatial domains and has the potential to identify biomarkers. This capability facilitates cancer diagnosis and broadens its applicability to investigate gene-disease associations. AVAILABILITY AND IMPLEMENTATION: The source code and data information has been deposited at https://github.com/Microbiods/HE2Gene. Xingjian Chen, Jiecong Lin, Weidun Xie, Zetian Zheng, Ka-Chun Wong |
Bioinform. | 7 |
| 2024 | Unraveling spatial domain characterization in spatially resolved transcriptomics with robust graph contrastive clusteringabstractMOTIVATION: Spatial transcriptomics can quantify gene expression and its spatial distribution in tissues, thus revealing molecular mechanisms of cellular interactions underlying tissue heterogeneity, tissue regeneration, and spatially localized disease mechanisms. However, existing spatial clustering methods often fail to exploit the full potential of spatial information, resulting in inaccurate identification of spatial domains. RESULTS: In this article, we develop a deep graph contrastive clustering framework, stDGCC, that accurately uncovers underlying spatial domains via explicitly modeling spatial information and gene expression profiles from spatial transcriptomics data. The stDGCC framework proposes a spatially informed graph node embedding model to preserve the topological information of spots and to learn the informative and discriminative characterization of spatial transcriptomics data through self-supervised contrastive learning. By simultaneously optimizing the contrastive learning loss, reconstruction loss, and Kullback-Leibler divergence loss, stDGCC achieves joint optimization of feature learning and topology structure preservation in an end-to-end manner. We validate the effectiveness of stDGCC on various spatial transcriptomics datasets acquired from different platforms, each with varying spatial resolutions. Our extensive experiments demonstrate the superiority of stDGCC over various state-of-the-art clustering methods in accurately identifying cellular-level biological structures. AVAILABILITY AND IMPLEMENTATION: Code and data are available from https://github.com/TimE9527/stDGCC and https://figshare.com/projects/stDGCC/186525. Yingxi Zhang, Zhuohan Yu, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 3 |
| 2024 | A localized decomposition evolutionary algorithm for imbalanced multi-objective optimization
Yulong Ye, Qiuzhen Lin, Ka-Chun Wong, Jianqiang Li 0001, Zhong Ming 0001, Carlos A. Coello Coello |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Evolving pathway activation from cancer gene expression data using nature-inspired ensemble optimization
Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Expert Syst. Appl. | 4 |
| 2024 | A complementary fused method using GRU and XGBoost models for long-term solar energy hourly forecasting
Yaojian Xu, Shaifeng Zheng, Qingling Zhu, Ka-Chun Wong, Qiuzhen Lin |
Expert Syst. Appl. | 4 |
| 2024 | CAGAN: Constrained neural architecture search for GANs
Yeming Yang, Xinzhi Zhang 0008, Qingling Zhu, Weineng Chen, Ka-Chun Wong, Qiuzhen Lin |
Knowl. Based Syst. | 5 |
| 2024 | Exhaustive Exploitation of Nature-Inspired Computation for Cancer Screening in an Ensemble MannerabstractAccurate screening of cancer types is crucial for effective cancer detection and precise treatment selection. However, the association between gene expression profiles and tumors is often limited to a small number of biomarker genes. While computational methods using nature-inspired algorithms have shown promise in selecting predictive genes, existing techniques are limited by inefficient search and poor generalization across diverse datasets. This study presents a framework termed Evolutionary Optimized Diverse Ensemble Learning (EODE) to improve ensemble learning for cancer classification from gene expression data. The EODE methodology combines an intelligent grey wolf optimization algorithm for selective feature space reduction, guided random injection modeling for ensemble diversity enhancement, and subset model optimization for synergistic classifier combinations. Extensive experiments were conducted across 35 gene expression benchmark datasets encompassing varied cancer types. Results demonstrated that EODE obtained significantly improved screening accuracy over individual and conventionally aggregated models. The integrated optimization of advanced feature selection, directed specialized modeling, and cooperative classifier ensembles helps address key challenges in current nature-inspired approaches. This provides an effective framework for robust and generalized ensemble learning with gene expression biomarkers. Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Core-Periphery Detection Based on Masked Bayesian Nonnegative Matrix FactorizationabstractCore–periphery structure is an essential mesoscale feature in complex networks. Previous researches mostly focus on discriminative approaches, while in this work we propose a generative model called masked Bayesian nonnegative matrix factorization. We build the model using two pair affiliation matrices to indicate core–periphery pair associations and using a mask matrix to highlight connections to core nodes. We propose an approach to infer the model parameters and prove the convergence of variables with our approach. Besides the abilities as traditional approaches, it is able to identify core scores with overlapping core–periphery pairs. We verify the effectiveness of our method using randomly generated networks and real-world networks. Experimental results demonstrate that the proposed method outperforms traditional approaches. Zhonghao Wang 0002, Ru Yuan, Jiaye Fu, Ka-Chun Wong, Chengbin Peng 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Toward Evolutionary Multitask Convolutional Neural Architecture SearchabstractEvolutionary neural architecture search (ENAS) methods have been successfully used to design convolutional neural network (CNN) architectures automatically. These methods have achieved excellent performance in creating a specific neural architecture for a single task but are less efficient for multiple tasks. Existing ENAS frameworks always repeatedly perform the search from scratch for each task, even though these tasks may be solved by similar CNN architectures. This work presents an evolutionary multi-task convolutional neural architecture search (MTNAS) framework to enable efficient architecture searches in multi-task scenarios by incorporating architectural similarities. The proposed MTNAS constructs architectures for different tasks simultaneously by implementing a knowledge-sharing mechanism among multiple search processes. Specifically, promising architectures found in one search process can be transferred and reused to generate high-quality architectures for others. Furthermore, we devise an adaptive strategy to dynamically adjust the frequency of knowledge transfer, aiming to alleviate the potential effect of negative transfer. Extensive experiments demonstrate that MTNAS can outperform state-of-the-art NAS methods or achieve comparable performance in different tasks but with 2× less search cost. Zhenkun Wang 0001, Liang Feng 0001, Songbai Liu, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 5 |
| 2024 | Attention-Like Multimodality Fusion With Data Augmentation for Diagnosis of Mental Disorders Using MRIabstractThe globally rising prevalence of mental disorders leads to shortfalls in timely diagnosis and therapy to reduce patients' suffering. Facing such an urgent public health problem, professional efforts based on symptom criteria are seriously overstretched. Recently, the successful applications of computer-aided diagnosis approaches have provided timely opportunities to relieve the tension in healthcare services. Particularly, multimodal representation learning gains increasing attention thanks to the high temporal and spatial resolution information extracted from neuroimaging fusion. In this work, we propose an efficient multimodality fusion framework to identify multiple mental disorders based on the combination of functional and structural magnetic resonance imaging. A multioutput conditional generative adversarial network (GAN) is developed to address the scarcity of multimodal data for augmentation. Based on the augmented training data, the multiheaded gating fusion model is proposed for classification by extracting the complementary features across different modalities. The experiments demonstrate that the proposed model can achieve robust accuracies of 75.1 ± 1.5 %, 72.9 ± 1.1 %, and 87.2 ± 1.5 % for autism spectrum disorder (ASD), attention deficit/hyperactivity disorder, and schizophrenia, respectively. In addition, the interpretability of our model is expected to enable the identification of remarkable neuropathology diagnostic biomarkers, leading to well-informed therapeutic decisions. Rui Liu 0038, Zhi-an Huang, Yao Hu 0001, Zexuan Zhu 0001, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Spatial-Temporal Co-Attention Learning for Diagnosis of Mental Disorders From Resting-State fMRI DataabstractNeuroimaging techniques have been widely adopted to detect the neurological brain structures and functions of the nervous system. As an effective noninvasive neuroimaging technique, functional magnetic resonance imaging (fMRI) has been extensively used in computer-aided diagnosis (CAD) of mental disorders, e.g., autism spectrum disorder (ASD) and attention deficit/hyperactivity disorder (ADHD). In this study, we propose a spatial-temporal co-attention learning (STCAL) model for diagnosing ASD and ADHD from fMRI data. In particular, a guided co-attention (GCA) module is developed to model the intermodal interactions of spatial and temporal signal patterns. A novel sliding cluster attention module is designed to address global feature dependency of self-attention mechanism in fMRI time series. Comprehensive experimental results demonstrate that our STCAL model can achieve competitive accuracies of 73.0 ± 4.5%, 72.0 ± 3.8%, and 72.5 ± 4.2% on the ABIDE I, ABIDE II, and ADHD-200 datasets, respectively. Moreover, the potential for feature pruning based on the co-attention scores is validated by the simulation experiment. The clinical interpretation analysis of STCAL can allow medical professionals to concentrate on the discriminative regions of interest and key time frames from fMRI data. Rui Liu 0038, Zhi-an Huang, Yao Hu 0001, Zexuan Zhu 0001, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Unsupervised Deep Embedded Fusion Representation of Single-Cell TranscriptomicsabstractCell clustering is a critical step in analyzing single-cell RNA sequencing (scRNA-seq) data, which allows us to characterize the cellular heterogeneity of transcriptional profiling at the single-cell level. Single-cell deep embedded representation models have recently become popular since they can learn feature representation and clustering simultaneously. However, the model still suffers from a variety of significant challenges, including the massive amount of data, pervasive dropout events, and complicated noise patterns in transcriptional profiling. Here, we propose a Single-Cell Deep Embedding Fusion Representation (scDEFR) model, which develop a deep embedded fusion representation to learn fused heterogeneous latent embedding that contains both the transcriptome gene-level information and the cell topology information. We first fuse them layer by layer to obtain compressed representations of intercellular relationships and transcriptome information. After that, the zero-inflated negative binomial model (ZINB)-based decoder is proposed to capture the global probabilistic structure of the data and reconstruct the final gene expression information and cell graph. Finally, by simultaneously integrating the clustering loss, crossentropy loss, ZINB loss, and the cell graph reconstruction loss, scDEFR can optimize clustering performance and learn the latent representation in fused information under a joint mutual supervised strategy. We conducted extensive and comprehensive experiments on 15 single-cell RNA-seq datasets from different sequencing platforms to demonstrate the superiority of scDEFR over a variety of state-of-the-art methods. Yanchi Su, Zhuohan Yu, Ka-Chun Wong, Xiangtao Li |
AAAI | 5 |
| 2023 | MDM: Molecular Diffusion Model for 3D Molecule GenerationabstractMolecule generation, especially generating 3D molecular geometries from scratch (i.e., 3D de novo generation), has become a fundamental task in drug design. Existing diffusion based 3D molecule generation methods could suffer from unsatisfactory performances, especially when generating large molecules. At the same time, the generated molecules lack enough diversity. This paper proposes a novel diffusion model to address those two challenges. First, interatomic relations are not included in molecules' 3D point cloud representations. Thus, it is difficult for existing generative models to capture the potential interatomic forces and abundant local constraints. To tackle this challenge, we propose to augment the potential interatomic forces and further involve dual equivariant encoders to encode interatomic forces of different strengths. Second, existing diffusion-based models essentially shift elements in geometry along the gradient of data density. Such a process lacks enough exploration in the intermediate steps of the Langevin dynamics. To address this issue, we introduce a distributional controlling variable in each diffusion/reverse step to enforce thorough explorations and further improve generation diversity. Extensive experiments on multiple benchmarks demonstrate that the proposed model significantly outperforms existing methods for both unconditional and conditional generation tasks. We also conduct case studies to help understand the physicochemical properties of the generated molecules. The codes are available at https://github.com/tencent-ailab/MDM. Hengtong Zhang, Tingyang Xu, Ka-Chun Wong |
AAAI | 4 |
| 2023 | Deep transfer learning for clinical decision-making based on high-throughput data: comprehensive survey with benchmark resultsabstractThe rapid growth of omics-based data has revolutionized biomedical research and precision medicine, allowing machine learning models to be developed for cutting-edge performance. However, despite the wealth of high-throughput data available, the performance of these models is hindered by the lack of sufficient training data, particularly in clinical research (in vivo experiments). As a result, translating this knowledge into clinical practice, such as predicting drug responses, remains a challenging task. Transfer learning is a promising tool that bridges the gap between data domains by transferring knowledge from the source to the target domain. Researchers have proposed transfer learning to predict clinical outcomes by leveraging pre-clinical data (mouse, zebrafish), highlighting its vast potential. In this work, we present a comprehensive literature review of deep transfer learning methods for health informatics and clinical decision-making, focusing on high-throughput molecular data. Previous reviews mostly covered image-based transfer learning works, while we present a more detailed analysis of transfer learning papers. Furthermore, we evaluated original studies based on different evaluation settings across cross-validations, data splits and model architectures. The result shows that those transfer learning methods have great potential; high-throughput sequencing data and state-of-the-art deep learning models lead to significant insights and conclusions. Additionally, we explored various datasets in transfer learning papers with statistics and visualization. Muhammad Toseef, Olutomilayo Olayemi Petinrin, Fuzhou Wang, Saifur Rahaman, Xiangtao Li, Ka-Chun Wong |
Briefings Bioinform. | 7 |
| 2023 | Automated exploitation of deep learning for cancer patient stratification across multiple typesabstractMOTIVATION: Recent frameworks based on deep learning have been developed to identify cancer subtypes from high-throughput gene expression profiles. Unfortunately, the performance of deep learning is highly dependent on its neural network architectures which are often hand-crafted with expertise in deep neural networks, meanwhile, the optimization and adjustment of the network are usually costly and time consuming. RESULTS: To address such limitations, we proposed a fully automated deep neural architecture search model for diagnosing consensus molecular subtypes from gene expression data (DNAS). The proposed model uses ant colony algorithm, one of the heuristic swarm intelligence algorithms, to search and optimize neural network architecture, and it can automatically find the optimal deep learning model architecture for cancer diagnosis in its search space. We validated DNAS on eight colorectal cancer datasets, achieving the average accuracy of 95.48%, the average specificity of 98.07%, and the average sensitivity of 96.24%, respectively. Without the loss of generality, we investigated the general applicability of DNAS further on other cancer types from different platforms including lung cancer and breast cancer, and DNAS achieved an area under the curve of 95% and 96%, respectively. In addition, we conducted gene ontology enrichment and pathological analysis to reveal interesting insights into cancer subtype identification and characterization across multiple cancer types. AVAILABILITY AND IMPLEMENTATION: The source code and data can be downloaded from https://github.com/userd113/DNAS-main. And the web server of DNAS is publicly accessible at 119.45.145.120:5001. Shijie Fan, Shaochuan Li, Yingwei Zhao, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 6 |
| 2023 | scBGEDA: deep single-cell clustering analysis via a dual denoising autoencoder with bipartite graph ensemble clusteringabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) is an increasingly popular technique for transcriptomic analysis of gene expression at the single-cell level. Cell-type clustering is the first crucial task in the analysis of scRNA-seq data that facilitates accurate identification of cell types and the study of the characteristics of their transcripts. Recently, several computational models based on a deep autoencoder and the ensemble clustering have been developed to analyze scRNA-seq data. However, current deep autoencoders are not sufficient to learn the latent representations of scRNA-seq data, and obtaining consensus partitions from these feature representations remains under-explored. RESULTS: To address this challenge, we propose a single-cell deep clustering model via a dual denoising autoencoder with bipartite graph ensemble clustering called scBGEDA, to identify specific cell populations in single-cell transcriptome profiles. First, a single-cell dual denoising autoencoder network is proposed to project the data into a compressed low-dimensional space and that can learn feature representation via explicit modeling of synergistic optimization of the zero-inflated negative binomial reconstruction loss and denoising reconstruction loss. Then, a bipartite graph ensemble clustering algorithm is designed to exploit the relationships between cells and the learned latent embedded space by means of a graph-based consensus function. Multiple comparison experiments were conducted on 20 scRNA-seq datasets from different sequencing platforms using a variety of clustering metrics. The experimental results indicated that scBGEDA outperforms other state-of-the-art methods on these datasets, and also demonstrated its scalability to large-scale scRNA-seq datasets. Moreover, scBGEDA was able to identify cell-type specific marker genes and provide functional genomic analysis by quantifying the influence of genes on cell clusters, bringing new insights into identifying cell types and characterizing the scRNA-seq data from different perspectives. AVAILABILITY AND IMPLEMENTATION: The source code of scBGEDA is available at https://github.com/wangyh082/scBGEDA. The software and the supporting data can be downloaded from https://figshare.com/articles/software/scBGEDA/19657911. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yunhe Wang 0002, Zhuohan Yu, Shaochuan Li, Chuang Bian, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 6 |
| 2023 | Chromothripsis detection with multiple myeloma patients based on deep graph learningabstractMOTIVATION: Chromothripsis, associated with poor clinical outcomes, is prognostically vital in multiple myeloma. The catastrophic event is reported to be detectable prior to the progression of multiple myeloma. As a result, chromothripsis detection can contribute to risk estimation and early treatment guidelines for multiple myeloma patients. However, manual diagnosis remains the gold standard approach to detect chromothripsis events with the whole-genome sequencing technology to retrieve both copy number variation (CNV) and structural variation data. Meanwhile, CNV data are much easier to obtain than structural variation data. Hence, in order to reduce the reliance on human experts' efforts and structural variation data extraction, it is necessary to establish a reliable and accurate chromothripsis detection method based on CNV data. RESULTS: To address those issues, we propose a method to detect chromothripsis solely based on CNV data. With the help of structure learning, the intrinsic relationship-directed acyclic graph of CNV features is inferred to derive a CNV embedding graph (i.e. CNV-DAG). Subsequently, a neural network based on Graph Transformer, local feature extraction, and non-linear feature interaction, is proposed with the embedding graph as the input to distinguish whether the chromothripsis event occurs. Ablation experiments, clustering, and feature importance analysis are also conducted to enable the proposed model to be explained by capturing mechanistic insights. AVAILABILITY AND IMPLEMENTATION: The source code and data are freely available at https://github.com/luvyfdawnYu/CNV_chromothripsis. Jixiang Yu, Nanjun Chen, Zetian Zheng, Ming Gao 0008, Ka-Chun Wong |
Bioinform. | 6 |
| 2023 | 2019 India International Congress on Computational Intelligence
Suash Deb, Ka-Chun Wong, Thomas Hanne |
Neural Comput. Appl. | 2 |
| 2023 | Hyperspectral Image Denoising via Weighted Multidirectional Low-Rank Tensor RecoveryabstractRecently, low-rank tensor recovery methods based on subspace representation have received increased attention in the field of hyperspectral image (HSI) denoising. Unfortunately, those methods usually analyze the prior structural information within different dimensions indiscriminately, ignoring the differences between modes, leaving substantial room for improvement. In this article, we first consider the low-rank properties in the subspace and prove that the structure correlation across the nonlocal self-similarity mode is much stronger than in the spatial sparsity and spectral correlation modes. On that basis, we introduce a new multidirectional low-rank regularization, in which each mode is assigned a different weight to characterize its contribution to estimating the tensor rank. After that, integrating the proposed regularization with the subspace-based tensor recovery framework, an optimization model for HSI mixed noise removal is developed. The proposed model can be addressed efficiently via the alternating minimization algorithm. Extensive experiments implemented with synthetic and real data demonstrate that the proposed method significantly outperforms other state-of-the-art HSI denoising methods, which clearly indicates the effectiveness of the proposed approach in HSI denoising. Yanchi Su, Ka-Chun Wong, Yi Chang 0001, Xiangtao Li |
IEEE Trans. Cybern. | 3 |
| 2023 | Evolutionary Multitasking for Large-Scale Multiobjective OptimizationabstractEvolutionary transfer optimization (ETO) has been becoming a hot research topic in the field of evolutionary computation, which is based on the fact that knowledge learning and transfer across the related optimization exercises can improve the efficiency of others. However, rare studies employ ETO to solve large-scale multiobjective optimization problems (LMOPs). To fill this research gap, this article proposes a new multitasking ETO algorithm via a powerful transfer learning model to simultaneously solve multiple LMOPs. In particular, inspired by adversarial domain adaptation in transfer learning, a discriminative reconstruction network (DRN) model (containing an encoder, a decoder, and a classifier) is created for each LMOP. At each generation, the DRN is trained by the currently obtained nondominated solutions for all LMOPs via backpropagation with gradient descent. With this well-trained DRN model, the proposed algorithm can transfer the solutions of source LMOPs directly to the target LMOP for assisting its optimization, can evaluate the correlation between the source and target LMOPs to control the transfer of solutions, and can learn a dimensional-reduced Pareto-optimal subspace of the target LMOP to improve the efficiency of transfer optimization in the large-scale search space. Moreover, we propose a real-world multitasking LMOP suite to simulate the training of deep neural networks (DNNs) on multiple different classification tasks. Finally, the effectiveness of the proposed algorithm has been validated in this real-world problem suite and the other two synthetic problem suites. Songbai Liu, Qiuzhen Lin, Liang Feng 0001, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 4 |
| 2023 | Evolutionary Large-Scale Multiobjective Optimization: Benchmarks and AlgorithmsabstractEvolutionary large-scale multiobjective optimization (ELMO) has received increasing attention in recent years. This study has compared various existing optimizers for ELMO on different benchmarks, revealing that both benchmarks and algorithms for ELMO still need significant improvement. Thus, a new test suite and a new optimizer framework are proposed to further promote the research of ELMO. More realistic features are considered in the new benchmarks, such as mixed formulation of objective functions, mixed linkages in variables, and imbalanced contributions of variables to the objectives, which are challenging to the existing optimizers. To better tackle these benchmarks, a variable group-based learning strategy is embedded into the new optimizer framework for ELMO, which significantly improves the quality of reproduction in large-scale search space. The experimental results validate that the designed benchmarks can comprehensively evaluate the performance of existing optimizers for ELMO and the proposed optimizer shows distinct advantages in tackling these benchmarks. Songbai Liu, Qiuzhen Lin, Ka-Chun Wong, Qing Li 0001, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 3 |
| 2023 | An Immune-Inspired Resource Allocation Strategy for Many-Objective OptimizationabstractRecently, a number of resource allocation strategies have been proposed for evolutionary algorithms to efficiently tackle multiobjective optimization problems (MOPs). However, these methods mainly allocate computational resources based on the convergence improvement under the decomposition-based framework, which may become ineffective with the increased number of optimization objectives. To address this problem, this article suggests an immune-inspired resource allocation strategy, which breaks through the decomposition-based framework and can better balance convergence and diversity for many-objective optimization. In our method, the diversity distances of solutions are defined by the Euclidean distances of their projected points on the unit hyperplane. Then, based on the diversity distances, resource allocation is realized by using an immune cloning operator to encourage exploring sparse regions of the search space. Moreover, to provide high-quality solutions in coordination with this immune cloning operator, a novel archive update mechanism is designed. When compared to most well-known resource allocation strategies, our method is advantageous for many-objective optimization. The experimental results also validate the superiority of our method over several state-of-the-art evolutionary algorithms for solving two sets of complicated MOPs having 5 to 15 objectives. Qiuzhen Lin, Zhong Ming 0001, Ka-Chun Wong, Maoguo Gong, Carlos A. Coello Coello |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | ZINB-Based Graph Embedding Autoencoder for Single-Cell RNA-Seq InterpretationsabstractSingle-cell RNA sequencing (scRNA-seq) provides high-throughput information about the genome-wide gene expression levels at the single-cell resolution, bringing a precise understanding on the transcriptome of individual cells. Unfortunately, the rapidly growing scRNA-seq data and the prevalence of dropout events pose substantial challenges for cell type annotation. Here, we propose a single-cell model-based deep graph embedding clustering (scTAG) method, which simultaneously learns cell–cell topology representations and identifies cell clusters based on deep graph convolutional network. scTAG integrates the zero-inflated negative binomial (ZINB) model into a topology adaptive graph convolutional autoencoder to learn the low-dimensional latent representation and adopts Kullback–Leibler (KL) divergence for the clustering tasks. By simultaneously optimizing the clustering loss, ZINB loss, and the cell graph reconstruction loss, scTAG jointly optimizes cluster label assignment and feature learning with the topological structures preserved in an end-to-end manner. Extensive experiments on 16 single-cell RNA-seq datasets from diverse yet representative single-cell sequencing platforms demonstrate the superiority of scTAG over various state-of-the-art clustering methods. Zhuohan Yu, Yifu Lu, Yunhe Wang 0002, Fan Tang, Ka-Chun Wong, Xiangtao Li |
AAAI | 5 |
| 2022 | EGFI: drug-drug interaction extraction and generation with fusion of enriched entity and sentence informationabstractMOTIVATION: The rapid growth in literature accumulates diverse and yet comprehensive biomedical knowledge hidden to be mined such as drug interactions. However, it is difficult to extract the heterogeneous knowledge to retrieve or even discover the latest and novel knowledge in an efficient manner. To address such a problem, we propose EGFI for extracting and consolidating drug interactions from large-scale medical literature text data. Specifically, EGFI consists of two parts: classification and generation. In the classification part, EGFI encompasses the language model BioBERT which has been comprehensively pretrained on biomedical corpus. In particular, we propose the multihead self-attention mechanism and packed BiGRU to fuse multiple semantic information for rigorous context modeling. In the generation part, EGFI utilizes another pretrained language model BioGPT-2 where the generation sentences are selected based on filtering rules. RESULTS: We evaluated the classification part on 'DDIs 2013' dataset and 'DTIs' dataset, achieving the F1 scores of 0.842 and 0.720 respectively. Moreover, we applied the classification part to distinguish high-quality generated sentences and verified with the existing growth truth to confirm the filtered sentences. The generated sentences that are not recorded in DrugBank and DDIs 2013 dataset demonstrated the potential of EGFI to identify novel drug relationships. AVAILABILITY: Source code are publicly available at https://github.com/Layne-Huang/EGFI. Jiecong Lin, Xiangtao Li, Linqi Song, Zetian Zheng, Ka-Chun Wong |
Briefings Bioinform. | 6 |
| 2022 | CoaDTI: multi-modal co-attention based framework for drug-target interaction annotationabstractMOTIVATION: The identification of drug-target interactions (DTIs) plays a vital role for in silico drug discovery, in which the drug is the chemical molecule, and the target is the protein residues in the binding pocket. Manual DTI annotation approaches remain reliable; however, it is notoriously laborious and time-consuming to test each drug-target pair exhaustively. Recently, the rapid growth of labelled DTI data has catalysed interests in high-throughput DTI prediction. Unfortunately, those methods highly rely on the manual features denoted by human, leading to errors. RESULTS: Here, we developed an end-to-end deep learning framework called CoaDTI to significantly improve the efficiency and interpretability of drug target annotation. CoaDTI incorporates the Co-attention mechanism to model the interaction information from the drug modality and protein modality. In particular, CoaDTI incorporates transformer to learn the protein representations from raw amino acid sequences, and GraphSage to extract the molecule graph features from SMILES. Furthermore, we proposed to employ the transfer learning strategy to encode protein features by pre-trained transformer to address the issue of scarce labelled data. The experimental results demonstrate that CoaDTI achieves competitive performance on three public datasets compared with state-of-the-art models. In addition, the transfer learning strategy further boosts the performance to an unprecedented level. The extended study reveals that CoaDTI can identify novel DTIs such as reactions between candidate drugs and severe acute respiratory syndrome coronavirus 2-associated proteins. The visualization of co-attention scores can illustrate the interpretability of our model for mechanistic insights. AVAILABILITY: Source code are publicly available at https://github.com/Layne-Huang/CoaDTI. Jiecong Lin, Rui Liu 0038, Zetian Zheng, Lingkuan Meng, Xingjian Chen, Xiangtao Li, Ka-Chun Wong |
Briefings Bioinform. | 8 |
| 2022 | High-throughput single-cell RNA-seq data imputation and characterization with surrogate-assisted automated deep learningabstractSingle-cell RNA sequencing (scRNA-seq) technologies have been heavily developed to probe gene expression profiles at single-cell resolution. Deep imputation methods have been proposed to address the related computational challenges (e.g. the gene sparsity in single-cell data). In particular, the neural architectures of those deep imputation models have been proven to be critical for performance. However, deep imputation architectures are difficult to design and tune for those without rich knowledge of deep neural networks and scRNA-seq. Therefore, Surrogate-assisted Evolutionary Deep Imputation Model (SEDIM) is proposed to automatically design the architectures of deep neural networks for imputing gene expression levels in scRNA-seq data without any manual tuning. Moreover, the proposed SEDIM constructs an offline surrogate model, which can accelerate the computational efficiency of the architectural search. Comprehensive studies show that SEDIM significantly improves the imputation and clustering performance compared with other benchmark methods. In addition, we also extensively explore the performance of SEDIM in other contexts and platforms including mass cytometry and metabolic profiling in a comprehensive manner. Marker gene detection, gene ontology enrichment and pathological analysis are conducted to provide novel insights into cell-type identification and the underlying mechanisms. The source code is available at https://github.com/li-shaochuan/SEDIM. Xiangtao Li, Shaochuan Li, Shixiong Zhang 0002, Ka-Chun Wong |
Briefings Bioinform. | 5 |
| 2022 | DeepMotifSyn: a deep learning approach to synthesize heterodimeric DNA motifs
Jiecong Lin, Xingjian Chen, Shixiong Zhang 0002, Ka-Chun Wong |
Briefings Bioinform. | 5 |
| 2022 | Reducing healthcare disparities using multiple multiethnic data distributions with fine-tuning of transfer learningabstractHealthcare disparities in multiethnic medical data is a major challenge; the main reason lies in the unequal data distribution of ethnic groups among data cohorts. Biomedical data collected from different cancer genome research projects may consist of mainly one ethnic group, such as people with European ancestry. In contrast, the data distribution of other ethnic races such as African, Asian, Hispanic, and Native Americans can be less visible than the counterpart. Data inequality in the biomedical field is an important research problem, resulting in the diverse performance of machine learning models while creating healthcare disparities. Previous researches have reduced the healthcare disparities only using limited data distributions. In our study, we work on fine-tuning of deep learning and transfer learning models with different multiethnic data distributions for the prognosis of 33 cancer types. In previous studies, to reduce the healthcare disparities, only a single ethnic cohort was used as the target domain with one major source domain. In contrast, we focused on multiple ethnic cohorts as the target domain in transfer learning using the TCGA and MMRF CoMMpass study datasets. After performance comparison for experiments with new data distributions, our proposed model shows promising performance for transfer learning schemes compared to the baseline approach for old and new data distributation experiments. Muhammad Toseef, Xiangtao Li, Ka-Chun Wong |
Briefings Bioinform. | 3 |
| 2022 | HCRNet: high-throughput circRNA-binding event identification from CLIP-seq data using deep temporal convolutional networkabstractIdentifying genome-wide binding events between circular RNAs (circRNAs) and RNA-binding proteins (RBPs) can greatly facilitate our understanding of functional mechanisms within circRNAs. Thanks to the development of cross-linked immunoprecipitation sequencing technology, large amounts of genome-wide circRNA binding event data have accumulated, providing opportunities for designing high-performance computational models to discriminate RBP interaction sites and thus to interpret the biological significance of circRNAs. Unfortunately, there are still no computational models sufficiently flexible to accommodate circRNAs from different data scales and with various degrees of feature representation. Here, we present HCRNet, a novel end-to-end framework for identification of circRNA-RBP binding events. To capture the hierarchical relationships, the multi-source biological information is fused to represent circRNAs, including various natural language sequence features. Furthermore, a deep temporal convolutional network incorporating global expectation pooling was developed to exploit the latent nucleotide dependencies in an exhaustive manner. We benchmarked HCRNet on 37 circRNA datasets and 31 linear RNA datasets to demonstrate the effectiveness of our proposed method. To evaluate further the model's robustness, we performed HCRNet on a full-length dataset containing 740 circRNAs. Results indicate that HCRNet generally outperforms existing methods. In addition, motif analyses were conducted to exhibit the interpretability of HCRNet on circRNAs. All supporting source code and data can be downloaded from https://github.com/yangyn533/HCRNet and https://doi.org/10.6084/m9.figshare.16943722.v1. And the web server of HCRNet is publicly accessible at http://39.104.118.143:5001/. Zilong Hou, Hongli Ma, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Briefings Bioinform. | 7 |
| 2022 | GMHCC: high-throughput analysis of biomolecular data using graph-based multiple hierarchical consensus clusteringabstractMOTIVATION: Thanks to the development of high-throughput sequencing technologies, massive amounts of various biomolecular data have been accumulated to revolutionize the study of genomics and molecular biology. One of the main challenges in analyzing this biomolecular data is to cluster their subtypes into subpopulations to facilitate subsequent downstream analysis. Recently, many clustering methods have been developed to address the biomolecular data. However, the computational methods often suffer from many limitations such as high dimensionality, data heterogeneity and noise. RESULTS: In our study, we develop a novel Graph-based Multiple Hierarchical Consensus Clustering (GMHCC) method with an unsupervised graph-based feature ranking (FR) and a graph-based linking method to explore the multiple hierarchical information of the underlying partitions of the consensus clustering for multiple types of biomolecular data. Indeed, we first propose to use a graph-based unsupervised FR model to measure each feature by building a graph over pairwise features and then providing each feature with a rank. Subsequently, to maintain the diversity and robustness of basic partitions (BPs), we propose multiple diverse feature subsets to generate several BPs and then explore the hierarchical structures of the multiple BPs by refining the global consensus function. Finally, we develop a new graph-based linking method, which explicitly considers the relationships between clusters to generate the final partition. Experiments on multiple types of biomolecular data including 35 cancer gene expression datasets and eight single-cell RNA-seq datasets validate the effectiveness of our method over several state-of-the-art consensus clustering approaches. Furthermore, differential gene analysis, gene ontology enrichment analysis and KEGG pathway analysis are conducted, providing novel insights into cell developmental lineages and characterization mechanisms. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/yifuLu/GMHCC. The software and the supporting data can be downloaded from: https://figshare.com/articles/software/GMHCC/17111291. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yifu Lu, Zhuohan Yu, Yunhe Wang 0006, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 5 |
| 2022 | scWMC: weighted matrix completion-based imputation of scRNA-seq data via prior subspace informationabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) can provide insight into gene expression patterns at the resolution of individual cells, which offers new opportunities to study the behavior of different cell types. However, it is often plagued by dropout events, a phenomenon where the expression value of a gene tends to be measured as zero in the expression matrix due to various technical defects. RESULTS: In this article, we argue that borrowing gene and cell information across column and row subspaces directly results in suboptimal solutions due to the noise contamination in imputing dropout values. Thus, to impute more precisely the dropout events in scRNA-seq data, we develop a regularization for leveraging that imperfect prior information to estimate the true underlying prior subspace and then embed it in a typical low-rank matrix completion-based framework, named scWMC. To evaluate the performance of the proposed method, we conduct comprehensive experiments on simulated and real scRNA-seq data. Extensive data analysis, including simulated analysis, cell clustering, differential expression analysis, functional genomic analysis, cell trajectory inference and scalability analysis, demonstrate that our method produces improved imputation results compared to competing methods that benefits subsequent downstream analysis. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/XuYuanchi/scWMC and test data is available at https://doi.org/10.5281/zenodo.6832477. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yanchi Su, Fuzhou Wang, Shixiong Zhang 0002, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 5 |
| 2022 | EDCNN: identification of genome-wide RNA-binding proteins using evolutionary deep convolutional neural networkabstractMOTIVATION: RNA-binding proteins (RBPs) are a group of proteins associated with RNA regulation and metabolism, and play an essential role in mediating the maturation, transport, localization and translation of RNA. Recently, Genome-wide RNA-binding event detection methods have been developed to predict RBPs. Unfortunately, the existing computational methods usually suffer some limitations, such as high-dimensionality, data sparsity and low model performance. RESULTS: Deep convolution neural network has a useful advantage for solving high-dimensional and sparse data. To improve further the performance of deep convolution neural network, we propose evolutionary deep convolutional neural network (EDCNN) to identify protein-RNA interactions by synergizing evolutionary optimization with gradient descent to enhance deep conventional neural network. In particular, EDCNN combines evolutionary algorithms and different gradient descent models in a complementary algorithm, where the gradient descent and evolution steps can alternately optimize the RNA-binding event search. To validate the performance of EDCNN, an experiment is conducted on two large-scale CLIP-seq datasets, and results reveal that EDCNN provides superior performance to other state-of-the-art methods. Furthermore, time complexity analysis, parameter analysis and motif analysis are conducted to demonstrate the effectiveness of our proposed algorithm from several perspectives. AVAILABILITY AND IMPLEMENTATION: The EDCNN algorithm is available at GitHub: https://github.com/yaweiwang1232/EDCNN. Both the software and the supporting data can be downloaded from: https://figshare.com/articles/software/EDCNN/16803217. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 4 |
| 2022 | Exploring high-throughput biomolecular data with multiobjective robust continuous clustering
Yunhe Wang 0002, Ka-Chun Wong, Xiangtao Li |
Inf. Sci. | 2 |
| 2022 | Multiple source transfer learning for dynamic multiobjective optimization
Yulong Ye, Qiuzhen Lin, Lijia Ma, Ka-Chun Wong, Maoguo Gong, Carlos A. Coello Coello |
Inf. Sci. | 4 |
| 2022 | Leveraging Multi-source knowledge for Chinese clinical named entity recognition via relational graph convolutional network
Yang Xiang 0003, Ka-Chun Wong, Qingcai Chen, Jun Yan 0010, Buzhou Tang |
J. Biomed. Informatics | 4 |
| 2022 | Intrusion detection using multi-objective evolutionary convolutional neural network for Internet of Things in Fog computing
Yi Chen 0020, Qiuzhen Lin, Wenhong Wei, Junkai Ji, Ka-Chun Wong, Carlos A. Coello Coello |
Knowl. Based Syst. | 5 |
| 2022 | A self-adaptive weighted differential evolution approach for large-scale feature selection
Xubin Wang 0001, Yunhe Wang 0002, Ka-Chun Wong, Xiangtao Li |
Knowl. Based Syst. | 3 |
| 2022 | Knowledge guided Bayesian classification for dynamic multi-objective optimization
Yulong Ye, Qiuzhen Lin, Ka-Chun Wong, Jianqiang Li 0001, Zhong Ming 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Genome-wide identification and characterization of DNA enhancers with a stacked multivariate fusion frameworkabstractEnhancers are short non-coding DNA sequences outside of the target promoter regions that can be bound by specific proteins to increase a gene's transcriptional activity, which has a crucial role in the spatiotemporal and quantitative regulation of gene expression. However, enhancers do not have a specific sequence motifs or structures, and their scattered distribution in the genome makes the identification of enhancers from human cell lines particularly challenging. Here we present a novel, stacked multivariate fusion framework called SMFM, which enables a comprehensive identification and analysis of enhancers from regulatory DNA sequences as well as their interpretation. Specifically, to characterize the hierarchical relationships of enhancer sequences, multi-source biological information and dynamic semantic information are fused to represent regulatory DNA enhancer sequences. Then, we implement a deep learning-based sequence network to learn the feature representation of the enhancer sequences comprehensively and to extract the implicit relationships in the dynamic semantic information. Ultimately, an ensemble machine learning classifier is trained based on the refined multi-source features and dynamic implicit relations obtained from the deep learning-based sequence network. Benchmarking experiments demonstrated that SMFM significantly outperforms other existing methods using several evaluation metrics. In addition, an independent test set was used to validate the generalization performance of SMFM by comparing it to other state-of-the-art enhancer identification methods. Moreover, we performed motif analysis based on the contribution scores of different bases of enhancer sequences to the final identification results. Besides, we conducted interpretability analysis of the identified enhancer sequences based on attention weights of EnhancerBERT, a fine-tuned BERT model that provides new insights into exploring the gene semantic information likely to underlie the discovered enhancers in an interpretable manner. Finally, in a human placenta study with 4,562 active distal gene regulatory enhancers, SMFM successfully exposed tissue-related placental development and the differential mechanism, demonstrating the generalizability and stability of our proposed framework. Zilong Hou, Ka-Chun Wong, Xiangtao Li |
PLoS Comput. Biol. | 4 |
| 2022 | A Self-Guided Reference Vector Strategy for Many-Objective OptimizationabstractGenerally, decomposition-based evolutionary algorithms in many-objective optimization (MaOEA/Ds) have widely used reference vectors (RVs) to provide search directions and maintain diversity. However, their performance is highly affected by the matching degree on the shapes of the RVs and the Pareto front (PF). To address this problem, this article proposes a self-guided RV (SRV) strategy for MaOEA/Ds, aiming to extract RVs from the population using a modified k -means clustering method. To give a promising clustering result, an angle-based density measurement strategy is used to initialize the centroids, which are then adjusted to obtain the final clusters, aiming to properly reflect the population's distribution. Afterward, these centroids are extracted to obtain adaptive RVs for self-guiding the search process. To verify the effectiveness of this SRV strategy, it is embedded into three well-known MaOEA/Ds that originally use the fixed RVs. Moreover, a new strategy of embedding SRV into MaOEA/Ds is discussed when the RVs are adjusted at each generation. The simulation results validate the superiority of our SRV strategy, when tackling numerous many-objective optimization problems with regular and irregular PFs. Songbai Liu, Qiuzhen Lin, Ka-Chun Wong, Carlos A. Coello Coello, Jianqiang Li 0001, Zhong Ming 0001, Jun Zhang 0003 |
IEEE Trans. Cybern. | 3 |
| 2022 | Evolutionary Multiobjective Clustering Algorithms With Ensemble for Patient StratificationabstractPatient stratification has been studied widely to tackle subtype diagnosis problems for effective treatment. Due to the dimensionality curse and poor interpretability of data, there is always a long-lasting challenge in constructing a stratification model with high diagnostic ability and good generalization. To address these problems, this article proposes two novel evolutionary multiobjective clustering algorithms with ensemble (NSGA-II-ECFE and MOEA/D-ECFE) with four cluster validity indices used as the objective functions. First, an effective ensemble construction method is developed to enrich the ensemble diversity. After that, an ensemble clustering fitness evaluation (ECFE) method is proposed to evaluate the ensembles by measuring the consensus clustering under those four objective functions. To generate the consensus clustering, ECFE exploits the hybrid co-association matrix from the ensembles and then dynamically selects the suitable clustering algorithm on that matrix. Multiple experiments have been conducted to demonstrate the effectiveness of the proposed algorithm in comparison with seven clustering algorithms, twelve ensemble clustering approaches, and two multiobjective clustering algorithms on 55 synthetic datasets and 35 real patient stratification datasets. The experimental results demonstrate the competitive edges of the proposed algorithms over those compared methods. Furthermore, the proposed algorithm is applied to extend its advantages by identifying cancer subtypes from five cancer-related single-cell RNA-seq datasets. Yunhe Wang 0002, Xiangtao Li, Ka-Chun Wong, Yi Chang 0001, Shengxiang Yang |
IEEE Trans. Cybern. | 3 |
| 2022 | Particle Swarm Optimized Gaussian Process Classifier for Treatment Discontinuation Prediction in Multicohort Metastatic Castration-Resistant Prostate Cancer PatientsabstractProstate cancer is the second leading cancer in men, according to the WHO world cancer report. Its prevention and treatment demand proper attention. Despite numerous attempts for disease prevention, prostate tumours can still become metastatic by blood circulation to other organs. Several treatments have been adopted. However, findings show that the docetaxel treatment induces adverse reactions in patients. Particle Swarm Optimized Gaussian Process Classifier (PSO-GPC) is proposed to determine when to discontinue treatment. Based on three cohorts of prostate cancer patients, we propose and compare several classifiers for the best performance in determining treatment discontinuation. Given the data skewness and class imbalance, the models are evaluated based on both the area under receiver operating characteristics curve (AUC) and area under precision recall curve (AUPRC). With the AUCs ranging between 0.6717-0.8499, and AUPRCs ranging between 0.1392-0.5423, PSO-GPC performs better than the state-of-the-art. We have carried out statistical analysis for ranking methods and analyzed independent cohort data with PSO-GPC, demonstrating its unbiased performance. A proper determination of treatment discontinuation in metastatic castration-resistant prostate cancer patients will reduce the mortality rate in cancer patients. Olutomilayo Olayemi Petinrin, Xiangtao Li, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | SRG-Vote: Predicting Mirna-Gene Relationships via Embedding and LSTM EnsembleabstractTargeted therapy for one for a set of genes has made it possible to apply precision medicine for different patients due to the existence of tumor heterogeneity. However, how to regulate those genes are still problematic. One of the natural regulators of genes is microRNAs. Thus, a better understanding of the miRNA-gene interaction mechanism might contribute to future diagnosis, prevention, and cancer therapy. The interactions between microRNA and genes play an essential role in molecular genetics. The in-vivo experiments validating the relationships between them are time-consuming, money-costly, and labor-intensive. With the development of high-throughput technology, we dealt with tons of biological data. However, extracting features from tremendous raw data and making a mathematical model is still a challenging topic. Machine learning and deep learning algorithms have become powerful tools in dealing with biological data. Inspired by this, in this paper, we propose a model that combines features/embedding extraction methods, deep learning algorithms, and a voting system. We leverage doc2vec to generate sequential embedding from molecular sequences. The role2vec, GCN, and GMM for geometrical embedding were generated from the complex network from similarity and pair-wise datasets. For the deep learning algorithms, we leveraged LSTM and Bi-LSTM according to different embedding and features. Finally, we adopted a voting system to balance results from different data sources. The results have shown that our voting system could achieve a higher AUC than the existing benchmark. The case studies demonstrate that our model could reveal potential relationships between miRNAs and genes. The source code, features, and predictive results can be downloaded at https://github.com/Xshelton/SRG-vote. Weidun Xie, Zetian Zheng, Qiuzhen Lin, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Subclass-Specific Prognosis and Treatment Efficacy Inference in Head and Neck Squamous CarcinomaabstractExploring the prognostic classification and biomarkers in Head and Neck Squamous Carcinoma (HNSC) is of great clinical significance. We hybridized three prominent strategies to comprehensively characterize the molecular features of HNSC. We constructed a 15-gene signature to predict patients' death risk with an average AUC of 0.744 for 1-, 3-, and 5-year on TCGA-HNSC training set, and average AUCs of 0.636, 0.584, 0.755 in GSE65858, GSE-112026, CPTAC-HNSCC datasets, respectively. By combined with NMF clustering and consensus clustering of fraction of tumor immune cell infiltration (ICI) in the tumor microenvironment (TME), we captured a more refined biological characteristics of HNSC, and observed a prognosis heterogeneity in high tumor immunity patients. By matching tumor subset-specific expression signatures to drug-induced cell line expression profiles from large-scale pharmacogenomic databases in the OCTAD workspace, we identified a group of HNSC patients featured with poor prognosis and demonstrated that the individuals in this group are likely to receive increased drug sensitivity to reverse differentially expressed disease signature genes. This trend is especially highlighted among those with higher death risk and tumour immunity. Zetian Zheng, Weidun Xie, Xingjian Chen, Fuzhou Wang, Xiangtao Li, Qiuzhen Lin, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 8 |
| 2022 | Multiobjective Deep Clustering and its Applications in Single-cell RNA-seq DataabstractSingle-cell RNA sequencing is a transformative technology that enables us to study the heterogeneity of the tissue at the cellular level. Clustering is used as the key computational approach to group cells under the transcriptome profiles from single-cell RNA-seq data. However, accurate identification of distinct cell types is facing the challenge of high dimensionality, and it could cause uninformative clusters when clustering is directly applied on the original transcriptome. To address such challenge, an evolutionary multiobjective deep clustering (EMDC) algorithm is proposed to identify single-cell RNA-seq data in this study. First, EMDC removes redundant and irrelevant genes by applying the differential gene expression analysis to identify differentially expressed genes across biological conditions. After that, a deep autoencoder is proposed to project the high-dimensional data into different low-dimensional nonlinear embedding subspaces under different bottleneck layers. Then, the basic clustering algorithm is applied in those nonlinear embedding subspaces to generate some basic clustering results to produce the cluster ensemble. To lessen the unnecessary cost produced by those clusterings in the ensemble, the multiobjective evolutionary optimization is designed to prune the basic clustering results in the ensemble, unleashing its cell type discovery performance under three objective functions. Multiple experiments have been conducted on 30 synthetic single-cell RNA-seq datasets and six real single-cell RNA-seq datasets, which reveal that EMDC outperforms eight other clustering methods and three multiobjective optimization algorithms in cell type identification. In addition, we have conducted extensive comparisons to effectively demonstrate the impact of each component in our proposed EMDC. Yunhe Wang 0002, Chuang Bian, Ka-Chun Wong, Xiangtao Li, Shengxiang Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Metric Learning Based Vision Transformer for Product Matching
Wei Shao 0009, Fuzhou Wang, Weidun Xie, Ka-Chun Wong |
ICONIP (1) | 5 |
| 2021 | Human host status inference from temporal microbiome changes via recurrent neural networksabstractWith the rapid increase in sequencing data, human host status inference (e.g. healthy or sick) from microbiome data has become an important issue. Existing studies are mostly based on single-point microbiome composition, while it is rare that the host status is predicted from longitudinal microbiome data. However, single-point-based methods cannot capture the dynamic patterns between the temporal changes and host status. Therefore, it remains challenging to build good predictive models as well as scaling to different microbiome contexts. On the other hand, existing methods are mainly targeted for disease prediction and seldom investigate other host statuses. To fill the gap, we propose a comprehensive deep learning-based framework that utilizes longitudinal microbiome data as input to infer the human host status. Specifically, the framework is composed of specific data preparation strategies and a recurrent neural network tailored for longitudinal microbiome data. In experiments, we evaluated the proposed method on both semi-synthetic and real datasets based on different sequencing technologies and metagenomic contexts. The results indicate that our method achieves robust performance compared to other baseline and state-of-the-art classifiers and provides a significant reduction in prediction time. Xingjian Chen, Lingjing Liu, Jianyi Yang 0002, Ka-Chun Wong |
Briefings Bioinform. | 5 |
| 2021 | RNCE: network integration with reciprocal neighbors contextual encoding for multi-modal drug community study on cancer targetsabstractMining drug targets and mechanisms of action (MoA) for novel anticancer drugs from pharmacogenomic data is a path to enhance the drug discovery efficiency. Recent approaches have successfully attempted to discover targets/MoA by characterizing drug similarities and communities with integrative methods on multi-modal or multi-omics drug information. However, the sparse and imbalanced community size structure of the drug network is seldom considered in recent approaches. Consequently, we developed a novel network integration approach accounting for network structure by a reciprocal nearest neighbor and contextual information encoding (RNCE) approach. In addition, we proposed a tailor-made clustering algorithm to perform drug community detection on drug networks. RNCE and spectral clustering are proved to outperform state-of-the-art approaches in a series of tests, including network similarity tests and community detection tests on two drug databases. The observed improvement of RNCE can contribute to the field of drug discovery and the related multi-modal/multi-omics integrative studies. Availabilityhttps://github.com/WINGHARE/RNCE. Ka-Chun Wong |
Briefings Bioinform. | 2 |
| 2021 | iDeepSubMito: identification of protein submitochondrial localization with deep learningabstractMitochondria are membrane-bound organelles containing over 1000 different proteins involved in mitochondrial function, gene expression and metabolic processes. Accurate localization of those proteins in the mitochondrial compartments is critical to their operation. A few computational methods have been developed for predicting submitochondrial localization from the protein sequences. Unfortunately, most of these computational methods focus on employing biological features or evolutionary information to extract sequence features, which greatly limits the performance of subsequent identification. Moreover, the efficiency of most computational models is still under explored, especially the deep learning feature, which is promising but requires improvement. To address these limitations, we propose a novel computational method called iDeepSubMito to predict the location of mitochondrial proteins to the submitochondrial compartments. First, we adopted a coding scheme using the ProteinELMo to model the probability distribution over the protein sequences and then represent the protein sequences as continuous vectors. Then, we proposed and implemented convolutional neural network architecture based on the bidirectional LSTM with self-attention mechanism, to effectively explore the contextual information and protein sequence semantic features. To demonstrate the effectiveness of our proposed iDeepSubMito, we performed cross-validation on two datasets containing 424 proteins and 570 proteins respectively, and consisting of four different mitochondrial compartments (matrix, inner membrane, outer membrane and intermembrane regions). Experimental results revealed that our method outperformed other computational methods. In addition, we tested iDeepSubMito on the M187, M983 and MitoCarta3.0 to further verify the efficiency of our method. Finally, the motif analysis and the interpretability analysis were conducted to reveal novel insights into subcellular biological functions of mitochondrial proteins. iDeepSubMito source code is available on GitHub at https://github.com/houzl3416/iDeepSubMito. Zilong Hou, Ka-Chun Wong, Xiangtao Li |
Briefings Bioinform. | 4 |
| 2021 | Identification of pan-cancer Ras pathway activation with deep learningabstractThe identification of hidden responders is often an essential challenge in precision oncology. A recent attempt based on machine learning has been proposed for classifying aberrant pathway activity from multiomic cancer data. However, we note several critical limitations there, such as high-dimensionality, data sparsity and model performance. Given the central importance and broad impact of precision oncology, we propose nature-inspired deep Ras activation pan-cancer (NatDRAP), a deep neural network (DNN) model, to address those restrictions for the identification of hidden responders. In this study, we develop the nature-inspired deep learning model that integrates bulk RNA sequencing, copy number and mutation data from PanCanAltas to detect pan-cancer Ras pathway activation. In NatDRAP, we propose to synergize the nature-inspired artificial bee colony algorithm with different gradient-based optimizers in one framework for optimizing DNNs in a collaborative manner. Multiple experiments were conducted on 33 different cancer types across PanCanAtlas. The experimental results demonstrate that the proposed NatDRAP can provide superior performance over other benchmark methods with strong robustness towards diagnosing RAS aberrant pathway activity across different cancer types. In addition, gene ontology enrichment and pathological analysis are conducted to reveal novel insights into the RAS aberrant pathway activity identification and characterization. NatDRAP is written in Python and available at https://github.com/lixt314/NatDRAP1. Xiangtao Li, Shaochuan Li, Yunhe Wang 0002, Shixiong Zhang 0002, Ka-Chun Wong |
Briefings Bioinform. | 5 |
| 2021 | Deep embedded clustering with multiple objectives on scRNA-seq dataabstractIn recent years, single-cell RNA sequencing (scRNA-seq) technologies have been widely adopted to interrogate gene expression of individual cells; it brings opportunities to understand the underlying processes in a high-throughput manner. Deep embedded clustering (DEC) was demonstrated successful in high-dimensional sparse scRNA-seq data by joint feature learning and cluster assignment for identifying cell types simultaneously. However, the deep network architecture for embedding clustering is not trivial to optimize. Therefore, we propose an evolutionary multiobjective DEC by synergizing the multiobjective evolutionary optimization to simultaneously evolve the hyperparameters and architectures of DEC in an automatic manner. Firstly, a denoising autoencoder is integrated into the DEC to project the high-dimensional sparse scRNA-seq data into a low-dimensional space. After that, to guide the evolution, three objective functions are formulated to balance the model's generality and clustering performance for robustness. Meanwhile, migration and mutation operators are proposed to optimize the objective functions to select the suitable hyperparameters and architectures of DEC in the multiobjective framework. Multiple comparison analyses are conducted on twenty synthetic data and eight real data from different representative single-cell sequencing platforms to validate the effectiveness. The experimental results reveal that the proposed algorithm outperforms other state-of-the-art clustering methods under different metrics. Meanwhile, marker genes identification, gene ontology enrichment and pathology analysis are conducted to reveal novel insights into the cell type identification and characterization mechanisms. Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong |
Briefings Bioinform. | 3 |
| 2021 | iCircRBP-DHN: identification of circRNA-RBP interaction sites using deep hierarchical networkabstractCircular RNAs (circRNAs) are widely expressed in eukaryotes. The genome-wide interactions between circRNAs and RNA-binding proteins (RBPs) can be probed from cross-linking immunoprecipitation with sequencing data. Therefore, computational methods have been developed for identifying RBP binding sites on circRNAs. Unfortunately, those computational methods often suffer from the low discriminative power of feature representations, numerical instability and poor scalability. To address those limitations, we propose a novel computational method called iCircRBP-DHN using deep hierarchical network for discriminating circRNA-RBP binding sites. The network architecture can be regarded as a deep multi-scale residual network followed by bidirectional gated recurrent units (BiGRUs) with the self-attention mechanism, which can simultaneously extract local and global contextual information. Meanwhile, we propose novel encoding schemes by integrating CircRNA2Vec and the K-tuple nucleotide frequency pattern to represent different degrees of nucleotide dependencies. To validate the effectiveness of our proposed iCircRBP-DHN, we compared its performance with other computational methods on 37 circRNAs datasets and 31 linear RNAs datasets, respectively. The experimental results reveal that iCircRBP-DHN can achieve superior performance over those state-of-the-art algorithms. Moreover, we perform motif analysis on circRNAs bound by those different RBPs, demonstrating that our proposed CircRNA2Vec encoding scheme can be promising. The iCircRBP-DHN method is made available at https://github.com/houzl3416/iCircRBP-DHN. Zilong Hou, Zhiqiang Ma 0003, Xiangtao Li, Ka-Chun Wong |
Briefings Bioinform. | 5 |
| 2021 | Identification of haploinsufficient genes from epigenomic data using deep forestabstractHaploinsufficiency, wherein a single allele is not enough to maintain normal functions, can lead to many diseases including cancers and neurodevelopmental disorders. Recently, computational methods for identifying haploinsufficiency have been developed. However, most of those computational methods suffer from study bias, experimental noise and instability, resulting in unsatisfactory identification of haploinsufficient genes. To address those challenges, we propose a deep forest model, called HaForest, to identify haploinsufficient genes. The multiscale scanning is proposed to extract local contextual representations from input features under Linear Discriminant Analysis. After that, the cascade forest structure is applied to obtain the concatenated features directly by integrating decision-tree-based forests. Meanwhile, to exploit the complex dependency structure among haploinsufficient genes, the LightGBM library is embedded into HaForest to reveal the highly expressive features. To validate the effectiveness of our method, we compared it to several computational methods and four deep learning algorithms on five epigenomic data sets. The results reveal that HaForest achieves superior performance over the other algorithms, demonstrating its unique and complementary performance in identifying haploinsufficient genes. The standalone tool is available at https://github.com/yangyn533/HaForest. Shaochuan Li, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Briefings Bioinform. | 5 |
| 2021 | Elucidating transcriptomic profiles from single-cell RNA sequencing data using nature-inspired compressed sensingabstractGene-expression profiling can define the cell state and gene-expression pattern of cells at the genetic level in a high-throughput manner. With the development of transcriptome techniques, processing high-dimensional genetic data has become a major challenge in expression profiling. Thanks to the recent widespread use of matrix decomposition methods in bioinformatics, a computational framework based on compressed sensing was adopted to reduce dimensionality. However, compressed sensing requires an optimization strategy to learn the modular dictionaries and activity levels from the low-dimensional random composite measurements to reconstruct the high-dimensional gene-expression data. Considering this, here we introduce and compare four compressed sensing frameworks coming from nature-inspired optimization algorithms (CSCS, ABCCS, BACS and FACS) to improve the quality of the decompression process. Several experiments establish that the three proposed methods outperform benchmark methods on nine different datasets, especially the FACS method. We illustrate therefore, the robustness and convergence of FACS in various aspects; notably, time complexity and parameter analyses highlight properties of our proposed FACS. Furthermore, differential gene-expression analysis, cell-type clustering, gene ontology enrichment and pathology analysis are conducted, which bring novel insights into cell-type identification and characterization mechanisms from different perspectives. All algorithms are written in Python and available at https://github.com/Philyzh8/Nature-inspired-CS. Zhuohan Yu, Chuang Bian, Genggeng Liu, Shixiong Zhang 0002, Ka-Chun Wong, Xiangtao Li |
Briefings Bioinform. | 5 |
| 2021 | Early cancer detection from genome-wide cell-free DNA fragmentation via shuffled frog leaping algorithm and support vector machineabstractMOTIVATION: Early cancer detection is significant for patient mortality rate reduction. Although machine learning has been widely employed in that context, there are still deficiencies. In this work, we studied different machine learning algorithms for early cancer detection and proposed an Adaptive Support Vector Machine (ASVM) method by synergizing Shuffled Frog Leaping Algorithm and Support Vector Machine (SVM) in this study. RESULTS: Since ASVM regulates SVM for parameter adaption based on data characteristics, the experimental results reflected the robust generalization capability of ASVM on different datasets under different settings; for instance, ASVM can enhance the sensitivity by over 10% for early cancer detection compared with SVM. Besides, our proposed ASVM outperformed Grid Search + SVM and Random Search + SVM by significant margins in terms of the area under the ROC curve (AUC) (0.938 versus 0.922 versus 0.921). AVAILABILITY AND IMPLEMENTATION: The proposed algorithm and dataset are available at https://github.com/ElaineLIU-920/ASVM-for-Early-Cancer-Detection. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Linjing Liu, Xingjian Chen, Ka-Chun Wong |
Bioinform. | 3 |
| 2021 | CancerEMC: frontline non-invasive cancer screening from circulating protein biomarkers and mutations in cell-free DNAabstractMOTIVATION: The early detection of cancer through accessible blood tests can foster early patient interventions. Although there are developments in cancer detection from cell-free DNA (cfDNA), its accuracy remains speculative. Given its central importance with broad impacts, we aspire to address the challenge. METHOD: A bagging Ensemble Meta Classifier (CancerEMC) is proposed for early cancer detection based on circulating protein biomarkers and mutations in cfDNA from blood. CancerEMC is generally designed for both binary cancer detection and multi-class cancer type localization. It can address the class imbalance problem in multi-analyte blood test data based on robust oversampling and adaptive synthesis techniques. RESULTS: Based on the clinical blood test data, we observe that the proposed CancerEMC has outperformed other algorithms and state-of-the-arts studies (including CancerSEEK) for cancer detection. The results reveal that our proposed method (i.e. CancerEMC) can achieve the best performance result for both binary cancer classification with 99.17% accuracy (AUC = 0.999) and localized multiple cancer detection with 74.12% accuracy (AUC = 0.938). Addressing the data imbalance issue with oversampling techniques, the accuracy can be increased to 91.50% (AUC = 0.992), where the state-of-the-art method can only be estimated at 69.64% (AUC = 0.921). Similar results can also be observed on independent and isolated testing data. AVAILABILITY: https://github.com/saifurcubd/Cancer-Detection. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Saifur Rahaman, Xiangtao Li, Jun Yu 0004, Ka-Chun Wong |
Bioinform. | 4 |
| 2021 | Future DNA computing device and accompanied tool stack: Towards high-throughput computation
Shankai Yan, Ka-Chun Wong |
Future Gener. Comput. Syst. | 2 |
| 2021 | Categorical Matrix Completion With Active Learning for High-Throughput ScreeningabstractThe recent advances in wet-lab automation enable high-throughput experiments to be conducted seamlessly. In particular, the exhaustive enumeration of all possible conditions is always involved in high-throughput screening. Nonetheless, such a screening strategy is hardly believed to be optimal and cost-effective. By incorporating artificial intelligence, we design an open-source model based on categorical matrix completion and active machine learning to guide high throughput screening experiments. Specifically, we narrow our scope to the high-throughput screening for chemical compound effects on diverse protein sub-cellular locations. In the proposed model, we believe that exploration is more important than the exploitation in the long-run of high-throughput screening experiment, Therefore, we design several innovations to circumvent the existing limitations. In particular, categorical matrix completion is designed to accurately impute the missing experiments while margin sampling is also implemented for uncertainty estimation. The model is systematically tested on both simulated and real data. The simulation results reflect that our model can be robust to diverse scenarios, while the real data results demonstrate the wet-lab applicability of our model for high-throughput screening experiments. Lastly, we attribute the model success to its exploration ability by revealing the related matrix ranks and distinct experiment coverage comparisons. Junhui Hou, Ka-Chun Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Evolving Transcriptomic Profiles From Single-Cell RNA-Seq Data Using Nature-Inspired Multiobjective OptimizationabstractTranscriptomic profiling plays an important role in post-genomic analysis. Especially, the single-cell RNA-seq technology has advanced our understanding of gene expression from cell population level into individual cell level. Many computational methods have been proposed to decipher transcriptomic profiles from those RNA-seq data. However, most of the related algorithms suffer from realistic restrictions such as high dimensionality and premature convergence. In this paper, we propose and formulate an evolutionary multiobjective blind compressed sensing (EMOBCS) to address those problems for evolving transcriptomic profiles from single-cell RNA-seq data. In the proposed framework, to characterize various gene expression profile models, two objective functions including chi-squared kernel score and euclidean distance of different gene expression profiles are formulated. After that, multiobjective blind compressed sensing based on artificial bee colony is designed to optimize the two objective functions on single-cell RNA-seq data by proposing a rank probability model and two new search strategies into the cooperative convolution framework in an unbiased manner. To demonstrate its effectiveness, extensive experiments have been conducted, comparing the proposed algorithm with 14 algorithms including eight state-of-the-art algorithms and six different EMOBCS algorithms under different search strategies on 10 single-cell RNA-seq datasets and one case study. The experimental results reveal that the proposed algorithm is better than or comparable with those compared algorithms. Furthermore, we also conduct the time complexity analysis, convergence analysis, and parameter analysis to demonstrate various properties of EMOBCS. Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Evolving Multiobjective Cancer Subtype Diagnosis From Cancer Gene Expression DataabstractDetection and diagnosis of cancer are especially essential for early prevention and effective treatments. Many studies have been proposed to tackle the subtype diagnosis problems with those data, which often suffer from low diagnostic ability and bad generalization. This article studies a multiobjective PSO-based hybrid algorithm (MOPSOHA) to optimize four objectives including the number of features, the accuracy, and two entropy-based measures: the relevance and the redundancy simultaneously, diagnosing the cancer data with high classification power and robustness. First, we propose a novel binary encoding strategy to choose informative gene subsets to optimize those objective functions. Second, a mutation operator is designed to enhance the exploration capability of the swarm. Finally, a local search method based on the "best/1" mutation operator of differential evolutionary algorithm (DE) is employed to exploit the neighborhood area with sparse high-quality solutions since the base vector always approaches to some good promising areas. In order to demonstrate the effectiveness of MOPSOHA, it is tested on 41 cancer datasets including thirty-five cancer gene expression datasets and six independent disease datasets. Compared MOPSOHA with other state-of-the-art algorithms, the performance of MOPSOHA is superior to other algorithms in most of the benchmark datasets. Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Multiobjective Genome-Wide RNA-Binding Event Identification From CLIP-Seq DataabstractRNA-binding proteins (RBPs) are the master regulators of mRNA processing, which are vital players for the post-transcriptional control of gene expression. In recent years, crosslinking immunoprecipitation sequencing (CLIP-seq) technologies have enabled us to sequence massive amounts of genome-wide RNA-binding event data. Its increasing availability provides opportunities to identify protein-RNA interactions on a genome-wide scale. Genome-wide RNA-binding event detection methods have been developed to the understanding of the proteins' functions within cellular processes. Unfortunately, those methods often suffer from realistic restrictions, such as high costs, intensive computation, high dimensionality, numerical instability, and data sparsity. We present a computational method [multiobjective forest algorithm (MFA)] to identify protein-RNA interactions from CLIP-seq data by synergizing multiobjective biogeography-based optimization (BBO) with random forest (RF). Since most of the tree-structured classifiers in RF are unnecessarily bulky with extra time costs and memory consumption, multiobjective BBO is designed to prune the unsuitable tree-structured classifiers dynamically. Moreover, to direct the evolution dynamics of the MFA, two objective functions are formulated to balance model generality and complexity for robust performance. To validate our MFA method, we compare its performance across 31 large-scale CLIP-seq datasets. The experimental results demonstrate that MFA can obtain superior performance over the current state-of-the-art methods. Mechanistic insights are also revealed and discussed to explore the multifaceted aspects of MFA through data source importance analysis, matrix rank estimations, seeding component perturbations, and multiobjective optimization methodology comparisons. Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong |
IEEE Trans. Cybern. | 3 |
| 2021 | Nature-Inspired Compressed Sensing for Transcriptomic Profiling From Random Composite MeasurementsabstractTranscriptomic profiling is a high-throughput approach to measure gene expression levels under different experimental conditions at different timings. With the development of the related technologies such as single-cell RNA-Seq, the dimensions of gene expression data are increased to hundreds of thousands or more for high-resolution insights. There is a long-lasting challenge in exploiting the relations between transcriptomic profiles and random composite measurements. To address it, we proposed a mathematical framework based on differential evolution (global search) with the help of compressed sensing (local search) termed as DECS. Exploiting the inherent sparse nature of gene expression data, the proposed DECS method can learn the sparse module dictionaries and levels from the low-dimensional random composite measurements for reconstructing the high-dimensional gene expression data with significant orders of magnitude (e.g. 200 × ). Several experiments were conducted to compare DECS with three benchmark methods, demonstrating that the proposed DECS outperforms the benchmark methods and can recover most of the gene expression patterns. The underlying reasons are discussed and illustrated by revealing the related mechanistic insights through extensive benchmarks on nine GSE datasets and their sensitivity analysis. Shixiong Zhang 0002, Xiangtao Li, Qiuzhen Lin, Ka-Chun Wong |
IEEE Trans. Cybern. | 4 |
| 2020 | Context awareness and embedding for biomedical event extractionabstractMOTIVATION: Biomedical event extraction is fundamental for information extraction in molecular biology and biomedical research. The detected events form the central basis for comprehensive biomedical knowledge fusion, facilitating the digestion of massive information influx from the literature. Limited by the event context, the existing event detection models are mostly applicable for a single task. A general and scalable computational model is desiderated for biomedical knowledge management. RESULTS: We consider and propose a bottom-up detection framework to identify the events from recognized arguments. To capture the relations between the arguments, we trained a bidirectional long short-term memory network to model their context embedding. Leveraging the compositional attributes, we further derived the candidate samples for training event classifiers. We built our models on the datasets from BioNLP Shared Task for evaluations. Our method achieved the average F-scores of 0.81 and 0.92 on BioNLPST-BGI and BioNLPST-BB datasets, respectively. Comparing with seven state-of-the-art methods, our method nearly doubled the existing F-score performance (0.92 versus 0.56) on the BioNLPST-BB dataset. Case studies were conducted to reveal the underlying reasons. AVAILABILITY AND IMPLEMENTATION: https://github.com/cskyan/evntextrc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shankai Yan, Ka-Chun Wong |
Bioinform. | 2 |
| 2020 | Nature-inspired multiobjective patient stratification from cancer gene expression data
Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Inf. Sci. | 3 |
| 2020 | Verbal aggression detection on Twitter comments: convolutional neural network for short-text sentiment analysis
Shankai Yan, Ka-Chun Wong |
Neural Comput. Appl. | 3 |
| 2020 | Special issue of 2017 India International Congress on Computational Intelligence
Suash Deb, Ka-Chun Wong, Thomas Hanne |
Neural Comput. Appl. | 2 |
| 2020 | EDITORIAL: Special Issue of 2018 India International Congress on Computational Intelligence
Suash Deb, Ka-Chun Wong, Xin-She Yang 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Single-Cell RNA Sequencing Data Interpretation by Evolutionary Multiobjective ClusteringabstractIn recent years, single-cell RNA sequencing reveals diverse cell genetics at unprecedented resolutions. Such technological advances enable researchers to uncover the functionally distinct cell subtypes such as hematopoietic stem cell subpopulation identification. However, most of the related algorithms have been hindered by the high-dimensionality and sparse nature of single-cell RNA sequencing (RNA-seq) data. To address those problems, we propose a multiobjective evolutionary clustering based on adaptive non-negative matrix factorization (MCANMF) for multiobjective single-cell RNA-seq data clustering. First, adaptive non-negative matrix factorization is proposed to decompose data for feature extraction. After that, a multiobjective clustering algorithm based on learning vector quantization is proposed to analyze single-cell RNA-seq data. To validate the effectiveness of MCANMF, we benchmark MCANMF against 15 state-of-the-art methods including seven feature extraction methods, seven clustering methods, and the kernel-based similarity learning method on six published single-cell RNA sequencing datasets comprehensively. When compared with those 15 state-of-the-art methods, MCANMF performs better than the others on those single-cell RNA sequencing datasets according to multiple evaluation metrics. Moreover, the MCANMF component analysis, time complexity analysis, and parameter analysis are conducted to demonstrate various properties of our proposed algorithm. Xiangtao Li, Ka-Chun Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Nature-Inspired Multiobjective Epistasis Elucidation from Genome-Wide Association StudiesabstractIn recent years, the detection of epistatic interactions of multiple genetic variants on the causes of complex diseases brings a significant challenge in genome-wide association studies (GWAS). However, most of the existing methods still suffer from algorithmic limitations such as single-objective optimization, intensive computational requirement, and premature convergence. In this paper, we propose and formulate an epistatic interaction multi-objective artificial bee colony algorithm based on decomposition (EIMOABC/D) to address those problems for genetic interaction detection in genome-wide association studies. First, to direct the genetic interaction detection, two objective functions are formulated to characterize various epistatic models; rank probability model is proposed to sort each population into different nondomination levels based on the fast nondominated sorting approach. After that, the mutual information based local search algorithm is proposed to guide the population search for disease model evaluations in an unbiased manner. To validate the effectiveness of EIMOABC/D, we compare EIMOABC/D against seven state-of-the-art methods on 77 epistatic models including eight small-scale epistatic models with marginal effects, eight large-scale epistatic models with marginal effects, 60 large-scale epistatic models without any marginal effect, and one case study. The experimental results indicate that our proposed algorithm EIMOABC/D outperforms seven state-of-the-art methods on those epistatic models. Furthermore, time complexity analysis and parameter analysis are conducted to demonstrate various properties of our proposed algorithm. Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Deleterious Non-Synonymous Single Nucleotide Polymorphism Predictions on Human Transcription FactorsabstractTranscription factors (TFs) are the major components of human gene regulation. In particular, they bind onto specific DNA sequences and regulate neighborhood genes in different tissues at different developmental stages. Non-synonymous single nucleotide polymorphisms on its protein-coding sequences could result in undesired consequences in human. Therefore, it is necessary to develop methods for predicting any abnormality among those non-synonymous single nucleotide polymorphisms. To address it, we have developed and compared different strategies to predict deleterious non-synonymous single nucleotide polymorphisms (also known as missense mutations) on the protein-coding sequences of human TFs. Taking advantage of evolutionary conservation signals, we have developed and compared different classifiers with different feature sets as computed from different evolutionarily related sequence collections. The results indicate that the classic ensemble algorithm, Adaboost with decision stumps, with orthologous sequence collection, has performed the best (namely, TFmedic). We have further compared TFmedic with other state-of-the-arts methods (i.e., PolyPhen-2 and SIFT) on PolyPhen-2's own datasets, demonstrating that TFmedic can outperform the others. As applications, we have further applied TFmedic to all possible missense mutations on all human transcription factors; the proteome-wide results reveal interesting insights, consistent with the existing physiochemical knowledge. A case study with the actual 3D structure is conducted, revealing how TFmedic can be contributed to protein-DNA binding complex studies. Ka-Chun Wong, Shankai Yan, Qiuzhen Lin, Xiangtao Li, Chengbin Peng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | GESgnExt: Gene Expression Signature Extraction and Meta-Analysis on Gene Expression OmnibusabstractThe gene expression omnibus (GEO) repository harbours an exponentially increasing number of gene expression studies. The expression data, as well as the related metadata, provides an abundant resource for knowledge discovery. Each study in GEO focuses on the gene expression perturbation of a specific subject (e.g., gene, drug, and disease). The identification of those subjects and the associations among them are beneficial for further in-depth studies. However, they cannot be directly inferred from the studies. A unified representation of those subjects (i.e., gene expression signatures) is desired. We developed GESgnExt for the automatic construction of gene expression signatures. The resultant 6542 signatures are built on 1934 series and 35 919 samples from GEO. To evaluate its significance, we calculated the similarities among those signatures and compared the discovered associations against the existing interaction databases. The signatures connect the genes, drugs, and diseases, covering most of the experimentally validated interactions. Besides, we have discovered 3307 novel signatures and their related associations, complementing the existing signature knowledge. The biomedical relevance of GESgnExt is demonstrated further in multiple case studies, providing mechanistic insights into its knowledge discovery process. Shankai Yan, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Single-cell RNA-seq interpretations using evolutionary multiobjective ensemble pruningabstractMOTIVATION: In recent years, single-cell RNA sequencing enables us to discover cell types or even subtypes. Its increasing availability provides opportunities to identify cell populations from single-cell RNA-seq data. Computational methods have been employed to reveal the gene expression variations among multiple cell populations. Unfortunately, the existing ones can suffer from realistic restrictions such as experimental noises, numerical instability, high dimensionality and computational scalability. RESULTS: We propose an evolutionary multiobjective ensemble pruning algorithm (EMEP) that addresses those realistic restrictions. Our EMEP algorithm first applies the unsupervised dimensionality reduction to project data from the original high dimensions to low-dimensional subspaces; basic clustering algorithms are applied in those new subspaces to generate different clustering results to form cluster ensembles. However, most of those cluster ensembles are unnecessarily bulky with the expense of extra time costs and memory consumption. To overcome that problem, EMEP is designed to dynamically select the suitable clustering results from the ensembles. Moreover, to guide the multiobjective ensemble evolution, three cluster validity indices including the overall cluster deviation, the within-cluster compactness and the number of basic partition clusters are formulated as the objective functions to unleash its cell type discovery performance using evolutionary multiobjective optimization. We applied EMEP to 55 simulated datasets and seven real single-cell RNA-seq datasets, including six single-cell RNA-seq dataset and one large-scale dataset with 3005 cells and 4412 genes. Two case studies are also conducted to reveal mechanistic insights into the biological relevance of EMEP. We found that EMEP can achieve superior performance over the other clustering algorithms, demonstrating that EMEP can identify cell populations clearly. AVAILABILITY AND IMPLEMENTATION: EMEP is written in Matlab and available at https://github.com/lixt314/EMEP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong |
Bioinform. | 3 |
| 2019 | Synergizing CRISPR/Cas9 off-target predictions for ensemble insights and practical applicationsabstractMOTIVATION: The RNA-guided CRISPR/Cas9 system has been widely applied to genome editing. CRISPR/Cas9 system can effectively edit the on-target genes. Nonetheless, it has recently been demonstrated that many homologous off-target genomic sequences could be mutated, leading to unexpected gene-editing outcomes. Therefore, a plethora of tools were proposed for the prediction of off-target activities of CRISPR/Cas9. Nonetheless, each computational tool has its own advantages and drawbacks under diverse conditions. It is hardly believed that a single tool is optimal for all conditions. Hence, we would like to explore the ensemble learning potential on synergizing multiple tools with genomic annotations together to enhance its predictive abilities. RESULTS: We proposed an ensemble learning framework which synergizes multiple tools together to predict the off-target activities of CRISPR/Cas9 in different combinations. Interestingly, the ensemble learning using AdaBoost outperformed other individual off-target predictive tools. We also investigated the effect of evolutionary conservation (PhyloP and PhastCons) and chromatin annotations (ChromHMM and Segway) and found that only PhyloP can enhance the predictive capabilities further. Case studies are conducted to reveal ensemble insights into the off-target predictions, demonstrating how the current study can be applied in different genomic contexts. The best prediction predicted by AdaBoost is up to 0.9383 (AUC) and 0.2998 (PRC) that outperforms other classifiers. This is ascribable to the fact that AdaBoost introduces a new weak classifier (i.e. decision stump) in each iteration to learn the DNA sequences that were misclassified as off-targets until a small error rate is reached iteratively. AVAILABILITY AND IMPLEMENTATION: The source codes are freely available on GitHub at https://github.com/Alexzsx/CRISPR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shixiong Zhang 0002, Xiangtao Li, Qiuzhen Lin, Ka-Chun Wong |
Bioinform. | 4 |
| 2019 | Elucidating Genome-Wide Protein-RNA Interactions Using Differential EvolutionabstractRNA-binding proteins (RBPs) play an important role in the post-transcriptional control of RNAs, such as splicing, polyadenylation, mRNA stabilization, mRNA localization, and translation. Thanks to the recent breakthrough, non-negative matrix factorization (NMF) has been developed to combine multiple data sources to discover non-overlapping and class-specific RNA binding patterns. However, several challenges still exist in determining the number of latent dimensions in the factorization steps. In most circumstances, it is often assumed that the number of latent dimensions (or components) is given. Such trial-and-error procedures can be tedious in practice. In order to address this problem, differential evolution algorithm is proposed as the model selection method to choose the suitable number of ranks, which can adaptively decompose the input protein-RNA data matrix into different nonnegative components. Experimental results demonstrate that the proposed algorithms can improve the factorization quality over the recent state-of-the-arts. The effectiveness of the proposed algorithms are supported by comprehensive performance benchmarking on 31 genome-wide cross-linking immunoprecipitation (CLIP) coupled with high-throughput sequencing (CLIP-seq) datasets. In addition, time complexity analysis and parameter analysis are conducted to demonstrate the robustness of the proposed methods. Xiangtao Li, Ka-Chun Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | ToBio: Global Pathway Similarity Search Based on Topological and Biological FeaturesabstractPathway similarity search plays a vital role in the post-genomics era. Unfortunately, pathway similarity search involves the graph isomorphism problem which is NP-complete. Therefore, efficient search algorithms are desirable. In this work, we propose a novel global pathway similarity search approach named ToBio, which considers both topological and biological features for effective global pathway similarity search. Specifically, as motivated from nature, various topological and biological features including subgraph signature similarities, sequence similarities, and gene ontology similarities are considered in ToBio. Since different features carry different functional importance and dependences, we report three schemes of ToBio using different sets of features. In addition, to enhance the existing search algorithms for rigorous comparisons, post-processing pipelines are also proposed to investigate how different features can contribute to the search performance. ToBio and other state-of-the-art methods are benchmarked on the gold-standard pathway datasets from three species. The results demonstrate the competitive edges of ToBio over the state-of-the-arts ranging from the topological aspects to the biological aspects. Case studies have been conducted to reveal mechanistic insights into the unique search performance of ToBio. Jiao Zhang 0003, Sam Kwong, Ka-Chun Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | Evolutionary Multiobjective Clustering and Its Applications to Patient StratificationabstractPatient stratification has a major role in enabling efficient and personalized medicine. An important task in patient stratification is to discover disease subtypes for effective treatment. To achieve this goal, the research on clustering algorithms for patient stratification has brought attention from both academia and medical community over the past decades. However, existing clustering algorithms suffer from realistic restrictions such as experimental noises, high dimensionality, and poor interpretability. In particular, the existing clustering algorithms usually determine clustering quality using only one internal evaluation function. Unfortunately, it is obvious that one internal evaluation function is hard to be fitted and robust for all datasets. Therefore, in this paper, a novel multiobjective framework called multiobjective clustering algorithm by fast search and find of density peaks is proposed to address those limitations altogether. In the proposed framework, a parameter candidate population is evolved under multiple objectives to select features and evaluate clustering densities automatically. To guide the multiobjective evolution, five cluster validity indices including compactness, separation, Calinski-Harabasz index, Davies-Bouldin index, and Dunn index, are chosen as the objective functions, capturing multiple characteristics of the evolving clusters. Multiobjective differential evolution algorithm based on decomposition is adopted to optimize those five objective functions simultaneously. To demonstrate its effectiveness, extensive experiments have been conducted, comparing the proposed algorithm with 45 algorithms including nine state-of-the-art clustering algorithms, five multiobjective evolutionary algorithms, and 31 baseline algorithms under different objective subsets on 94 datasets featuring 35 real patient stratification datasets, 55 synthetic datasets based on a real human transcription regulation network model, and four other medical datasets. The numerical results reveal that the proposed algorithm can achieve better or competitive solutions than the others. Besides, time complexity analysis, convergence analysis, and parameter analysis are conducted to demonstrate the robustness of the proposed algorithm from different perspectives. Xiangtao Li, Ka-Chun Wong |
IEEE Trans. Cybern. | 2 |
| 2019 | A Clustering-Based Evolutionary Algorithm for Many-Objective Optimization ProblemsabstractThis paper suggests a novel clustering-based evolutionary algorithm for many-objective optimization problems. Its main idea is to classify the population into a number of clusters, which is expected to solve the difficulty of balancing convergence and diversity in high-dimensional objective space. The individuals showing high similarities on the vector angles are gathered into the same cluster, such that the population’s distribution can be well portrayed by the clusters. To efficiently find these clusters, partitional clustering is first used to classify the union population into${m}$main clusters based on the${m}$axis vectors (${m}$is the number of objectives), and then hierarchical clustering is further run on these${m}$main clusters to get${N}$final clusters (${N}$is the population size and${N>m}$). At last, in environmental selection, one individual from each of${N}$clusters closest to the axis vectors is selected to maintain diversity, while one individual from each of the other clusters is preferred by a simple convergence indicator to ensure convergence. When tackling some well-known test problems with 5–15 objectives, extensive experiments validate the superiority of our algorithm over six competitive many-objective EAs, especially on problems with incomplete and irregular Pareto-optimal fronts. Qiuzhen Lin, Songbai Liu, Ka-Chun Wong, Maoguo Gong, Carlos A. Coello Coello, Jianyong Chen, Jun Zhang 0003 |
IEEE Trans. Evol. Comput. | 3 |
| 2019 | An Effective Ensemble Framework for Multiobjective OptimizationabstractThis paper proposes an effective ensemble framework (EF) for tackling multiobjective optimization problems, by combining the advantages of various evolutionary operators and selection criteria that are run on multiple populations. A simple ensemble algorithm is realized as a prototype to demonstrate our proposed framework. Two mechanisms, namely competition and cooperation, are employed to drive the running of the ensembles. Competition is designed by adaptively running different evolutionary operators on multiple populations. The operator that better fits the problem’s characteristics will receive more computational resources, being rewarded by a decomposition-based credit assignment strategy. Cooperation is achieved by a cooperative selection of the offspring generated by different populations. In this way, the promising offspring from one population have chances to migrate into the other populations to enhance their convergence or diversity. Moreover, the population update information is further exploited to build an evolutionary potentiality model, which is used to guide the evolutionary process. Our experimental results show the superior performance of our proposed ensemble algorithms in solving most cases of a set of 31 test problems, which corroborates the advantages of our EF. Wenjun Wang 0003, Shaoqiang Yang, Qiuzhen Lin, Qingfu Zhang 0001, Ka-Chun Wong, Carlos A. Coello Coello, Jianyong Chen |
IEEE Trans. Evol. Comput. | 5 |
| 2019 | PathEmb: Random Walk Based Document Embedding for Global Pathway Similarity SearchabstractPathway analysis is a cornerstone of system biology. In particular, pathway similarity search plays a key role in establishing structural, functional, and evolutionary relationships between different biological entities. Given a query pathway as well as a database, a pathway similarity search aims to identify novel pathways that are homologous to the query pathway. Unfortunately, the pathway similarity search is computationally inefficient due to the NP-complete graph isomorphism problem. In this study, we introduce a novel algorithmic framework for pathway similarity search, named PathEmb (Pathway Embedding), which is analogous to the Skip-gram model where each pathway is represented as a "document." PathEmb exploits a second order random walk strategy to explore diverse pathway patterns. All signaling paths traversed from random walks are regarded as "sentences," which are constituted as a "document" afterwards. Then, the "document" pattern for the individual pathway is mapped into a low-dimensional feature space for downstream tasks. Furthermore, PathEmb is a topology-free pathway similarity search algorithm, which is feasible to handle any pathway with arbitrary structure. We have extensively evaluated PathEmb and other cutting-edge methods on three pathway datasets. The experimental results demonstrate that PathEmb outperforms the existing methods in terms of computational efficiency and search accuracy. The source code of PathEmb are freely available online https://github.com/zhangjiaobxy/PathEmb. Jiao Zhang 0003, Sam Kwong, Qiuzhen Lin, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 5 |
| 2018 | Off-target predictions in CRISPR-Cas9 gene editing using deep learningabstractMotivation: The prediction of off-target mutations in CRISPR-Cas9 is a hot topic due to its relevance to gene editing research. Existing prediction methods have been developed; however, most of them just calculated scores based on mismatches to the guide sequence in CRISPR-Cas9. Therefore, the existing prediction methods are unable to scale and improve their performance with the rapid expansion of experimental data in CRISPR-Cas9. Moreover, the existing methods still cannot satisfy enough precision in off-target predictions for gene editing at the clinical level. Results: To address it, we design and implement two algorithms using deep neural networks to predict off-target mutations in CRISPR-Cas9 gene editing (i.e. deep convolutional neural network and deep feedforward neural network). The models were trained and tested on the recently released off-target dataset, CRISPOR dataset, for performance benchmark. Another off-target dataset identified by GUIDE-seq was adopted for additional evaluation. We demonstrate that convolutional neural network achieves the best performance on CRISPOR dataset, yielding an average classification area under the ROC curve (AUC) of 97.2% under stratified 5-fold cross-validation. Interestingly, the deep feedforward neural network can also be competitive at the average AUC of 97.0% under the same setting. We compare the two deep neural network models with the state-of-the-art off-target prediction methods (i.e. CFD, MIT, CROP-IT, and CCTop) and three traditional machine learning models (i.e. random forest, gradient boosting trees, and logistic regression) on both datasets in terms of AUC values, demonstrating the competitive edges of the proposed algorithms. Additional analyses are conducted to investigate the underlying reasons from different perspectives. Availability and implementation: The example code are available at https://github.com/MichaelLinn/off_target_prediction. The related datasets are available at https://github.com/MichaelLinn/off_target_prediction/tree/master/data. Jiecong Lin, Ka-Chun Wong |
Bioinform. | 2 |
| 2018 | A scalable community detection algorithm for large graphs using stochastic block models
Chengbin Peng 0001, Ka-Chun Wong, Xiangliang Zhang 0001, David E. Keyes |
Intell. Data Anal. | 3 |
| 2018 | Adaptive multiple-elites-guided composite differential evolution algorithm with a shift mechanism
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Qiuzhen Lin, Ka-Chun Wong, Jianyong Chen, Jian Lu 0002 |
Inf. Sci. | 5 |
| 2018 | An adaptive immune-inspired multi-objective algorithm with multiple differential evolution strategies
Qiuzhen Lin, Yueping Ma, Jianyong Chen, Qingling Zhu, Carlos A. Coello Coello, Ka-Chun Wong, Fei Chen 0003 |
Inf. Sci. | 6 |
| 2018 | Special issue: the International Conference on Soft Computing and Machine Intelligence
Suash Deb, Thomas Hanne, Ka-Chun Wong |
Soft Comput. | 3 |
| 2018 | A Comparative Study for Identifying the Chromosome-Wide Spatial Clusters from High-Throughput Chromatin Conformation Capture DataabstractIn the past years, the high-throughput sequencing technologies have enabled massive insights into genomic annotations. In contrast, the full-scale three-dimensional arrangements of genomic regions are relatively unknown. Thanks to the recent breakthroughs in High-throughput Chromosome Conformation Capture (Hi-C) techniques, non-negative matrix factorization (NMF) has been adopted to identify local spatial clusters of genomic regions from Hi-C data. However, such non-negative matrix factorization entails a high-dimensional non-convex objective function to be optimized with non-negative constraints. We propose and compare more than ten optimization algorithms to improve the identification of local spatial clusters via NMF. To circumvent and optimize the high-dimensional, non-convex, and constrained objective function, we draw inspiration from the nature to perform in silico evolution. The proposed algorithms consist of a population of candidates to be evolved while the NMF acts as local search during the evolutions. The population based optimization algorithm coordinates and guides the non-negative matrix factorization toward global optima. Experimental results show that the proposed algorithms can improve the quality of non-negative matrix factorization over the recent state-of-the-arts. The effectiveness and robustness of the proposed algorithms are supported by comprehensive performance benchmarking on chromosome-wide Hi-C contact maps of yeast and human. In addition, time complexity analysis, convergence analysis, parameter analysis, biological case studies, and gene ontology similarity analysis are conducted to demonstrate the robustness of the proposed methods from different perspectives. Xiangtao Li, Ka-Chun Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | A Diversity-Enhanced Resource Allocation Strategy for Decomposition-Based Multiobjective Evolutionary AlgorithmabstractThe multiobjective evolutionary algorithm (MOEA) based on decomposition transforms a multiobjective optimization problem into a set of aggregated subproblems and then optimizes them collaboratively. Since these subproblems usually have different degrees of difficulty, resource allocation (RA) strategies have been reported to enhance performance, attempting to dynamically assign proper amounts of computational resources for the solution of each of these subproblems. However, existing schemes for decomposition-based MOEAs fully rely on the relative improvement of the aggregated functions to do this. This paper proposes a diversity-enhanced RA strategy for this kind of MOEA, depending on both relative improvement on aggregated function value and solution density around each subproblem to assign computational resources. Thus, one subproblem surrounded with fewer solutions in its neighboring area and more relative improvement on the aggregated function value will be allocated a higher probability for evolution. Our experimental results show the advantages of our proposed strategy over two popular RA strategies available for decomposition-based MOEAs, on tackling a set of complicated benchmark problems. Qiuzhen Lin, Genmiao Jin, Yueping Ma, Ka-Chun Wong, Carlos A. Coello Coello, Jianqiang Li 0001, Jianyong Chen, Jun Zhang 0003 |
IEEE Trans. Cybern. | 4 |
| 2018 | Particle Swarm Optimization With a Balanceable Fitness Estimation for Many-Objective Optimization ProblemsabstractRecently, it was found that most multiobjective particle swarm optimizers (MOPSOs) perform poorly when tackling many-objective optimization problems (MaOPs). This is mainly because the loss of selection pressure that occurs when updating the swarm. The number of nondominated individuals is substantially increased and the diversity maintenance mechanisms in MOPSOs always guide the particles to explore sparse regions of the search space. This behavior results in the final solutions being distributed loosely in objective space, but far away from the true Pareto-optimal front. To avoid the above scenario, this paper presents a balanceable fitness estimation method and a novel velocity update equation, to compose a novel MOPSO (NMPSO), which is shown to be more effective to tackle MaOPs. Moreover, an evolutionary search is further run on the external archive in order to provide another search pattern for evolution. The DTLZ and WFG test suites with 4-10 objectives are used to assess the performance of NMPSO. Our experiments indicate that NMPSO has superior performance over four current MOPSOs, and over four competitive multiobjective evolutionary algorithms (SPEA2-SDE, NSGA-III, MOEA/DD, and SRA), when solving most of the test problems adopted. Qiuzhen Lin, Songbai Liu, Qingling Zhu, Chaoyu Tang, Ruizhen Song, Jianyong Chen, Carlos A. Coello Coello, Ka-Chun Wong, Jun Zhang 0003 |
IEEE Trans. Evol. Comput. | 8 |
| 2018 | Multiobjective Patient Stratification Using Evolutionary Multiobjective OptimizationabstractOne of the main challenges in modern medic-ine is to stratify patients for personalized care. Many different clustering methods have been proposed to solve the problem in both quantitative and biologically meaningful manners. However, existing clustering algorithms suffer from numerous restrictions such as experimental noises, high dimensionality, and poor interpretability. To overcome those limitations altogether, we propose and formulate a multiobjective framework based on evolutionary multiobjective optimization to balance the feature relevance and redundancy for patient stratification. To demonstrate the effectiveness of our proposed algorithms, we benchmark our algorithms across 55 synthetic datasets based on a real human transcription regulation network model, 35 real cancer gene expression datasets, and two case studies. Experimental results suggest that the proposed algorithms perform better than the recent state-of-the-arts. In addition, time complexity analysis, convergence analysis, and parameter analysis are conducted to demonstrate the robustness of the proposed methods from different perspectives. Finally, the t-Distributed Stochastic Neighbor Embedding (t-SNE) is applied to project the selected feature subsets onto two or three dimensions to visualize the high-dimensional patient stratification data. Xiangtao Li, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | MotifHyades: expectation maximization for de novo DNA motif pair discovery on paired sequencesabstractMOTIVATION: In higher eukaryotes, protein-DNA binding interactions are the central activities in gene regulation. In particular, DNA motifs such as transcription factor binding sites are the key components in gene transcription. Harnessing the recently available chromatin interaction data, computational methods are desired for identifying the coupling DNA motif pairs enriched on long-range chromatin-interacting sequence pairs (e.g. promoter-enhancer pairs) systematically. RESULTS: To fill the void, a novel probabilistic model (namely, MotifHyades) is proposed and developed for de novo DNA motif pair discovery on paired sequences. In particular, two expectation maximization algorithms are derived for efficient model training with linear computational complexity. Under diverse scenarios, MotifHyades is demonstrated faster and more accurate than the existing ad hoc computational pipeline. In addition, MotifHyades is applied to discover thousands of DNA motif pairs with higher gold standard motif matching ratio, higher DNase accessibility and higher evolutionary conservation than the previous ones in the human K562 cell line. Lastly, it has been run on five other human cell lines (i.e. GM12878, HeLa-S3, HUVEC, IMR90, and NHEK), revealing another thousands of novel DNA motif pairs which are characterized across a broad spectrum of genomic features on long-range promoter-enhancer pairs. AVAILABILITY AND IMPLEMENTATION: The matrix-algebra-optimized versions of MotifHyades and the discovered DNA motif pairs can be found in http://bioinfo.cs.cityu.edu.hk/MotifHyades. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ka-Chun Wong |
Bioinform. | 1 |
| 2017 | NSSRF: global network similarity search with subgraph signatures and its applicationsabstractMOTIVATION: The exponential growth of biological network database has increasingly rendered the global network similarity search (NSS) computationally intensive. Given a query network and a network database, it aims to find out the top similar networks in the database against the query network based on a topological similarity measure of interest. With the advent of big network data, the existing search methods may become unsuitable since some of them could render queries unsuccessful by returning empty answers or arbitrary query restrictions. Therefore, the design of NSS algorithm remains challenging under the dilemma between accuracy and efficiency. RESULTS: We propose a global NSS method based on regression, denotated as NSSRF, which boosts the search speed without any significant sacrifice in practical performance. As motivated from the nature, subgraph signatures are heavily involved. Two phases are proposed in NSSRF: offline model building phase and similarity query phase. In the offline model building phase, the subgraph signatures and cosine similarity scores are used for efficient random forest regression (RFR) model training. In the similarity query phase, the trained regression model is queried to return similar networks. We have extensively validated NSSRF on biological pathways and molecular structures; NSSRF demonstrates competitive performance over the state-of-the-arts. Remarkably, NSSRF works especially well for large networks, which indicates that the proposed approach can be promising in the era of big data. Case studies have proven the efficiencies and uniqueness of NSSRF which could be missed by the existing state-of-the-arts. AVAILABILITY AND IMPLEMENTATION: The source code of two versions of NSSRF are freely available for downloading at https://github.com/zhangjiaobxy/nssrfBinary and https://github.com/zhangjiaobxy/nssrfPackage . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiao Zhang 0003, Sam Kwong, Yuheng Jia, Ka-Chun Wong |
Bioinform. | 4 |
| 2017 | A scalable community detection algorithm for large graphs using stochastic block modelsabstractCommunity detection in graphs is widely used in social and biological networks, and the stochastic block model is a powerful probabilistic tool for describing graphs with community structures. However, in the era of “big data”, traditional inference algorithms for such a model are increasingly limi ted due to their high time complexity and poor scalability. In this paper, we propose a multi-stage maximum likelihood approach to recover the latent parameters of the stochastic block model, in time linear with respect to the number of edges. We also propose a parallel algorithm based on message passing. Our algorithm can overlap communication and computation, providing speedup without compromising accuracy as the number of processors grows. For example, to process a real-world graph with about 1.3 million nodes and 10 million edges, our algorithm requires about 6 seconds on 64 cores of a contemporary commodity Linux cluster. Experiments demonstrate that the algorithm can produce high quality results on both benchmark and real-world graphs. An example of finding more meaningful communities is illustrated consequently in comparison with a popular modularity maximization algorithm. Chengbin Peng 0001, Ka-Chun Wong, Xiangliang Zhang 0001, David E. Keyes |
Intell. Data Anal. | 3 |
| 2017 | A novel artificial bee colony algorithm with an adaptive population size for numerical function optimization
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Qiuzhen Lin, Zhenkun Wen, Ka-Chun Wong, Jianyong Chen |
Inf. Sci. | 7 |
| 2017 | Elucidating high-dimensional cancer hallmark annotation via enriched ontology
Shankai Yan, Ka-Chun Wong |
J. Biomed. Informatics | 2 |
| 2017 | Evolving Transcription Factor Binding Site Models From Protein Binding Microarray DataabstractProtein binding microarray (PBM) is a high-throughput platform that can measure the DNA binding preference of a protein in a comprehensive and unbiased manner. In this paper, we describe the PBM motif model building problem. We apply several evolutionary computation methods and compare their performance with the interior point method, demonstrating their performance advantages. In addition, given the PBM domain knowledge, we propose and describe a novel method called kmerGA which makes domain-specific assumptions to exploit PBM data properties to build more accurate models than the other models built. The effectiveness and robustness of kmerGA is supported by comprehensive performance benchmarking on more than 200 datasets, time complexity analysis, convergence analysis, parameter analysis, and case studies. To demonstrate its utility further, kmerGA is applied to two real world applications: 1) PBM rotation testing and 2) ChIP-Seq peak sequence prediction. The results support the biological relevance of the models learned by kmerGA, and thus its real world applicability. Ka-Chun Wong, Chengbin Peng 0001, Yue Li 0017 |
IEEE Trans. Cybern. | 1 |
| 2017 | An External Archive-Guided Multiobjective Particle Swarm Optimization AlgorithmabstractThe selection of swarm leaders (i.e., the personal best and global best), is important in the design of a multiobjective particle swarm optimization (MOPSO) algorithm. Such leaders are expected to effectively guide the swarm to approach the true Pareto optimal front. In this paper, we present a novel external archive-guided MOPSO algorithm (AgMOPSO), where the leaders for velocity update are all selected from the external archive. In our algorithm, multiobjective optimization problems (MOPs) are transformed into a set of subproblems using a decomposition approach, and then each particle is assigned accordingly to optimize each subproblem. A novel archive-guided velocity update method is designed to guide the swarm for exploration, and the external archive is also evolved using an immune-based evolutionary strategy. These proposed approaches speed up the convergence of AgMOPSO. The experimental results fully demonstrate the superiority of our proposed AgMOPSO in solving most of the test problems adopted, in terms of two commonly used performance measures. Moreover, the effectiveness of our proposed archive-guided velocity update method and immune-based evolutionary strategy is also experimentally validated on more than 30 test MOPs. Qingling Zhu, Qiuzhen Lin, Weineng Chen, Ka-Chun Wong, Carlos A. Coello Coello, Jianqiang Li 0001, Jianyong Chen, Jun Zhang 0003 |
IEEE Trans. Cybern. | 4 |
| 2016 | A cone order sequence based multi-objective evolutionary algorithmabstractA cone order sequence based MOEA (CS-MOEA) is proposed to deal with the multi-objective optimization problems. Instead of only using the Pareto dominance, it constructs a sequence of cone order to balance the search diversity and convergence. By gradually increasing the open angle of the cone order, it approximates the Pareto cone gradually. A simple formula for judging the θ-cone dominance is derived, which is easy to be computed. Moreover, an energy model is introduced for the selection of individuals to maintain population diversity. Experiments on more than 10 problems (i.e. zdt and dtlz benchmark problem sets) demonstrate that the proposed method is competitive, compared with Stable Matching MOEA/D (STM-MOEA/D) and MOEA/D-DE. Yueming Lyu, Qingfu Zhang 0001, Ka-Chun Wong |
CEC | 3 |
| 2016 | Identification of coupling DNA motif pairs on long-range chromatin interactions in human K562 cellsabstractMOTIVATION: The protein-DNA interactions between transcription factors (TFs) and transcription factor binding sites (TFBSs, also known as DNA motifs) are critical activities in gene transcription. The identification of the DNA motifs is a vital task for downstream analysis. Unfortunately, the long-range coupling information between different DNA motifs is still lacking. To fill the void, as the first-of-its-kind study, we have identified the coupling DNA motif pairs on long-range chromatin interactions in human. RESULTS: The coupling DNA motif pairs exhibit substantially higher DNase accessibility than the background sequences. Half of the DNA motifs involved are matched to the existing motif databases, although nearly all of them are enriched with at least one gene ontology term. Their motif instances are also found statistically enriched on the promoter and enhancer regions. Especially, we introduce a novel measurement called motif pairing multiplicity which is defined as the number of motifs that are paired with a given motif on chromatin interactions. Interestingly, we observe that motif pairing multiplicity is linked to several characteristics such as regulatory region type, motif sequence degeneracy, DNase accessibility and pairing genomic distance. Taken into account together, we believe the coupling DNA motif pairs identified in this study can shed lights on the gene transcription mechanism under long-range chromatin interactions. AVAILABILITY AND IMPLEMENTATION: The identified motif pair data is compressed and available in the supplementary materials associated with this manuscript. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001 |
Bioinform. | 1 |
| 2016 | A Comparison Study for DNA Motif Modeling on Protein Binding MicroarrayabstractTranscription factor binding sites (TFBSs) are relatively short (5-15 bp) and degenerate. Identifying them is a computationally challenging task. In particular, protein binding microarray (PBM) is a high-throughput platform that can measure the DNA binding preference of a protein in a comprehensive and unbiased manner; for instance, a typical PBM experiment can measure binding signal intensities of a protein to all possible DNA k-mers (k = 8∼10). Since proteins can often bind to DNA with different binding intensities, one of the major challenges is to build TFBS (also known as DNA motif) models which can fully capture the quantitative binding affinity data. To learn DNA motif models from the non-convex objective function landscape, several optimization methods are compared and applied to the PBM motif model building problem. In particular, representative methods from different optimization paradigms have been chosen for modeling performance comparison on hundreds of PBM datasets. The results suggest that the multimodal optimization methods are very effective for capturing the binding preference information from PBM data. In particular, we observe a general performance improvement if choosing di-nucleotide modeling over mono-nucleotide modeling. In addition, the models learned by the best-performing method are applied to two independent applications: PBM probe rotation testing and ChIP-Seq peak sequence prediction, demonstrating its biological applicability. Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001, Hau-San Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | A Scalable Community Detection Algorithm for Large Graphs Using Stochastic Block Models
Chengbin Peng 0001, Ka-Chun Wong, Xiangliang Zhang 0001, David E. Keyes |
IJCAI | 3 |
| 2015 | Active Learning Based on Single-Hidden Layer Feed-Forward Neural NetworkabstractIn this paper, we propose two stream-based active learning algorithms for single-hidden layer feed-forward neural networks (SLFNs) trained by extreme learning machine (ELM). Uncertainty and inconsistency are adopted as two sample selection criteria. Uncertainty reflects the nondeterminacy of a sample among different decision classes, which is calculated by information entropy or Gini-index. Inconsistency reflects the disagreement of the sample between its conditional features and decision labels, which is calculated by the lower approximations in fuzzy rough sets. Experimental results demonstrate that inconsistency-based strategy is more effective than uncertainty based strategy for SLFNs under stream-based environment. Ran Wang 0001, Sam Kwong, Qingshan Jiang, Ka-Chun Wong |
SMC | 4 |
| 2015 | Data Analytics for Protein-DNA Binding InteractionsabstractDetermining the protein-DNA binding specificity is an important step in understanding genetic codes. With a large amount of protein-DNA complexes, mature statistical and data mining techniques, and efficient computational power, a fundamental and comprehensive protein-DNA binding sequence analysis is conducted and described in this work. In particular, two different types of analysis are proposed and described. Firstly, statistical analysis is conducted to give holistic insights into the protein-DNA binding sequences. Secondly, data mining techniques are applied to extract interesting sequence patterns which takes into account both sides (protein and DNA sides). The results demonstrate that there are statistically enriched sequence patterns among the protein-DNA binding sequences. Nonetheless, it also confirms that there is not any general principle in protein-DNA binding in a big data analytics manner. To address that, contemporary data mining methods are introduced to discover advanced sequence patterns. The patterns are validated with an external database, revealing biological insights into protein-DNA binding interactions. Ka-Chun Wong |
SMC | 1 |
| 2015 | SignalSpider: probabilistic pattern discovery on multiple normalized ChIP-Seq signal profilesabstractMOTIVATION: Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-Seq) measures the genome-wide occupancy of transcription factors in vivo. Different combinations of DNA-binding protein occupancies may result in a gene being expressed in different tissues or at different developmental stages. To fully understand the functions of genes, it is essential to develop probabilistic models on multiple ChIP-Seq profiles to decipher the combinatorial regulatory mechanisms by multiple transcription factors. RESULTS: In this work, we describe a probabilistic model (SignalSpider) to decipher the combinatorial binding events of multiple transcription factors. Comparing with similar existing methods, we found SignalSpider performs better in clustering promoter and enhancer regions. Notably, SignalSpider can learn higher-order combinatorial patterns from multiple ChIP-Seq profiles. We have applied SignalSpider on the normalized ChIP-Seq profiles from the ENCODE consortium and learned model instances. We observed different higher-order enrichment and depletion patterns across sets of proteins. Those clustering patterns are supported by Gene Ontology (GO) enrichment, evolutionary conservation and chromatin interaction enrichment, offering biological insights for further focused studies. We also proposed a specific enrichment map visualization method to reveal the genome-wide transcription factor combinatorial patterns from the models built, which extend our existing fine-scale knowledge on gene regulation to a genome-wide level. AVAILABILITY AND IMPLEMENTATION: The matrix-algebra-optimized executables and source codes are available at the authors' websites: http://www.cs.toronto.edu/∼wkc/SignalSpider. Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001, Zhaolei Zhang |
Bioinform. | 1 |
| 2015 | Probabilistic Inference on Multiple Normalized Signal Profiles from Next Generation Sequencing: Transcription Factor Binding SitesabstractWith the prevalence of chromatin immunoprecipitation (ChIP) with sequencing (ChIP-Seq) technology, massive ChIP-Seq data has been accumulated. The ChIP-Seq technology measures the genome-wide occupancy of DNA-binding proteins in vivo. It is well-known that different DNA-binding protein occupancies may result in a gene being regulated in different conditions (e.g. different cell types). To fully understand a gene's function, it is essential to develop probabilistic models on multiple ChIP-Seq profiles for deciphering the gene transcription causalities. In this work, we propose and describe two probabilistic models. Assuming the conditional independence of different DNA-binding proteins' occupancies, the first method (SignalRanker) is developed as an intuitive method for ChIP-Seq genome-wide signal profile inference. Unfortunately, such an assumption may not always hold in some gene regulation cases. Thus, we propose and describe another method (FullSignalRanker) which does not make the conditional independence assumption. The proposed methods are compared with other existing methods on ENCODE ChIP-Seq datasets, demonstrating its regression and classification ability. The results suggest that FullSignalRanker is the best-performing method for recovering the signal ranks on the promoter and enhancer regions. In addition, FullSignalRanker is also the best-performing method for peak sequence classification. We envision that SignalRanker and FullSignalRanker will become important in the era of next generation sequencing. FullSignalRanker program is available on the following website: http://www.cs.toronto.edu/~wkc/FullSignalRanker/. Ka-Chun Wong, Chengbin Peng 0001, Yue Li 0017 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | A probabilistic approach to explore human miRNA targetome by integrating miRNA-overexpression data and sequence informationabstractMOTIVATION: Systematic identification of microRNA (miRNA) targets remains a challenge. The miRNA overexpression coupled with genome-wide expression profiling is a promising new approach and calls for a new method that integrates expression and sequence information. RESULTS: We developed a probabilistic scoring method called targetScore. TargetScore infers miRNA targets as the transformed fold-changes weighted by the Bayesian posteriors given observed target features. To this end, we compiled 84 datasets from Gene Expression Omnibus corresponding to 77 human tissue or cells and 113 distinct transfected miRNAs. Comparing with other methods, targetScore achieves significantly higher accuracy in identifying known targets in most tests. Moreover, the confidence targets from targetScore exhibit comparable protein downregulation and are more significantly enriched for Gene Ontology terms. Using targetScore, we explored oncomir-oncogenes network and predicted several potential cancer-related miRNA-messenger RNA interactions. AVAILABILITY AND IMPLEMENTATION: TargetScore is available at Bioconductor: http://www.bioconductor.org/packages/devel/bioc/html/TargetScore.html. Yue Li 0017, Anna Goldenberg, Ka-Chun Wong, Zhaolei Zhang |
Bioinform. | 3 |
| 2014 | Mirsynergy: detecting synergistic miRNA regulatory modules by overlapping neighbourhood expansionabstractMOTIVATION: Identification of microRNA regulatory modules (MiRMs) will aid deciphering aberrant transcriptional regulatory network in cancer but is computationally challenging. Existing methods are stochastic or require a fixed number of regulatory modules. RESULTS: We propose Mirsynergy, an efficient deterministic overlapping clustering algorithm adapted from a recently developed framework. Mirsynergy operates in two stages: it first forms MiRMs based on co-occurring microRNA (miRNA) targets and then expands each MiRM by greedily including (excluding) mRNAs into (from) the MiRM to maximize the synergy score, which is a function of miRNA-mRNA and gene-gene interactions. Using expression data for ovarian, breast and thyroid cancer from The Cancer Genome Atlas, we compared Mirsynergy with internal controls and existing methods. Mirsynergy-MiRMs exhibit significantly higher functional enrichment and more coherent miRNA-mRNA expression anti-correlation. Based on Kaplan-Meier survival analysis, we proposed several prognostically promising MiRMs and envisioned their utility in cancer research. AVAILABILITY AND IMPLEMENTATION: Mirsynergy is implemented/available as an R/Bioconductor package at www.cs.utoronto.ca/∼yueli/Mirsynergy.html. Yue Li 0017, Cheng Liang 0001, Ka-Chun Wong, Jiawei Luo 0001, Zhaolei Zhang |
Bioinform. | 3 |
| 2014 | SNPdryad: predicting deleterious non-synonymous human SNPs using only orthologous protein sequencesabstractMOTIVATION: The recent advances in genome sequencing have revealed an abundance of non-synonymous polymorphisms among human individuals; subsequently, it is of immense interest and importance to predict whether such substitutions are functional neutral or have deleterious effects. The accuracy of such prediction algorithms depends on the quality of the multiple-sequence alignment, which is used to infer how an amino acid substitution is tolerated at a given position. Because of the scarcity of orthologous protein sequences in the past, the existing prediction algorithms all include sequences of protein paralogs in the alignment, which can dilute the conservation signal and affect prediction accuracy. However, we believe that, with the sequencing of a large number of mammalian genomes, it is now feasible to include only protein orthologs in the alignment and improve the prediction performance. RESULTS: We have developed a novel prediction algorithm, named SNPdryad, which only includes protein orthologs in building a multiple sequence alignment. Among many other innovations, SNPdryad uses different conservation scoring schemes and uses Random Forest as a classifier. We have tested SNPdryad on several datasets. We found that SNPdryad consistently outperformed other methods in several performance metrics, which is attributed to the exclusion of paralogous sequence. We have run SNPdryad on the complete human proteome, generating prediction scores for all the possible amino acid substitutions. AVAILABILITY AND IMPLEMENTATION: The algorithm and the prediction results can be accessed from the Web site: http://snps.ccbr.utoronto.ca:8080/SNPdryad/ CONTACT: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Ka-Chun Wong, Zhaolei Zhang |
Bioinform. | 1 |
| 2012 | Multiplicative Algorithms for Constrained Non-negative Matrix FactorizationabstractNon-negative matrix factorization (NMF) provides the advantage of parts-based data representation through additive only combinations. It has been widely adopted in areas like item recommending, text mining, data clustering, speech denoising, etc. In this paper, we provide an algorithm that allows the factorization to have linear or approximately linear constraints with respect to each factor. We prove that if the constraint function is linear, algorithms within our multiplicative framework will converge. This theory supports a large variety of equality and inequality constraints, and can facilitate application of NMF to a much larger domain. Taking the recommender system as an example, we demonstrate how a specialized weighted and constrained NMF algorithm can be developed to fit exactly for the problem, and the tests justify that our constraints improve the performance for both weighted and unweighted NMF algorithms under several different metrics. In particular, on the Movie lens data with 94% of items, the Constrained NMF improves recall rate 3% compared to SVD50 and 45% compared to SVD150, which were reported as the best two in the top-N metric. Chengbin Peng 0001, Ka-Chun Wong, Alyn P. Rockwood, Xiangliang Zhang 0001, Jinling Jiang, David E. Keyes |
ICDM | 2 |
| 2012 | A novel web-based system for tropical cyclone analysis and predictionabstractA web-based system is developed for the analysis and prediction of tropical cyclones, particularly their landfalls and recurvatures. To facilitate accessibility to the system, its development is based on Google Maps application programming interface (API), Java and client/server architecture. In addition to the construction of a powerful query system for the multi-source, multi-scale and multi-level tropical cyclone database, data mining approach and dynamic modelling approach have been implemented and integrated for effective and efficient analysis, prediction and visualization of tropical cyclone movements. The system can be accessed worldwide by researchers, professionals and the general public. It is thus a powerful system for research, real-life application and knowledge dissemination. Its extensibility and user-friendliness pave the road for further development and enable more in-depth analysis and real-time operation. Yee Leung, Man Hon Wong 0001, Ka-Chun Wong, Wei Zhang 0048, Kwong-Sak Leung |
Int. J. Geogr. Inf. Sci. | 3 |
| 2012 | Evolutionary multimodal optimization using the principle of locality
Ka-Chun Wong, Chun-Ho Wu, Ricky K. P. Mok, Chengbin Peng 0001, Zhaolei Zhang |
Inf. Sci. | 1 |
| 2011 | Discovering approximate-associated sequence patterns for protein-DNA interactionsabstractMOTIVATION: The bindings between transcription factors (TFs) and transcription factor binding sites (TFBSs) are fundamental protein-DNA interactions in transcriptional regulation. Extensive efforts have been made to better understand the protein-DNA interactions. Recent mining on exact TF-TFBS-associated sequence patterns (rules) has shown great potentials and achieved very promising results. However, exact rules cannot handle variations in real data, resulting in limited informative rules. In this article, we generalize the exact rules to approximate ones for both TFs and TFBSs, which are essential for biological variations. RESULTS: A progressive approach is proposed to address the approximation to alleviate the computational requirements. Firstly, similar TFBSs are grouped from the available TF-TFBS data (TRANSFAC database). Secondly, approximate and highly conserved binding cores are discovered from TF sequences corresponding to each TFBS group. A customized algorithm is developed for the specific objective. We discover the approximate TF-TFBS rules by associating the grouped TFBS consensuses and TF cores. The rules discovered are evaluated by matching (verifying with) the actual protein-DNA binding pairs from Protein Data Bank (PDB) 3D structures. The approximate results exhibit many more verified rules and up to 300% better verification ratios than the exact ones. The customized algorithm achieves over 73% better verification ratios than traditional methods. Approximate rules (64-79%) are shown statistically significant. Detailed variation analysis and conservation verification on NCBI records demonstrate that the approximate rules reveal both the flexible and specific protein-DNA interactions accurately. The approximate TF-TFBS rules discovered show great generalized capability of exploring more informative binding rules. Tak-Ming Chan, Ka-Chun Wong, Kin-Hong Lee, Man Hon Wong 0001, Terrence Chi-Kong Lau, Stephen Kwok-Wing Tsui, Kwong-Sak Leung |
Bioinform. | 2 |
| 2011 | Generalizing and learning protein-DNA binding sequence representations by an evolutionary algorithm
Ka-Chun Wong, Chengbin Peng 0001, Man Hon Wong 0001, Kwong-Sak Leung |
Soft Comput. | 1 |
| 2010 | Effect of Spatial Locality on an Evolutionary Algorithm for Multimodal Optimization
Ka-Chun Wong, Kwong-Sak Leung, Man Hon Wong 0001 |
EvoApplications (1) | 1 |
| 2010 | Protein structure prediction on a lattice model via multimodal optimization techniquesabstractThis paper considers the protein structure prediction problem as a multimodal optimization problem. In particular, de novo protein structure prediction problems on the 3D Hydrophobic-Polar (HP) lattice model are tackled by evolutionary algorithms using multimodal optimization techniques. In addition, a new mutation approach and performance metric are proposed for the problem. The experimental results indicate that the proposed algorithms are more effective than the state-of-the-arts algorithms, even though they are simple. Ka-Chun Wong, Kwong-Sak Leung, Man Hon Wong 0001 |
GECCO | 1 |
| 2009 | An evolutionary algorithm with species-specific explosion for multimodal optimizationabstractThis paper presents an evolutionary algorithm, which we call Evolutionary Algorithm with Species-specific Explosion (EASE), for multimodal optimization. EASE is built on the Species Conserving Genetic Algorithm (SCGA), and the design is improved in several ways. In particular, it not only identifies species seeds, but also exploits the species seeds to create multiple mutated copies in order to further converge to the respective optimum for each species. Experiments were conducted to compare EASE and SCGA on four benchmark functions. Cross-comparison with recent rival techniques on another five benchmark functions was also reported. The results reveal that EASE has a competitive edge over the other algorithms tested. Ka-Chun Wong, Kwong-Sak Leung, Man Hon Wong 0001 |
GECCO | 1 |