EDBT 2026 Demo / reviewers in the wild / expert
Yushan Qiu
dblp:140/3580
· DBLP profile ↗
25ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-9393-3648ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMCC: A Novel Clustering Method for Single- and Multi-Omics Data Based on Co-Regularized Network FusionabstractClustering is a common technique for statistical data analysis and is essential for developing precision medicine. Numerous computational methods have been proposed for integrating multi-omics data to identify cancer subtypes. However, most existing clustering models based on network fusion fail to preserve the consistency of the distribution of the data before and after fusion. Motivated by this observation, we would like to measure and minimize the distribution difference between networks, which may not be in the same space, to improve the performance of data fusion. We were therefore motivated to develop a flexible clustering model, based on network fusion, that minimizes the distribution difference between the data before and after fusion by co-regularization; the model can be applied to both single- and multi-omics data. We propose a new network fusion model for single- and multi-omics data clustering for identifying cancer or cell subtypes based on co-regularized network fusion (SMCC). SMCC integrates low-rank subspace representation and entropy to fuse networks. In addition, it measures and minimizes the distribution difference between the similarity networks and the fusion network by co-regularization. The model can both reduce the noise interference in the source data and make the statistical characteristics of the fusion result closer to those of the source data. We evaluated the clustering performance of SMCC across 16 real single- and multi-omics dataset. The experimental results demonstrated that SMCC is superior to 17 state-of-the-art clustering methods. Moreover, it is effective for identifying cancer or cell subtypes, thereby promoting the development of precision medicine. Sha Tian, Yushan Qiu, Quan Zou 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | GSTRPCA: irregular tensor singular value decomposition for single-cell multi-omics data clusteringabstractSingle-cell multi-omics refers to the various types of biological data at the single-cell level. These data have enabled insight and resolution to cellular phenotypes, biological processes, and developmental stages. Current advances hold high potential for breakthroughs by integrating multiple different omics layers. However, singlecell multi-omics data usually have different feature dimensions and direct or indirect relationships. How to keep the data structure of these different data and extract hidden relationships is a major challenge for omics data integration, and effective integration models are urgently needed. In this paper, we propose an irregular tensor decomposition model (GSTRPCA) based on tensor robust principal component analysis (TRPCA). We developed a weighted threshold model for the decomposition of irregular tensor data by combining low-rank and sparsity constraints, which requires that the low-dimensional embeddings of the data remain lowrank and sparse. The major advantage of the GSTRPCA algorithm is its ability to keep the original data structure and explore hidden related features among omics data. For GSTRPCA, we also designed an effective algorithm that theoretically guarantees global convergence for the tensor decomposition. The computational experiments on irregular tensor datasets demonstrate that GSTRPCA significantly outperformed the state-of-the-art methods and hence confirm the superiority of GSTRPCA in clustering single-cell multiomics data. To our knowledge, this is the first tensor decomposition method for irregular tensor data to keep the data structure and hence improve the clustering performance for single-cell multi-omics data. GSTRPCA is a Matlabbased algorithm, and the code is available from https://github.com/GGL-B/GSTRPCA. Lu-Bin Cui, Guiliang Guo, Michael Kwok-Po Ng, Quan Zou 0001, Yushan Qiu |
Briefings Bioinform. | 5 |
| 2025 | Ensemble machine learning-based pre-trained annotation approach for scRNA-seq data using gradient boosting with genetic optimizerabstractSingle-cell RNA sequencing (scRNA-seq) has revolutionized the study of gene expression by allowing researchers to analyze the transcriptomes of individual cells. This technology provides unprecedented insights into cellular heterogeneity, cellular states, and biological processes at a single-cell resolution. The problem of single-cell RNA annotation involves assigning meaningful labels or annotations to each cell in the scRNA-seq dataset, indicating its corresponding cell type, state, or biological function. Current annotation methods are often challenged by limited source data quality, sensitivity to batch effects, and poor adaptability to uncharacterized cell types. We propose an ensemble machine learning-based pre-trained annotation framework that integrates gradient boosting and genetic optimization for robust feature selection. The proposed method uses ensemble learning to enhance annotation accuracy under data scarcity, addressing limitations in existing supervised methods by leveraging a combination of multiple annotated datasets and feature alignment strategies. Through comprehensive benchmarking across varied biological contexts, we demonstrate that the proposed approach significantly improves annotation accuracy and generalization across different scRNA-seq platforms, especially under conditions of reduced reference data. Results confirm its versatility and resilience in accurately annotating cell types, even under reduced data conditions, establishing it as a powerful tool for cell-type classification in scRNA-seq data. Osama Elnahas, Waleed M. Ead, Yushan Qiu |
BMC Bioinform. | 3 |
| 2025 | APDCA: An accurate and effective method for predicting associations between RBPs and AS-events during epithelial-mesenchymal transitionabstractMOTIVATION: Epithelial-mesenchymal transition (EMT) plays a key role in cancer metastasis by promoting changes in adhesion and motility. RNA-binding proteins (RBPs) regulate alternative splicing (AS) during EMT, enabling a single gene to produce multiple protein isoforms that affect tumor progression. Disruption of RBP-AS interactions may disrupt the progress of diseases like cancer. Despite the importance of RBP-AS relationships in EMT, few computational methods predict these associations. Existing models struggle in sparse settings with limited known associations. To improve performance, we incorporate both sparsity constraints and heterogeneous biological data to infer RBP-AS associations. RESULT: We propose a new method based on Accelerated Proximal DC Algorithm (APDCA) for predicting RBP-AS associations. In particular, APDCA combines sparse low-rank matrix factorization with a Difference-of-Convex (DC) optimization framework and uses extrapolation to improve convergence. A key feature of APDCA is the use of a sparsity constraint, which filters out noise and highlights key associations. In addition, integrating multiple related data sources with direct or indirect relationships can help in reaching a more comprehensive view of RBPs and AS events and to reduce the impact of false positives associated with individual data sources. we prove that our proposed algorithm is convergent under some conditions and the experimental results have illustrated that APDCA outperforms six baseline methods in both AUC and AUPR. A case study on the RBP QKI shows that the top predictions are verified by the OncoSplicing database. Thus, APDCA provides a fast, interpretable, and scalable tool for discovering post-transcriptional regulatory interactions. Yangsong He, Zheng-Jian Bai, Wai-Ki Ching, Quan Zou 0001, Yushan Qiu |
PLoS Comput. Biol. | 5 |
| 2025 | PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clusteringabstractThe development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency. Mingzhu Liu, Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
PLoS Comput. Biol. | 3 |
| 2024 | scMNMF: a novel method for single-cell multi-omics clustering based on matrix factorizationabstractMOTIVATION: The technology for analyzing single-cell multi-omics data has advanced rapidly and has provided comprehensive and accurate cellular information by exploring cell heterogeneity in genomics, transcriptomics, epigenomics, metabolomics and proteomics data. However, because of the high-dimensional and sparse characteristics of single-cell multi-omics data, as well as the limitations of various analysis algorithms, the clustering performance is generally poor. Matrix factorization is an unsupervised, dimensionality reduction-based method that can cluster individuals and discover related omics variables from different blocks. Here, we present a novel algorithm that performs joint dimensionality reduction learning and cell clustering analysis on single-cell multi-omics data using non-negative matrix factorization that we named scMNMF. We formulate the objective function of joint learning as a constrained optimization problem and derive the corresponding iterative formulas through alternating iterative algorithms. The major advantage of the scMNMF algorithm remains its capability to explore hidden related features among omics data. Additionally, the feature selection for dimensionality reduction and cell clustering mutually influence each other iteratively, leading to a more effective discovery of cell types. We validated the performance of the scMNMF algorithm using two simulated and five real datasets. The results show that scMNMF outperformed seven other state-of-the-art algorithms in various measurements. AVAILABILITY AND IMPLEMENTATION: scMNMF code can be found at https://github.com/yushanqiu/scMNMF. Yushan Qiu, Quan Zou 0001 |
Briefings Bioinform. | 1 |
| 2024 | scTPC: a novel semisupervised deep clustering model for scRNA-seq dataabstractMOTIVATION: Continuous advancements in single-cell RNA sequencing (scRNA-seq) technology have enabled researchers to further explore the study of cell heterogeneity, trajectory inference, identification of rare cell types, and neurology. Accurate scRNA-seq data clustering is crucial in single-cell sequencing data analysis. However, the high dimensionality, sparsity, and presence of "false" zero values in the data can pose challenges to clustering. Furthermore, current unsupervised clustering algorithms have not effectively leveraged prior biological knowledge, making cell clustering even more challenging. RESULTS: This study investigates a semisupervised clustering model called scTPC, which integrates the triplet constraint, pairwise constraint, and cross-entropy constraint based on deep learning. Specifically, the model begins by pretraining a denoising autoencoder based on a zero-inflated negative binomial distribution. Deep clustering is then performed in the learned latent feature space using triplet constraints and pairwise constraints generated from partial labeled cells. Finally, to address imbalanced cell-type datasets, a weighted cross-entropy loss is introduced to optimize the model. A series of experimental results on 10 real scRNA-seq datasets and five simulated datasets demonstrate that scTPC achieves accurate clustering with a well-designed framework. AVAILABILITY AND IMPLEMENTATION: scTPC is a Python-based algorithm, and the code is available from https://github.com/LF-Yang/Code or https://zenodo.org/records/10951780. Yushan Qiu, Lingfei Yang, Hao Jiang 0009, Quan Zou 0001 |
Bioinform. | 1 |
| 2024 | AGML: Adaptive Graph-Based Multi-Label Learning for Prediction of RBP and as Event Associations During EMTabstractIncreasing evidence has indicated that RNA-binding proteins (RBPs) play an essential role in mediating alternative splicing (AS) events during epithelial-mesenchymal transition (EMT). However, due to the substantial cost and complexity of biological experiments, how AS events are regulated and influenced remains largely unknown. Thus, it is important to construct effective models for inferring hidden RBP-AS event associations during EMT process. In this paper, a novel and efficient model was developed to identify AS event-related candidate RBPs based on Adaptive Graph-based Multi-Label learning (AGML). In particular, we propose to adaptively learn a new affinity graph to capture the intrinsic structure of data for both RBPs and AS events. Multi-view similarity matrices are employed for maintaining the intrinsic structure and guiding the adaptive graph learning. We then simultaneously update the RBP and AS event associations that are predicted from both spaces by applying multi-label learning. The experimental results have shown that our AGML achieved AUC values of 0.9521 and 0.9873 by 5-fold and leave-one-out cross-validations, respectively, indicating the superiority and effectiveness of our proposed model. Furthermore, AGML can serve as an efficient and reliable tool for uncovering novel AS events-associated RBPs and is applicable for predicting the associations between other biological entities. Yushan Qiu, Wai-Ki Ching, Hongmin Cai, Hao Jiang 0009, Quan Zou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | SSNMDI: a novel joint learning model of semi-supervised non-negative matrix factorization and data imputation for clustering of single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) technology attracts extensive attention in the biomedical field. It can be used to measure gene expression and analyze the transcriptome at the single-cell level, enabling the identification of cell types based on unsupervised clustering. Data imputation and dimension reduction are conducted before clustering because scRNA-seq has a high 'dropout' rate, noise and linear inseparability. However, independence of dimension reduction, imputation and clustering cannot fully characterize the pattern of the scRNA-seq data, resulting in poor clustering performance. Herein, we propose a novel and accurate algorithm, SSNMDI, that utilizes a joint learning approach to simultaneously perform imputation, dimensionality reduction and cell clustering in a non-negative matrix factorization (NMF) framework. In addition, we integrate the cell annotation as prior information, then transform the joint learning into a semi-supervised NMF model. Through experiments on 14 datasets, we demonstrate that SSNMDI has a faster convergence speed, better dimensionality reduction performance and a more accurate cell clustering performance than previous methods, providing an accurate and robust strategy for analyzing scRNA-seq data. Biological analysis are also conducted to validate the biological significance of our method, including pseudotime analysis, gene ontology and survival analysis. We believe that we are among the first to introduce imputation, partial label information, dimension reduction and clustering to the single-cell field. AVAILABILITY AND IMPLEMENTATION: The source code for SSNMDI is available at https://github.com/yushanqiu/SSNMDI. Yushan Qiu, Chang Yan, Quan Zou 0001 |
Briefings Bioinform. | 1 |
| 2023 | ConSpaS: a contrastive learning framework for identifying spatial domains by integrating local and global similaritiesabstractSpatial transcriptomics is a rapidly growing field that aims to comprehensively characterize tissue organization and architecture at single-cell or sub-cellular resolution using spatial information. Such techniques provide a solid foundation for the mechanistic understanding of many biological processes in both health and disease that cannot be obtained using traditional technologies. Several methods have been proposed to decipher the spatial context of spots in tissue using spatial information. However, when spatial information and gene expression profiles are integrated, most methods only consider the local similarity of spatial information. As they do not consider the global semantic structure, spatial domain identification methods encounter poor or over-smoothed clusters. We developed ConSpaS, a novel node representation learning framework that precisely deciphers spatial domains by integrating local and global similarities based on graph autoencoder (GAE) and contrastive learning (CL). The GAE effectively integrates spatial information using local similarity and gene expression profiles, thereby ensuring that cluster assignment is spatially continuous. To improve the characterization of the global similarity of gene expression data, we adopt CL to consider the global semantic information. We propose an augmentation-free mechanism to construct global positive samples and use a semi-easy sampling strategy to define negative samples. We validated ConSpaS on multiple tissue types and technology platforms by comparing it with existing typical methods. The experimental results confirmed that ConSpaS effectively improved the identification accuracy of spatial domains with biologically meaningful spatial patterns, and denoised gene expression data while maintaining the spatial expression pattern. Furthermore, our proposed method better depicted the spatial trajectory by integrating local and global similarities. Siyao Wu, Yushan Qiu, Xiaoqing Cheng |
Briefings Bioinform. | 2 |
| 2023 | PCB: A pseudotemporal causality-based Bayesian approach to identify EMT-associated regulatory relationships of AS events and RBPs during breast cancer progressionabstractDuring breast cancer metastasis, the developmental process epithelial-mesenchymal (EM) transition is abnormally activated. Transcriptional regulatory networks controlling EM transition are well-studied; however, alternative RNA splicing also plays a critical regulatory role during this process. Alternative splicing was proved to control the EM transition process, and RNA-binding proteins were determined to regulate alternative splicing. A comprehensive understanding of alternative splicing and the RNA-binding proteins that regulate it during EM transition and their dynamic impact on breast cancer remains largely unknown. To accurately study the dynamic regulatory relationships, time-series data of the EM transition process are essential. However, only cross-sectional data of epithelial and mesenchymal specimens are available. Therefore, we developed a pseudotemporal causality-based Bayesian (PCB) approach to infer the dynamic regulatory relationships between alternative splicing events and RNA-binding proteins. Our study sheds light on facilitating the regulatory network-based approach to identify key RNA-binding proteins or target alternative splicing events for the diagnosis or treatment of cancers. The data and code for PCB are available at: http://hkumath.hku.hk/~wkc/PCB(data+code).zip. Liangjie Sun, Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
PLoS Comput. Biol. | 2 |
| 2023 | scHOIS: Determining Cell Heterogeneity Through Hierarchical Clustering Based on Optimal Imputation StrategyabstractAdvances in single-cell RNA sequencing (scRNA-seq) technology provide an unbiased and high-throughput analysis of each cell at single-cell resolution, and further facilitate the development of cellular heterogeneity analysis. Despite the promise of scRNA-seq, the data generated by this method are sparse and noisy because of the presence of dropout events, which can greatly impact downstream analyses such as differential gene expression, cell type annotation, and linage trajectory reconstruction. The development of effective and robust computational methods to address both dropout and clustering are thus urgently needed. In this study, we propose a flexible, accurate two-stage algorithm for single cell heterogeneity analysis via hierarchical clustering based on an optimal imputation strategy, called scHOIS. At the first stage, masked non-negative matrix factorization is applied to approximate the original observed scRNA-seq data, with optimal rank determined by variance analysis. At the second stage, hierarchical clustering is applied to group the imputed cells using Pearson correlation to measure similarity, with the optimal number of clusters determined by integrating three classical indexes. We performed extensive experiments on real-world datasets, which showed that scHOIS effectively and robustly distinguished cellular differences and that the clustering performance of this algorithm was superior to that of other state-of-the-art methods. Xiaoqing Cheng, Chang Yan, Hao Jiang 0009, Yushan Qiu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | MDICC: novel method for multi-omics data integration and cancer subtype identificationabstractEach type of cancer usually has several subtypes with distinct clinical implications, and therefore the discovery of cancer subtypes is an important and urgent task in disease diagnosis and therapy. Using single-omics data to predict cancer subtypes is difficult because genomes are dysregulated and complicated by multiple molecular mechanisms, and therefore linking cancer genomes to cancer phenotypes is not an easy task. Using multi-omics data to effectively predict cancer subtypes is an area of much interest; however, integrating multi-omics data is challenging. Here, we propose a novel method of multi-omics data integration for clustering to identify cancer subtypes (MDICC) that integrates new affinity matrix and network fusion methods. Our experimental results show the effectiveness and generalization of the proposed MDICC model in identifying cancer subtypes, and its performance was better than those of currently available state-of-the-art clustering methods. Furthermore, the survival analysis demonstrates that MDICC delivered comparable or even better results than many typical integrative methods. Sha Tian, Yushan Qiu, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2022 | A high-order norm-product regularized multiple kernel learning framework for kernel optimization
Hao Jiang 0009, Dong Shen 0002, Wai-Ki Ching, Yushan Qiu |
Inf. Sci. | 4 |
| 2021 | HOMC: A Hierarchical Clustering Algorithm Based on Optimal Low Rank Matrix Completion for Single Cell Analysis
Xiaoqing Cheng, Chang Yan, Hao Jiang 0009, Yushan Qiu |
ICIC (3) | 4 |
| 2021 | Matrix factorization-based data fusion for the prediction of RNA-binding proteins and alternative splicing event associations during epithelial-mesenchymal transitionabstractMOTIVATION: The epithelial-mesenchymal transition (EMT) is a cellular-developmental process activated during tumor metastasis. Transcriptional regulatory networks controlling EMT are well studied; however, alternative RNA splicing also plays a critical regulatory role during this process. Unfortunately, a comprehensive understanding of alternative splicing (AS) and the RNA-binding proteins (RBPs) that regulate it during EMT remains largely unknown. Therefore, a great need exists to develop effective computational methods for predicting associations of RBPs and AS events. Dramatically increasing data sources that have direct and indirect information associated with RBPs and AS events have provided an ideal platform for inferring these associations. RESULTS: In this study, we propose a novel method for RBP-AS target prediction based on weighted data fusion with sparse matrix tri-factorization (WDFSMF in short) that simultaneously decomposes heterogeneous data source matrices into low-rank matrices to reveal hidden associations. WDFSMF can select and integrate data sources by assigning different weights to those sources, and these weights can be assigned automatically. In addition, WDFSMF can identify significant RBP complexes regulating AS events and eliminate noise and outliers from the data. Our proposed method achieves an area under the receiver operating characteristic curve (AUC) of $90.78\%$, which shows that WDFSMF can effectively predict RBP-AS event associations with higher accuracy compared with previous methods. Furthermore, this study identifies significant RBPs as complexes for AS events during EMT and provides solid ground for further investigation into RNA regulation during EMT and metastasis. WDFSMF is a general data fusion framework, and as such it can also be adapted to predict associations between other biological entities. Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
Briefings Bioinform. | 1 |
| 2021 | Prediction of RNA-binding protein and alternative splicing event associations during epithelial-mesenchymal transition based on inductive matrix completionabstractMOTIVATION: The developmental process of epithelial-mesenchymal transition (EMT) is abnormally activated during breast cancer metastasis. Transcriptional regulatory networks that control EMT have been well studied; however, alternative RNA splicing plays a vital regulatory role during this process and the regulating mechanism needs further exploration. Because of the huge cost and complexity of biological experiments, the underlying mechanisms of alternative splicing (AS) and associated RNA-binding proteins (RBPs) that regulate the EMT process remain largely unknown. Thus, there is an urgent need to develop computational methods for predicting potential RBP-AS event associations during EMT. RESULTS: We developed a novel model for RBP-AS target prediction during EMT that is based on inductive matrix completion (RAIMC). Integrated RBP similarities were calculated based on RBP regulating similarity, and RBP Gaussian interaction profile (GIP) kernel similarity, while integrated AS event similarities were computed based on AS event module similarity and AS event GIP kernel similarity. Our primary objective was to complete missing or unknown RBP-AS event associations based on known associations and on integrated RBP and AS event similarities. In this paper, we identify significant RBPs for AS events during EMT and discuss potential regulating mechanisms. Our computational results confirm the effectiveness and superiority of our model over other state-of-the-art methods. Our RAIMC model achieved AUC values of 0.9587 and 0.9765 based on leave-one-out cross-validation (CV) and 5-fold CV, respectively, which are larger than the AUC values from the previous models. RAIMC is a general matrix completion framework that can be adopted to predict associations between other biological entities. We further validated the prediction performance of RAIMC on the genes CD44 and MAP3K7. RAIMC can identify the related regulating RBPs for isoforms of these two genes. AVAILABILITY AND IMPLEMENTATION: The source code for RAIMC is available at https://github.com/yushanqiu/RAIMC. CONTACT: [email protected] online. Yushan Qiu, Wai-Ki Ching, Quan Zou 0001 |
Briefings Bioinform. | 1 |
| 2021 | Unsupervised Learning Framework With Multidimensional Scaling in Predicting Epithelial-Mesenchymal TransitionsabstractClustering tumor metastasis samples from gene expression data at the whole genome level remains an arduous challenge, in particular, when the number of experimental samples is small and the number of genes is huge. We focus on the prediction of the epithelial-mesenchymal transition (EMT), which is an underlying mechanism of tumor metastasis, here, rather than tumor metastasis itself, to avoid confounding effects of uncertainties derived from various factors. In this paper, we propose a novel model in predicting EMT based on multidimensional scaling (MDS) strategies and integrating entropy and random matrix detection strategies to determine the optimal reduced number of dimension in low dimensional space. We verified our proposed model with the gene expression data for EMT samples of breast cancer and the experimental results demonstrated the superiority over state-of-the-art clustering methods. Furthermore, we developed a novel feature extraction method for selecting the significant genes and predicting the tumor metastasis. The source code is available at "https://github.com/yushanqiu/yushan.qiu-szu.edu.cn". Yushan Qiu, Hao Jiang 0009, Wai-Ki Ching |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Drug Side-Effect Profiles Prediction: From Empirical to Structural Risk MinimizationabstractThe identification of drug side-effects is considered to be an important step in drug design, which could not only shorten the time but also reduce the cost of drug development. In this paper, we investigate the relationship between the potential side-effects of drug candidates and their chemical structures. The preliminary Regularized Regression (RR) model for drug side-effects prediction has promising features in the efficiency of model training and the existence of a closed form solution. It performs better than other state-of-the-art methods, in terms of minimum accuracy and average accuracy. In order to dig inside how drug structure will associate with side effect, we further propose weighted GTS (Generalized T-Student Kernel: WGTS) SVM model from a structural risk minimization perspective. The SVM model proposed in this paper provides a better understanding of drug side-effects in the process of drug development. The usefulness of the WGTS model lies in the superior performance in a cross validation setting on 888 approved drugs with 1385 side-effects profiling from SIDER database. This work is expected to shed light on intriguing studies that predict potential un-identifying side-effects and suggest how we can avoid drug side-effects by the removal of some distinguished chemical structures. Hao Jiang 0009, Yushan Qiu, Wenpin Hou, Xiaoqing Cheng, Man Yi Yim, Wai-Ki Ching |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | On predicting epithelial mesenchymal transition by integrating RNA-binding proteins and correlation data via L1/2-regularization method
Yushan Qiu, Hao Jiang 0009, Wai-Ki Ching, Michael Kwok-Po Ng |
Artif. Intell. Medicine | 1 |
| 2018 | Novel Method for Singleton and Cyclic Attractor Observability in Boolean Networks
Yushan Qiu, Shaobo Tan, Dongqi Li, Ada Chaeli van der Zijp-Tan, Ada Fong, Glen M. Borchert, Jingshan Huang |
BIBM | 1 |
| 2016 | Unconstrained optimization in projection method for indefinite SVMsabstractPositive semi-definiteness is a critical property in Support Vector Machine (SVM) methods to ensure efficient solutions through convex quadratic programming. In this paper, we introduce a projection matrix on indefinite kernels to formulate a positive semi-definite one. The proposed model can be regarded as a generalized version of the spectrum method (denoising method and flipping method) by varying parameter λ. In particular, our suggested optimal λ under the Bregman matrix divergence theory can be obtained using unconstrained optimization. Experimental results on 4 real world data sets ranging from glycan classification to cancer prediction show that the proposed model can achieve better or competitive performance when compared to the related indefinite kernel methods. This may suggest a new way in motif extractions or cancer predictions. Hao Jiang 0009, Wai-Ki Ching, Yushan Qiu, Xiaoqing Cheng |
BIBM | 3 |
| 2016 | A systematic framework to derive N-glycan biosynthesis process and the automated construction of glycosylation networksabstractBACKGROUND: Abnormalities in glycan biosynthesis have been conclusively related to various diseases, whereas the complexity of the glycosylation process has impeded the quantitative analysis of biochemical experimental data for the identification of glycoforms contributing to disease. To overcome this limitation, the automatic construction of glycosylation reaction networks in silico is a critical step. RESULTS: In this paper, a framework K2014 is developed to automatically construct N-glycosylation networks in MATLAB with the involvement of the 27 most-known enzyme reaction rules of 22 enzymes, as an extension of previous model KB2005. A toolbox named Glycosylation Network Analysis Toolbox (GNAT) is applied to define network properties systematically, including linkages, stereochemical specificity and reaction conditions of enzymes. Our network shows a strong ability to predict a wider range of glycans produced by the enzymes encountered in the Golgi Apparatus in human cell expression systems. CONCLUSIONS: Our results demonstrate a better understanding of the underlying glycosylation process and the potential of systems glycobiology tools for analyzing conventional biochemical or mass spectrometry-based experimental data quantitatively in a more realistic and practical way. Wenpin Hou, Yushan Qiu, Nobuyuki Hashimoto, Wai-Ki Ching, Kiyoko F. Aoki-Kinoshita |
BMC Bioinform. | 2 |
| 2016 | Exact Identification of the Structure of a Probabilistic Boolean Network from SamplesabstractWe study the number of samples required to uniquely determine the structure of a probabilistic Boolean network (PBN), where PBNs are probabilistic extensions of Boolean networks. We show via theoretical analysis and computational analysis that the structure of a PBN can be exactly identified with high probability from a relatively small number of samples for interesting classes of PBNs of bounded indegree. On the other hand, we also show that there exist classes of PBNs for which it is impossible to uniquely determine the structure of a PBN from samples. Xiaoqing Cheng, Tomoya Mori, Yushan Qiu, Wai-Ki Ching, Tatsuya Akutsu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2015 | On observability of attractors in Boolean NetworksabstractBoolean network (BN) is a popular mathematical model for revealing the behavior of a genetic regulatory network, and observability plays a vital role in understanding the underlying network feature. However, the observability of attractor cycles, which is an interesting and important problem, has not been addressed in the literature. In this paper, we first proposed a novel problem on attractor observability in BNs. Identification of the minimum set of consecutive nodes can be used to determine uniquely the attractor cycle from the others in the network. We then develop a linear-time algorithm to identify the desired set of nodes. The proposed approaches are demonstrated and verified by numerical examples. The computational results are given to illustrate both the efficiency and effectiveness of our proposed methods. Yushan Qiu, Xiaoqing Cheng, Wai-Ki Ching, Hao Jiang 0009, Tatsuya Akutsu |
BIBM | 1 |