Ao Li 0001

dblp:54/2788-1 · DBLP profile ↗
← Back
49ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0001-9910-8967ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 14 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Cross-stain knowledge distillation for low-cost lung cancer programmed death ligand-1 assessment with multi-granularity multiple instance learning
Yi Shi 0013, Chong Ge, Anli Zhang, Ao Li 0001
Eng. Appl. Artif. Intell.5
2026 Understanding and tackling the modality imbalance problem in multimodal survival prediction
Chicheng Zhou, Yi Shi 0013, Anli Zhang, Ao Li 0001
Pattern Recognit.5
2026 VGFA: Variation-Robust Graph-Level Feature Alignment for Domain Adaptive Nuclei Detection
abstract
Accurate detection of nuclei is a crucial step in advancing pathology image analysis for disease diagnosis and treatment. However, significant domain discrepancies exist among pathology images, which severely degrade the performance of detection models. Despite promising results from existing domain adaptation approaches, they may overlook the detrimental impact of intra-domain variation (IDV) at both cell and image scales. The IDV issue typically manifests as dramatic differences in nuclei composition, morphology, and spatial arrangement, which occurs not only across different pathology images but also within a single image. This inherent heterogeneity produces highly complex feature distributions, ultimately making cross-domain alignment significantly more arduous. Moreover, the presence of IDV further injects noise into pseudo-labels, reducing the signal-to-noise ratio in the feature space and complicating the alignment process. To tackle these challenges, we propose a novel variation-robust graph-level feature alignment (VGFA) framework for unsupervised domain adaptive nuclei detection. Specifically, our method first incorporates a prior-based nuclei graph pruning scheme that harnesses nuclei spatial contextual priors and dynamically eliminates unreliable nodes from the nuclei graph. Then, a local-global nuclei encoding network is designed to learn nuclei graph representations that holistically encapsulate the consistent traits among various nuclei, thereby mitigating challenges posed by cell-scale IDV. Moreover, VGFA leverages a nuclei graph discrepancy loss that is resilient to image-scale IDV, achieving effective feature alignment in cross-domain graph feature space. Extensive experiments across different adaptation scenarios demonstrate that our VGFA framework achieves state-of-the-art performance, outperforming existing feature alignment methods in domain adaptive nuclei detection.
Aiqiu Wu, Anli Zhang, Ao Li 0001
IEEE Trans. Medical Imaging5
2025 Domain-Specific Interactive Prompting for Generalized Nuclei Classification
abstract
Accurate nuclei classification serves as a critical cornerstone for disease diagnosis and treatment, yet challenged by the heterogeneity of tissue types, staining procedures, and imaging techniques. Recently, vision-language models (VLMs) have demonstrated impressive success in the natural image field and advanced potential in medical imaging. However, the adaptation of VLMs to nuclei classification still poses several challenges, including limited generalization capability and coarse image-text feature alignment. In this paper, we propose SIGNPrompt, a domain-Specific Interactive Prompt learning framework for Generalized Nuclei classification. Specifically, to unleash the generalization capability of VLM-based models, we introduce a prior-guided domain adapting module that integrates nuclei prior information from the large language model (LLM), enabling flexible and robust adaptation to the inherent heterogeneity across pathological domains. Moreover, we develop a multi-modal interactive prompting mechanism to refine image-text feature alignment by leveraging the interdependence between visual and language prompting, thus enhancing the discriminability of nuclei categories. In addition, a simple yet effective noise-adding strategy is proposed to mitigate the overfitting problem in prompt learning. Extensive experiments on diverse public benchmarks and challenging zero-shot scenarios validate that SIGNPrompt consistently outperforms state-of-the-art (SOTA) methods in both accuracy and generalization.
Aiqiu Wu, Ao Li 0001
ACM Multimedia4
2025 Seeking multi-view commonality and peculiarity: A novel decoupling method for lung cancer subtype classification
Ziyu Gao, Yin Luo, Chi Cao, Houzhou Jiang, Ao Li 0001
Expert Syst. Appl.7
2025 Agnostic-Specific Modality Learning for Cancer Survival Prediction From Multiple Data
abstract
Cancer is a pressing public health problem and one of the main causes of mortality worldwide. The development of advanced computational methods for predicting cancer survival is pivotal in aiding clinicians to formulate effective treatment strategies and improve patient quality of life. Recent advances in survival prediction methods show that integrating diverse information from various cancer-related data, such as pathological images and genomics, is crucial for improving prediction accuracy. Despite promising results of existing approaches, there are great challenges of modality gap and semantic redundancy presented in multiple cancer data, which could hinder the comprehensive integration and pose substantial obstacles to further enhancing cancer survival prediction. In this study, we propose a novel agnostic-specific modality learning (ASML) framework for accurate cancer survival prediction. To bridge the modality gap and provide a comprehensive view of distinct data modalities, we employ an agnostic-specific learning strategy to learn the commonality across modalities and the uniqueness of each modality. Moreover, a cross-modal fusion network is exerted to integrate multimodal information by modeling modality correlations and diminish semantic redundancy in a divide-and-conquer manner. Extensive experiment results on three TCGA datasets demonstrate that ASML reaches better performance than other existing cancer survival prediction methods for multiple data.
Yi Shi 0013, Ao Li 0001
IEEE J. Biomed. Health Informatics4
2025 MIF: Multi-Shot Interactive Fusion Model for Cancer Survival Prediction Using Pathological Image and Genomic Data
abstract
Accurate cancer survival prediction is crucial for oncologists to determine therapeutic plan, which directly influences the treatment efficacy and survival outcome of patient. Recently, multimodal fusion-based prognostic methods have demonstrated effectiveness for survival prediction by fusing diverse cancer-related data from different medical modalities, e.g., pathological images and genomic data. However, these works still face significant challenges. First, most approaches attempt multimodal fusion by simple one-shot fusion strategy, which is insufficient to explore complex interactions underlying in highly disparate multimodal data. Second, current methods for investigating multimodal interactions face the capability-efficiency dilemma, which is the difficult balance between powerful modeling capability and applicable computational efficiency, thus impeding effective multimodal fusion. In this study, to encounter these challenges, we propose an innovative multi-shot interactive fusion method named MIF for precise survival prediction by utilizing pathological and genomic data. Particularly, a novel multi-shot fusion framework is introduced to promote multimodal fusion by decomposing it into successive fusing stages, thus delicately integrating modalities in a progressive way. Moreover, to address the capacity-efficiency dilemma, various affinity-based interactive modules are introduced to synergize the multi-shot framework. Specifically, by harnessing comprehensive affinity information as guidance for mining interactions, the proposed interactive modules can efficiently generate low-dimensional discriminative multimodal representations. Extensive experiments on different cancer datasets unravel that our method not only successfully achieves state-of-the-art performance by performing effective multimodal fusion, but also possesses high computational efficiency compared to existing survival prediction methods.
Yi Shi 0013, Ao Li 0001, Xun Chen 0001
IEEE J. Biomed. Health Informatics5
2025 Tackling Tumor Heterogeneity Issue: Transformer-Based Multiple Instance Enhancement Learning for Predicting EGFR Mutation via CT Images
abstract
Accurate and non-invasive prediction of epidermal growth factor receptor (EGFR) mutation is crucial for the diagnosis and treatment of non-small cell lung cancer (NSCLC). While computed tomography (CT) imaging shows promise in identifying EGFR mutation, current prediction methods heavily rely on fully supervised learning, which overlooks the substantial heterogeneity of tumors and therefore leads to suboptimal results. To tackle tumor heterogeneity issue, this study introduces a novel weakly supervised method named TransMIEL, which leverages multiple instance learning techniques for accurate EGFR mutation prediction. Specifically, we first propose an innovative instance enhancement learning (IEL) strategy that strengthens the discriminative power of instance features for complex tumor CT images by exploring self-derived soft pseudo-labels. Next, to improve tumor representation capability, we design a spatial-aware transformer (SAT) that fully captures inter-instance relationships of different pathological subregions to mirror the diagnostic processes of radiologists. Finally, an instance adaptive gating (IAG) module is developed to effectively emphasize the contribution of informative instance features in heterogeneous tumors, facilitating dynamic instance feature aggregation and increasing model generalization performance. Experimental results demonstrate that TransMIEL significantly outperforms existing fully and weakly supervised methods on both public and in-house NSCLC datasets. Additionally, visualization results show that our approach can highlight intra-tumor and peri-tumor areas relevant to EGFR mutation status. Therefore, our method holds significant potential as an effective tool for EGFR prediction and offers a novel perspective for future research on tumor heterogeneity.
Yulin Fang, Qilong Song, Chi Cao, Ziyu Gao, Biao Song, Xuhong Min, Ao Li 0001
IEEE Trans. Medical Imaging8
2024 Cross-Domain Nuclei Detection in Histopathology Images Using Graph-Based Nuclei Feature Alignment
abstract
As powerful tools deep neural networks have been successfully adopted for nuclei detection in histopathology images, whereas require the same probability distribution between training and testing data. However, domain shift among histopathology images widely exists in real-world applications and severely deteriorates the detection performance of deep neural networks. Despite encouraging results of existing domain adaptation methods, there remain challenges for cross-domain nuclei detection task. First, in view of the tiny size of nuclei, it is actually very difficult to obtain sufficient nuclei features, thus leading to a negative influence for feature alignment. Second, due to unavailable annotations in target domain, some extracted features contain background pixels and are thereby indiscriminative, which can largely confuse the alignment procedure. To address these challenges, in this paper, we propose an end-to-end graph-based nuclei feature alignment (GNFA) method for boosting cross-domain nuclei detection. Concretely, sufficient nuclei features are generated from nuclei graph convolutional network (NGCN) by aggregating information of adjacent nuclei upon construction of nuclei graph for successful alignment. In addition, importance learning module (ILM) is designed to further select discriminative nuclei features for mitigating negative influence of background pixels in target domain during alignment. By utilizing sufficient and discriminative node features generated from GNFA, our method can successfully perform feature alignment and effectively alleviate domain shift problem for nuclei detection. Extensive experiments of multiple adaptation scenarios reveal that our method achieves state-of-the-art performance in cross-domain nuclei detection compared with existing domain adaptation methods.
Xiaoya Zhu, Gang Meng, Ao Li 0001
IEEE J. Biomed. Health Informatics7
2023 CARL: Cross-Aligned Representation Learning for Multi-view Lung Cancer Histology Classification
Yin Luo, Wei Liu 0235, Qilong Song, Xuhong Min, Ao Li 0001
MICCAI (5)7
2023 CAMR: cross-aligned multimodal representation learning for cancer survival prediction
abstract
MOTIVATION: Accurately predicting cancer survival is crucial for helping clinicians to plan appropriate treatments, which largely improves the life quality of cancer patients and spares the related medical costs. Recent advances in survival prediction methods suggest that integrating complementary information from different modalities, e.g. histopathological images and genomic data, plays a key role in enhancing predictive performance. Despite promising results obtained by existing multimodal methods, the disparate and heterogeneous characteristics of multimodal data cause the so-called modality gap problem, which brings in dramatically diverse modality representations in feature space. Consequently, detrimental modality gaps make it difficult for comprehensive integration of multimodal information via representation learning and therefore pose a great challenge to further improvements of cancer survival prediction. RESULTS: To solve the above problems, we propose a novel method called cross-aligned multimodal representation learning (CAMR), which generates both modality-invariant and -specific representations for more accurate cancer survival prediction. Specifically, a cross-modality representation alignment learning network is introduced to reduce modality gaps by effectively learning modality-invariant representations in a common subspace, which is achieved by aligning the distributions of different modality representations through adversarial training. Besides, we adopt a cross-modality fusion module to fuse modality-invariant representations into a unified cross-modality representation for each patient. Meanwhile, CAMR learns modality-specific representations which complement modality-invariant representations and therefore provides a holistic view of the multimodal data for cancer survival prediction. Comprehensive experiment results demonstrate that CAMR can successfully narrow modality gaps and consistently yields better performance than other survival prediction methods using multimodal data. AVAILABILITY AND IMPLEMENTATION: CAMR is freely available at https://github.com/wxq-ustc/CAMR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xingqi Wu, Yi Shi 0013, Ao Li 0001
Bioinform.4
2023 Dual consistency semi-supervised nuclei detection via global regularization and local adversarial learning
Xiaoya Zhu, Gang Meng, Ao Li 0001
Neurocomputing6
2022 Global and local attentional feature alignment for domain adaptive nuclei detection in histopathology images
Xiaoya Zhu, Ao Li 0001, Gang Meng
Artif. Intell. Medicine3
2022 HFBSurv: hierarchical multimodal fusion with factorized bilinear models for cancer survival prediction
abstract
MOTIVATION: Cancer survival prediction can greatly assist clinicians in planning patient treatments and improving their life quality. Recent evidence suggests the fusion of multimodal data, such as genomic data and pathological images, is crucial for understanding cancer heterogeneity and enhancing survival prediction. As a powerful multimodal fusion technique, Kronecker product has shown its superiority in predicting survival. However, this technique introduces a large number of parameters that may lead to high computational cost and a risk of overfitting, thus limiting its applicability and improvement in performance. Another limitation of existing approaches using Kronecker product is that they only mine relations for one single time to learn multimodal representation and therefore face significant challenges in deeply mining rich information from multimodal data for accurate survival prediction. RESULTS: To address the above limitations, we present a novel hierarchical multimodal fusion approach named HFBSurv by employing factorized bilinear model to fuse genomic and image features step by step. Specifically, with a multiple fusion strategy HFBSurv decomposes the fusion problem into different levels and each of them integrates and passes information progressively from the low level to the high level, thus leading to the more specialized fusion procedure and expressive multimodal representation. In this hierarchical framework, both modality-specific and cross-modality attentional factorized bilinear modules are designed to not only capture and quantify complex relations from multimodal data, but also dramatically reduce computational complexity. Extensive experiments demonstrate that our method performs an effective hierarchical fusion of multimodal data and achieves consistently better performance than other methods for survival prediction. AVAILABILITY AND IMPLEMENTATION: HFBSurv is freely available at https://github.com/Liruiqing-ustc/HFBSurv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xingqi Wu, Ao Li 0001
Bioinform.3
2022 TSDLPP: A Novel Two-Stage Deep Learning Framework For Prognosis Prediction Based on Whole Slide Histopathological Images
abstract
Recently, digital pathology image-based prognosis prediction has become a hot topic in healthcare research to make early decisions on therapy and improve the treatment quality of patients. Therefore, there has been a recent surge of interest in designing deep learning method solving the problem of prognosis prediction with digital pathology images. However, whole slide histopathological images (WSIs) based prognosis prediction is still a challenge due to the large size of pathological images, the heterogeneity of tumors and the high cost of region of interests (ROIs) labeling. In this study, we design a novel two-stage deep learning framework for prognosis prediction (TSDLPP) based on WSIs. Our proposed framework consists of two-stage paradigms: 1) training tissue decomposition network (TDNet) to divide WSIs into cancerous and non-cancerous regions, 2) integrating general prognosis-related densely connected CNN (GPR-DCCNN) and morphology-specific prognosis-related densely connected CNNs (MSPR-DCCNNs) to extract different level features of pathological images. In the end, we apply TSDLPP to the prognosis prediction of breast cancer using The Cancer Genome Atlas (TCGA) datasets. Experiment results demonstrate that TSDLPP obtains superior performance of prognosis prediction compared with the existing state-of-arts methods.
Yu Liu 0113, Ao Li 0001, Jiangshu Liu, Gang Meng
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 Reconstruction-Assisted Feature Encoding Network for Histologic Subtype Classification of Non-Small Cell Lung Cancer
abstract
Accurate histological subtype classification between adenocarcinoma (ADC) and squamous cell carcinoma (SCC) using computed tomography (CT) images is of great importance to assist clinicians in determining treatment and therapy plans for non-small cell lung cancer (NSCLC) patients. Although current deep learning approaches have achieved promising progress in this field, they are often difficult to capture efficient tumor representations due to inadequate training data, and in consequence show limited performance. In this study, we propose a novel and effective reconstruction-assisted feature encoding network (RAFENet) for histological subtype classification by leveraging an auxiliary image reconstruction task to enable extra guidance and regularization for enhanced tumor feature representations. Different from existing reconstruction-assisted methods that directly use generalizable features obtained from shared encoder for primary task, a dedicated task-aware encoding module is utilized in RAFENet to perform refinement of generalizable features. Specifically, a cascade of cross-level non-local blocks are introduced to progressively refine generalizable features at different levels with the aid of lower-level task-specific information, which can successfully learn multi-level task-specific features tailored to histological subtype classification. Moreover, in addition to widely adopted pixel-wise reconstruction loss, we introduce a powerful semantic consistency loss function to explicitly supervise the training of RAFENet, which combines both feature consistency loss and prediction consistency loss to ensure semantic invariance during image reconstruction. Extensive experimental results show that RAFENet effectively addresses the difficult issues that cannot be resolved by existing reconstruction-based methods and consistently outperforms other state-of-the-art methods on both public and in-house NSCLC datasets. Supplementary material is available at https://github.com/lhch1994/Rafenet_sup_material.
Haichun Li, Qilong Song, Dongqi Gui, Xuhong Min, Ao Li 0001
IEEE J. Biomed. Health Informatics6
2021 Instance-Aware Feature Alignment for Cross-Domain Cell Nuclei Detection in Histopathology Images
Xiaoya Zhu, Gang Meng, Junsheng Zhang, Ao Li 0001
MICCAI (8)6
2021 PedMiner: a tool for linkage analysis-based identification of disease-associated variants using family based whole-exome sequencing data
abstract
With the advances of next-generation sequencing technology, the field of disease research has been revolutionized. However, pinpointing the disease-causing variants from millions of revealed variants is still a tough task. Here, we have reviewed the existing linkage analysis tools and presented PedMiner, a web-based application designed to narrow down candidate variants from family based whole-exome sequencing (WES) data through linkage analysis. PedMiner integrates linkage analysis, variant annotation and prioritization in one automated pipeline. It provides graphical visualization of the linked regions along with comprehensive annotation of variants and genes within these linked regions. This efficient and comprehensive application will be helpful for the scientific community working on Mendelian inherited disorders using family based WES data.
Jianteng Zhou, Jianing Gao, Daren Zhao, Ao Li 0001, Furhan Iqbal, Qinghua Shi, Yuanwei Zhang
Briefings Bioinform.5
2021 GPDBN: deep bilinear network integrating both genomic data and pathological images for breast cancer prognosis prediction
abstract
MOTIVATION: Breast cancer is a very heterogeneous disease and there is an urgent need to design computational methods that can accurately predict the prognosis of breast cancer for appropriate therapeutic regime. Recently, deep learning-based methods have achieved great success in prognosis prediction, but many of them directly combine features from different modalities that may ignore the complex inter-modality relations. In addition, existing deep learning-based methods do not take intra-modality relations into consideration that are also beneficial to prognosis prediction. Therefore, it is of great importance to develop a deep learning-based method that can take advantage of the complementary information between intra-modality and inter-modality by integrating data from different modalities for more accurate prognosis prediction of breast cancer. RESULTS: We present a novel unified framework named genomic and pathological deep bilinear network (GPDBN) for prognosis prediction of breast cancer by effectively integrating both genomic data and pathological images. In GPDBN, an inter-modality bilinear feature encoding module is proposed to model complex inter-modality relations for fully exploiting intrinsic relationship of the features across different modalities. Meanwhile, intra-modality relations that are also beneficial to prognosis prediction, are captured by two intra-modality bilinear feature encoding modules. Moreover, to take advantage of the complementary information between inter-modality and intra-modality relations, GPDBN further combines the inter- and intra-modality bilinear features by using a multi-layer deep neural network for final prognosis prediction. Comprehensive experiment results demonstrate that the proposed GPDBN significantly improves the performance of breast cancer prognosis prediction and compares favorably with existing methods. AVAILABILITYAND IMPLEMENTATION: GPDBN is freely available at https://github.com/isfj/GPDBN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhiqin Wang, Ao Li 0001
Bioinform.4
2021 PhosIDN: an integrated deep neural network for improving protein phosphorylation site prediction by combining sequence and protein-protein interaction information
abstract
MOTIVATION: Phosphorylation is one of the most studied post-translational modifications, which plays a pivotal role in various cellular processes. Recently, deep learning methods have achieved great success in prediction of phosphorylation sites, but most of them are based on convolutional neural network that may not capture enough information about long-range dependencies between residues in a protein sequence. In addition, existing deep learning methods only make use of sequence information for predicting phosphorylation sites, and it is highly desirable to develop a deep learning architecture that can combine heterogeneous sequence and protein-protein interaction (PPI) information for more accurate phosphorylation site prediction. RESULTS: We present a novel integrated deep neural network named PhosIDN, for phosphorylation site prediction by extracting and combining sequence and PPI information. In PhosIDN, a sequence feature encoding sub-network is proposed to capture not only local patterns but also long-range dependencies from protein sequences. Meanwhile, useful PPI features are also extracted in PhosIDN by a PPI feature encoding sub-network adopting a multi-layer deep neural network. Moreover, to effectively combine sequence and PPI information, a heterogeneous feature combination sub-network is introduced to fully exploit the complex associations between sequence and PPI features, and their combined features are used for final prediction. Comprehensive experiment results demonstrate that the proposed PhosIDN significantly improves the prediction performance of phosphorylation sites and compares favorably with existing general and kinase-specific phosphorylation site prediction methods. AVAILABILITY AND IMPLEMENTATION: PhosIDN is freely available at https://github.com/ustchangyuanyang/PhosIDN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hangyuan Yang, Xing-Ming Zhao, Ao Li 0001
Bioinform.5
2021 DP-SLAM: A visual SLAM with moving probability towards dynamic environments
Ao Li 0001, Meng Xu 0004, Zonghai Chen
Inf. Sci.1
2020 Inferring subgroup-specific driver genes from heterogeneous cancer samples via subspace learning with subgroup indication
abstract
MOTIVATION: Detecting driver genes from gene mutation data is a fundamental task for tumorigenesis research. Due to the fact that cancer is a heterogeneous disease with various subgroups, subgroup-specific driver genes are the key factors in the development of precision medicine for heterogeneous cancer. However, the existing driver gene detection methods are not designed to identify subgroup specificities of their detected driver genes, and therefore cannot indicate which group of patients is associated with the detected driver genes, which is difficult to provide specifically clinical guidance for individual patients. RESULTS: By incorporating the subspace learning framework, we propose a novel bioinformatics method called DriverSub, which can efficiently predict subgroup-specific driver genes in the situation where the subgroup annotations are not available. When evaluated by simulation datasets with known ground truth and compared with existing methods, DriverSub yields the best prediction of driver genes and the inference of their related subgroups. When we apply DriverSub on the mutation data of real heterogeneous cancers, we can observe that the predicted results of DriverSub are highly enriched for experimentally validated known driver genes. Moreover, the subgroups inferred by DriverSub are significantly associated with the annotated molecular subgroups, indicating its capability of predicting subgroup-specific driver genes. AVAILABILITY AND IMPLEMENTATION: The source code is publicly available at https://github.com/JianingXi/DriverSub. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianing Xi, Xiguo Yuan, Ao Li 0001, Xuelong Li 0001, Qinghua Huang
Bioinform.4
2020 SCSsim: an integrated tool for simulating single-cell genome sequencing data
abstract
MOTIVATION: Allele dropout (ADO) and unbalanced amplification of alleles are main technical issues of single-cell sequencing (SCS), and effectively emulating these issues is necessary for reliably benchmarking SCS-based bioinformatics tools. Unfortunately, currently available sequencing simulators are free of whole-genome amplification involved in SCS technique and therefore not suited for generating SCS datasets. We develop a new software package (SCSsim) that can efficiently simulate SCS datasets in a parallel fashion with minimal user intervention. SCSsim first constructs the genome sequence of single cell by mimicking a complement of genomic variations under user-controlled manner, and then amplifies the genome according to MALBAC technique and finally yields sequencing reads from the amplified products based on inferred sequencing profiles. Comprehensive evaluation in simulating different ADO rates, variation detection efficiency and genome coverage demonstrates that SCSsim is a very useful tool in mimicking single-cell sequencing data with high efficiency. AVAILABILITY AND IMPLEMENTATION: SCSsim is freely available at https://github.com/qasimyu/scssim. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhenhua Yu 0002, Fang Du, Xuehong Sun, Ao Li 0001
Bioinform.4
2020 Dual-Layer Strengthened Collaborative Topic Regression Modeling for Predicting Drug Sensitivity
abstract
An effective way to facilitate the development of modern oncology precision medicine is the systematical analysis of the known drug sensitivities that have emerged in recent years. Meanwhile, the screening of drug response in cancer cell lines provides an estimable genomic and pharmacological data towards high accuracy prediction. Existing works primarily utilize genomic or functional genomic features to classify or regress the drug response. Here in this work, by the migration and extension of the conventional merchandise recommendation methods, we introduce an innovation model on accurate drug sensitivity prediction by using dual-layer strengthened collaborative topic regression (DS-CTR), which incorporates not only the graphic model to jointly learn drugs and cell lines feature from pharmacogenomics data but also drug and cell line similarity network model to strengthen the correlation of the prediction results. Using Genomics of Drug Sensitivity in Cancer project (GDSC) as benchmark datasets, the 5-fold cross-validation experiment demonstrates that DS-CTR model significantly improves drug response prediction performance compared with four categories of state-of-the-art algorithms as for both Receiver Operator Curve (ROC) and the Area Under Receiver Operator Curve (AUC). By uncovering the unknown cell-drug associations with advanced literature evidences, our novel model DS-CTR is validated and supported. The model also provides the possibility to make the discovery of new anti-cancer therapeutics in the preclinical trials cheaper and faster.
Jianing Xi, Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 HetRCNA: A Novel Method to Identify Recurrent Copy Number Alternations from Heterogeneous Tumor Samples Based on Matrix Decomposition Framework
abstract
A common strategy to discovering cancer associated copy number aberrations (CNAs) from a cohort of cancer samples is to detect recurrent CNAs (RCNAs). Although the previous methods can successfully identify communal RCNAs shared by nearly all tumor samples, detecting subgroup-specific RCNAs and their related subgroup samples from cancer samples with heterogeneity is still invalid for these existing approaches. In this paper, we introduce a novel integrated method called HetRCNA, which can identify statistically significant subgroup-specific RCNAs and their related subgroup samples. Based on matrix decomposition framework with weight constraint, HetRCNA can successfully measure the subgroup samples by coefficients of left vectors with weight constraint and subgroup-specific RCNAs by coefficients of the right vectors and significance test. When we evaluate HetRCNA on simulated dataset, the results show that HetRCNA gives the best performances among the competing methods and is robust to the noise factors of the simulated data. When HetRCNA is applied on a real breast cancer dataset, our approach successfully identifies a bunch of RCNA regions and the result is highly correlated with the results of the other two investigated approaches. Notably, the genomic regions identified by HetRCNA harbor many breast cancer related genes reported by previous researches.
Jianing Xi, Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 LPGNMF: Predicting Long Non-Coding RNA and Protein Interaction Using Graph Regularized Nonnegative Matrix Factorization
abstract
Long non-coding RNAs (lncRNA) play crucial roles in a variety of biological processes and complex diseases. Massive studies have indicated that lncRNAs interact with related proteins to exert regulation of cellular biological processes. Because it is time-consuming and expensive to determine lncRNA-protein interaction by experiment, more accurate predictions of interaction by computational methods are imperative. We propose a novel computational approach, predicting lncRNA-protein interaction using graph regularized nonnegative matrix factorization (LPGNMF), to discover unobserved lncRNA-protein association. First, we calculate lncRNA similarity and protein similarity by integrating the lncRNA expression information and gene ontology information. Subsequently, we utilize graph regularized nonnegative matrix factorization framework to predict potential interactions for all lncRNA simultaneously. In the cross validation test, LPGNMF achieves an AUC of 85.2 percent, higher than those of other compared methods. In addition, novel lncRNA-protein interactions detected by LPGNMF are validated by literatures or database. The results indicate that our method is effective to discover potential lncRNA-protein interaction.
Jianing Xi, Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 A Novel MKL Method for GBM Prognosis Prediction by Integrating Histopathological Image and Multi-Omics Data
abstract
Glioblastoma multiforme (GBM) is one of the most malignant brain tumors with very short prognosis expectation. To improve patients' clinical treatment and their life quality after surgery, researches have developed tremendous in silico models and tools for predicting GBM prognosis based on molecular datasets and have earned great success. However, pathology still plays the most critical role in cancer diagnosis and prognosis in the clinic at present. Recent advancement of storing and processing histopathological images has drawn attention of researchers. Models based on histopathological images are developed, which show great potential for computer-aided pathological diagnoses. But models based on both molecular and histopathological images that could predict GBM prognosis with high accuracy are not present yet. In our previous research, we used the simple MKL method to integrate multi-omics data to improve GBM prognosis prediction successfully. In this paper, we have developed a novel multiple kernel learning (MKL) method, named histopathological integrating multiple kernel learning (HI-MKL), that could integrate both histopathological images and multi-omics data efficiently. By using datasets from The Cancer Genome Atlas project, we have built a system that could predict the GBM prognosis with high accuracy. Our research shows that HI-MKL is an accurate, robust, and generalized MKL method, which performs well in a GBM prognosis task.
Ao Li 0001
IEEE J. Biomed. Health Informatics2
2019 DeepPhos: prediction of protein phosphorylation sites with deep learning
abstract
MOTIVATION: Phosphorylation is the most studied post-translational modification, which is crucial for multiple biological processes. Recently, many efforts have been taken to develop computational predictors for phosphorylation site prediction, but most of them are based on feature selection and discriminative classification. Thus, it is useful to develop a novel and highly accurate predictor that can unveil intricate patterns automatically for protein phosphorylation sites. RESULTS: In this study we present DeepPhos, a novel deep learning architecture for prediction of protein phosphorylation. Unlike multi-layer convolutional neural networks, DeepPhos consists of densely connected convolutional neuron network blocks which can capture multiple representations of sequences to make final phosphorylation prediction by intra block concatenation layers and inter block concatenation layers. DeepPhos can also be used for kinase-specific prediction varying from group, family, subfamily and individual kinase level. The experimental results demonstrated that DeepPhos outperforms competitive predictors in general and kinase-specific phosphorylation site prediction. AVAILABILITY AND IMPLEMENTATION: The source code of DeepPhos is publicly deposited at https://github.com/USTCHIlab/DeepPhos. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fenglin Luo, Yu Liu 0113, Xing-Ming Zhao, Ao Li 0001
Bioinform.5
2019 A novel approach for drug response prediction in cancer cell lines via network representation learning
abstract
MOTIVATION: Prediction of cancer patient's response to therapeutic agent is important for personalized treatment. Because experimental verification of reactions between large cohort of patients and drugs is time-intensive, expensive and impractical, preclinical prediction model based on large-scale pharmacogenomic of cancer cell line is highly expected. However, most of the existing computational studies are primarily based on genomic profiles of cancer cell lines while ignoring relationships among genes and failing to capture functional similarity of cell lines. RESULTS: In this study, we present a novel approach named NRL2DRP, which integrates protein-protein interactions and captures similarity of cell lines' functional contexts, to predict drug responses. Through integrating genomic aberrations and drug responses information with protein-protein interactions, we construct a large response-related network, where the neighborhood structure of cell line provides a functional context to its therapeutic responses. Representation vectors of cell lines are extracted through network representation learning method, which could preserve vertices' neighborhood similarity and serve as features to build predictor for drug responses. The predictive performance of NRL2DRP is verified by cross-validation on GDSC dataset and methods comparison, where NRL2DRP achieves AUC > 79% for half drugs and outperforms previous methods. The validity of NRL2DRP is also supported by its effectiveness on uncovering accurate novel relationships between cell lines and drugs. Lots of newly predicted drug responses are confirmed by reported experimental evidences. AVAILABILITY AND IMPLEMENTATION: The code and documentation are available on https://github.com/USTC-HIlab/NRL2DRP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianghong Yang, Ao Li 0001, Yongqiang Li 0002, Xiangqian Guo
Bioinform.2
2019 LSCDFS-MKL: A multiple kernel based method for lung squamous cell carcinomas disease-free survival prediction with pathological and genomic data
abstract
Lung squamous cell carcinoma (SCC) is a fatal disease in both male and female, for which current treatments are inadequate. Surgical resection is regarded as the cornerstone of treatment for patients with lung SCC, but even for the same stage patients, the wide spectrum of disease-free survival (DFS) times exits. Therefore, how to improve the DFS prediction performance of lung SCC becomes one major research area. In this study, we proposed a novel method called LSCDFS-MKL, which was on the basis of multiple kernel learning to predict DFS of lung SCC. In LSCDFS-MKL, we first efficiently integrated pathological images and genomic data (copy number aberration, gene expression, protein expression) from lung SCC. The results of LSCDFS-MKL between different types of data show that the features extracted from pathological images play an important role in DFS prediction of lung SCC. Then we compared our method LSCDFS-MKL with other existing methods and performance analysis indicates that LSCDFS-MKL has a significantly better performance than other prediction methods. After that, we applied the proposed method on different stage stratums and the performance demonstrates that LSCDFS-MKL remains efficient in DFS prediction of lung SCC patients. Finally, we performed LSCDFS-MKL on an independent validation dataset and the accuracy of DFS prediction achieves 100%, which is promising.
Aoshuang Zhang, Ao Li 0001
J. Biomed. Informatics2
2019 A Multimodal Deep Neural Network for Human Breast Cancer Prognosis Prediction by Integrating Multi-Dimensional Data
abstract
Breast cancer is a highly aggressive type of cancer with very low median survival. Accurate prognosis prediction of breast cancer can spare a significant number of patients from receiving unnecessary adjuvant systemic treatment and its related expensive medical costs. Previous work relies mostly on selected gene expression data to create a predictive model. The emergence of deep learning methods and multi-dimensional data offers opportunities for more comprehensive analysis of the molecular characteristics of breast cancer and therefore can improve diagnosis, treatment, and prevention. In this study, we propose a Multimodal Deep Neural Network by integrating Multi-dimensional Data (MDNNMD) for the prognosis prediction of breast cancer. The novelty of the method lies in the design of our method's architecture and the fusion of multi-dimensional data. The comprehensive performance evaluation results show that the proposed method achieves a better performance than the prediction methods with single-dimensional data and other existing approaches. The source code implemented by TensorFlow 1.0 deep learning library can be downloaded from the Github: https://github.com/USTC-HIlab/MDNNMD.
Dongdong Sun, Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Discovering mutated driver genes through a robust and sparse co-regularized matrix factorization framework with prior information from mRNA expression patterns and interaction network
abstract
BACKGROUND: Discovery of mutated driver genes is one of the primary objective for studying tumorigenesis. To discover some relatively low frequently mutated driver genes from somatic mutation data, many existing methods incorporate interaction network as prior information. However, the prior information of mRNA expression patterns are not exploited by these existing network-based methods, which is also proven to be highly informative of cancer progressions. RESULTS: To incorporate prior information from both interaction network and mRNA expressions, we propose a robust and sparse co-regularized nonnegative matrix factorization to discover driver genes from mutation data. Furthermore, our framework also conducts Frobenius norm regularization to overcome overfitting issue. Sparsity-inducing penalty is employed to obtain sparse scores in gene representations, of which the top scored genes are selected as driver candidates. Evaluation experiments by known benchmarking genes indicate that the performance of our method benefits from the two type of prior information. Our method also outperforms the existing network-based methods, and detect some driver genes that are not predicted by the competing methods. CONCLUSIONS: In summary, our proposed method can improve the performance of driver gene discovery by effectively incorporating prior information from interaction network and mRNA expression patterns into a robust and sparse co-regularized matrix factorization framework.
Jianing Xi, Ao Li 0001
BMC Bioinform.3
2018 A novel unsupervised learning model for detecting driver genes from pan-cancer data through matrix tri-factorization framework with pairwise similarities constraints
Jianing Xi, Ao Li 0001
Neurocomputing2
2017 Anaconda: AN automated pipeline for somatic COpy Number variation Detection and Annotation from tumor exome sequencing data
abstract
BACKGROUND: Copy number variations (CNVs) are the main genetic structural variations in cancer genome. Detecting CNVs in genetic exome region is efficient and cost-effective in identifying cancer associated genes. Many tools had been developed accordingly and yet these tools lack of reliability because of high false negative rate, which is intrinsically caused by genome exonic bias. RESULTS: To provide an alternative option, here, we report Anaconda, a comprehensive pipeline that allows flexible integration of multiple CNV-calling methods and systematic annotation of CNVs in analyzing WES data. Just by one command, Anaconda can generate CNV detection result by up to four CNV detecting tools. Associated with comprehensive annotation analysis of genes involved in shared CNV regions, Anaconda is able to deliver a more reliable and useful report in assistance with CNV-associate cancer researches. CONCLUSION: Anaconda package and manual can be freely accessed at http://mcg.ustc.edu.cn/bsc/ANACONDA/ .
Jianing Gao, Changlin Wan, Ao Li 0001, Qiguang Zang, Rongjun Ban, Asim Ali, Zhenghua Yu, Qinghua Shi, Xiaohua Jiang, Yuanwei Zhang
BMC Bioinform.4
2017 A Heterogeneous Network Based Method for Identifying GBM-Related Genes by Integrating Multi-Dimensional Data
abstract
The emergence of multi-dimensional data offers opportunities for more comprehensive analysis of the molecular characteristics of human diseases and therefore improving diagnosis, treatment, and prevention. In this study, we proposed a heterogeneous network based method by integrating multi-dimensional data (HNMD) to identify GBM-related genes. The novelty of the method lies in that the multi-dimensional data of GBM from TCGA dataset that provide comprehensive information of genes, are combined with protein-protein interactions to construct a weighted heterogeneous network, which reflects both the general and disease-specific relationships between genes. In addition, a propagation algorithm with resistance is introduced to precisely score and rank GBM-related genes. The results of comprehensive performance evaluation show that the proposed method significantly outperforms the network based methods with single-dimensional data and other existing approaches. Subsequent analysis of the top ranked genes suggests they may be functionally implicated in GBM, which further corroborates the superiority of the proposed method. The source code and the results of HNMD can be downloaded from the following URL: http://bioinformatics.ustc.edu.cn/hnmd/ .
Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 IsomiR Bank: a research resource for tracking IsomiRs
abstract
UNLABELLED: : Next-Generation Sequencing (NGS) technology has revealed that microRNAs (miRNAs) are capable of exhibiting frequent differences from their corresponding mature reference sequences, generating multiple variants: the isoforms of miRNAs (isomiRs). These isomiRs mainly originate via the imprecise and alternative cleavage during the pre-miRNA processing and post-transcriptional modifications that influence miRNA stability, their sub-cellular localization and target selection. Although several tools for the identification of isomiR have been reported, no bioinformatics resource dedicated to gather isomiRs from public NGS data and to provide functional analysis of these isomiRs is available to date. Thus, a free online database, IsomiR Bank has been created to integrate isomiRs detected by our previously published algorithm CPSS. In total, 2727 samples (Small RNA NGS data downloaded from ArrayExpress) from eight species (Arabidopsis thaliana, Drosophila melanogaster, Danio rerio, Homo sapiens, Mus musculus, Oryza sativa, Solanum lycopersicum and Zea mays) are analyzed. At present, 308 919 isomiRs from 4706 mature miRNAs are collected into IsomiR Bank. In addition, IsomiR Bank provides target prediction and enrichment analysis to evaluate the effects of isomiRs on target selection. AVAILABILITY AND IMPLEMENTATION: IsomiR Bank is implemented in PHP/PERL + MySQL + R format and can be freely accessed at http://mcg.ustc.edu.cn/bsc/isomir/ CONTACTS: : [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuanwei Zhang, Qiguang Zang, Rongjun Ban, Yifan Yang 0001, QiaoMei Hao, Furhan Iqbal, Ao Li 0001, Qinghua Shi
Bioinform.10
2016 CloneCNA: detecting subclonal somatic copy number alterations in heterogeneous tumor samples from whole-exome sequencing data
abstract
BACKGROUND: Copy number alteration is a main genetic structural variation that plays an important role in tumor initialization and progression. Accurate detection of copy number alterations is necessary for discovering cancer-causing genes. Whole-exome sequencing has become a widely used technology in the last decade for detecting various types of genomic aberrations in cancer genomes. However, there are several major issues encountered in these detection problems, including normal cell contamination, tumor aneuploidy, and intra-tumor heterogeneity. Especially, deciphering the intra-tumor heterogeneity is imperative for identifying clonal and subclonal copy number alterations. RESULTS: We introduce CloneCNA, a novel bioinformatics tool for efficiently addressing these issues and automatically detecting clonal and subclonal somatic copy number alterations from heterogeneous tumor samples. CloneCNA fully explores the log ratio of read counts between paired tumor-normal samples and tumor B allele frequency of germline heterozygous SNP positions, further employs efficient statistical models to quantitatively represent copy number status of tumor sample containing multiple clones. We examine CloneCNA on simulated heterogeneous and real tumor samples, and the results demonstrate that CloneCNA has higher power to detect copy number alterations than existing methods. CONCLUSIONS: CloneCNA, a novel algorithm is developed to efficiently and accurately identify somatic copy number alterations from heterogeneous tumor samples. We demonstrate the statistical framework of CloneCNA represents a remarkable advance for tumor whole-exome sequencing data. We expect that CloneCNA will promote cancer-focused studies for investigating the role of clonal evolution and elucidating critical events benefiting tumor tumourigenesis and progression.
Zhenhua Yu 0002, Ao Li 0001
BMC Bioinform.2
2016 A text feature-based approach for literature mining of lncRNA-protein interactions
Ao Li 0001, Qiguang Zang, Dongdong Sun
Neurocomputing1
2016 Data mining in systems biology
Ao Li 0001, Xing-Ming Zhao, Shuigeng Zhou
Neurocomputing1
2016 Relevance search for predicting lncRNA-protein interactions based on heterogeneous network
Jianghong Yang, Ao Li 0001, Mengqu Ge
Neurocomputing2
2016 Discovering Recurrent Copy Number Aberrations in Complex Patterns via Non-Negative Sparse Singular Value Decomposition
abstract
Recurrent copy number aberrations (RCNAs) in multiple cancer samples are strongly associated with tumorigenesis, and RCNA discovery is helpful to cancer research and treatment. Despite the emergence of numerous RCNA discovering methods, most of them are unable to detect RCNAs in complex patterns that are influenced by complicating factors including aberration in partial samples, co-existing of gains and losses and normal-like tumor samples. Here, we propose a novel computational method, called non-negative sparse singular value decomposition (NN-SSVD), to address the RCNA discovering problem in complex patterns. In NN-SSVD, the measurement of RCNA is based on the aberration frequency in a part of samples rather than all samples, which can circumvent the complexity of different RCNA patterns. We evaluate NN-SSVD on synthetic dataset by comparison on detection scores and Receiver Operating Characteristics curves, and the results show that NN-SSVD outperforms existing methods in RCNA discovery and demonstrate more robustness to RCNA complicating factors. Applying our approach on a breast cancer dataset, we successfully identify a number of genomic regions that are strongly correlated with previous studies, which harbor a bunch of known breast cancer associated genes.
Jianing Xi, Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 Improve Glioblastoma Multiforme Prognosis Prediction by Using Feature Selection and Multiple Kernel Learning
abstract
Glioblastoma multiforme (GBM) is a highly aggressive type of brain cancer with very low median survival. In order to predict the patient's prognosis, researchers have proposed rules to classify different glioma cancer cell subtypes. However, survival time of different subtypes of GBM is often various due to different individual basis. Recent development in gene testing has evolved classic subtype rules to more specific classification rules based on single biomolecular features. These classification methods are proven to perform better than traditional simple rules in GBM prognosis prediction. However, the real power behind the massive data is still under covered. We believe a combined prediction model based on more than one data type could perform better, which will contribute further to clinical treatment of GBM. The Cancer Genome Atlas (TCGA) database provides huge dataset with various data types of many cancers that enables us to inspect this aggressive cancer in a new way. In this research, we have improved GBM prognosis prediction accuracy further by taking advantage of the minimum redundancy feature selection method (mRMR) and Multiple Kernel Machine (MKL) learning method. Our goal is to establish an integrated model which could predict GBM prognosis with high accuracy.
Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2015 RPdb: a database of experimentally verified cellular reprogramming records
abstract
UNLABELLED: Many cell lines can be reprogrammed to other cell lines by forced expression of a few transcription factors or by specifically designed culture methods, which have attracted a great interest in the field of regenerative medicine and stem cell research. Plenty of cell lines have been used to generate induced pluripotent stem cells (IPSCs) by expressing a group of genes and microRNAs. These IPSCs can differentiate into somatic cells to promote tissue regeneration. Similarly, many somatic cells can be directly reprogrammed to other cells without a stem cell state. All these findings are helpful in searching for new reprogramming methods and understanding the biological mechanism inside. However, to the best of our knowledge, there is still no database dedicated to integrating the reprogramming records. We built RPdb (cellular reprogramming database) to collect cellular reprogramming information and make it easy to access. All entries in RPdb are manually extracted from more than 2000 published articles, which is helpful for researchers in regenerative medicine and cell biology. AVAILABILITY AND IMPLEMENTATION: RPdb is freely available on the web at http://bioinformatics.ustc.edu.cn/rpdb with all major browsers supported. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ao Li 0001
Bioinform.4
2015 Identification of Genomic Aberrations in Cancer Subclones from Heterogeneous Tumor Samples
abstract
Tumor samples are usually heterogeneous, containing admixture of more than one kind of tumor subclones. Studies of genomic aberrations from heterogeneous tumor data are hindered by the mixed signal of tumor subclone cells. Most of the existing algorithms cannot distinguish contributions of different subclones from the measured single nucleotide polymorphism (SNP) array signals, which may cause erroneous estimation of genomic aberrations. Here, we have introduced a computational method, Cancer Heterogeneity Analysis from SNP-array Experiments (CHASE), to automatically detect subclone proportions and genomic aberrations from heterogeneous tumor samples. Our method is based on HMM, and incorporates EM algorithm to build a statistical model for modeling mixed signal of multiple tumor subclones. We tested the proposed approach on simulated datasets and two real datasets, and the results show that the proposed method can efficiently estimate tumor subclone proportions and recovery the genomic aberrations.
Hong Xia, Yuanning Liu, Ao Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2014 Prediction of human disease-specific phosphorylation sites with combined feature selection approach and support vector machine
abstract
Phosphorylation is a crucial post translational modification, which regulates almost all cellular process in life. It has long been recognized that protein phosphorylation has close relationship with diseases, and therefore many researches are undertaken to predict phosphorylation sites for disease treatment and drug design. However, despite the success achieved by these approaches, no method focuses on disease-associated phosphorylation sites prediction. Herein, for the first time we propose a novel approach that is specially designed to identify disease-specific phosphorylation sites based on SVM. Human disease-associated phosphorylation data is extracted from PhosphoSitePlus database and local sequences are derived for training. To take full advantage of sequence information, a combined feature selection method-based SVM (CFS-SVM) that incorporates mRMR filtering process and forward feature selection process is developed. With CFS-SVM, we successfully predict disease-specific phosphorylation sites. Performance evaluation shows that CFS-SVM is significantly better than the widely used classifiers, including Bayesian decision theory and k nearest neighbour. With the extremely high specificity of 99%, CFS-SVM can still achieve a high sensitivity. Besides, the analysis of corresponding kinases and selected features also shed light on understanding of the potential mechanism of disease-phosphorylation relationships and guide further experimental validations.
Xiaoyi Xu, Ao Li 0001
BIBM2
2014 CLImAT: accurate detection of copy number alteration and loss of heterozygosity in impure and aneuploid tumor samples using whole-genome sequencing data
abstract
MOTIVATION: Whole-genome sequencing of tumor samples has been demonstrated as an efficient approach for comprehensive analysis of genomic aberrations in cancer genome. Critical issues such as tumor impurity and aneuploidy, GC-content and mappability bias have been reported to complicate identification of copy number alteration and loss of heterozygosity in complex tumor samples. Therefore, efficient computational methods are required to address these issues. RESULTS: We introduce CLImAT (CNA and LOH Assessment in Impure and Aneuploid Tumors), a bioinformatics tool for identification of genomic aberrations from tumor samples using whole-genome sequencing data. Without requiring a matched normal sample, CLImAT takes integrated analysis of read depth and allelic frequency and provides extensive data processing procedures including GC-content and mappability correction of read depth and quantile normalization of B-allele frequency. CLImAT accurately identifies copy number alteration and loss of heterozygosity even for highly impure tumor samples with aneuploidy. We evaluate CLImAT on both simulated and real DNA sequencing data to demonstrate its ability to infer tumor impurity and ploidy and identify genomic aberrations in complex tumor samples. AVAILABILITY AND IMPLEMENTATION: The CLImAT software package can be freely downloaded at http://bioinformatics.ustc.edu.cn/CLImAT/.
Zhenhua Yu 0002, Yuanning Liu, Ao Li 0001
Bioinform.5
2013 PKIS: computational identification of protein Kinases for experimentally discovered protein Phosphorylation sites
abstract
BACKGROUND: Dynamic protein phosphorylation is an essential regulatory mechanism in various organisms. In this capacity, it is involved in a multitude of signal transduction pathways. Kinase-specific phosphorylation data lay the foundation for reconstruction of signal transduction networks. For this reason, precise annotation of phosphorylated proteins is the first step toward simulating cell signaling pathways. However, the vast majority of kinase-specific phosphorylation data remain undiscovered and existing experimental methods and computational phosphorylation site (P-site) prediction tools have various limitations with respect to addressing this problem. RESULTS: To address this issue, a novel protein kinase identification web server, PKIS, is here presented for the identification of the protein kinases responsible for experimentally verified P-sites at high specificity, which incorporates the composition of monomer spectrum (CMS) encoding strategy and support vector machines (SVMs). Compared to widely used P-site prediction tools including KinasePhos 2.0, Musite, and GPS2.1, PKIS largely outperformed these tools in identifying protein kinases associated with known P-sites. In addition, PKIS was used on all the P-sites in Phospho.ELM that currently lack kinase information. It successfully identified 14 potential SYK substrates with 36 known P-sites. Further literature search showed that 5 of them were indeed phosphorylated by SYK. Finally, an enrichment analysis was performed and 6 significant SYK-related signal pathways were identified. CONCLUSIONS: In general, PKIS can identify protein kinases for experimental phosphorylation sites efficiently. It is a valuable bioinformatics tool suitable for the study of protein phosphorylation. The PKIS web server is freely available at http://bioinformatics.ustc.edu.cn/pkis.
Liang Zou, Mang Wang 0001, Ao Li 0001
BMC Bioinform.5
2006 Missing value estimation for DNA microarray gene expression data by Support Vector Regression imputation and orthogonal coding scheme
abstract
BACKGROUND: Gene expression profiling has become a useful biological resource in recent years, and it plays an important role in a broad range of areas in biology. The raw gene expression data, usually in the form of large matrix, may contain missing values. The downstream analysis methods that postulate complete matrix input are thus not applicable. Several methods have been developed to solve this problem, such as K nearest neighbor impute method, Bayesian principal components analysis impute method, etc. In this paper, we introduce a novel imputing approach based on the Support Vector Regression (SVR) method. The proposed approach utilizes an orthogonal coding input scheme, which makes use of multi-missing values in one row of a certain gene expression profile and imputes the missing value into a much higher dimensional space, to obtain better performance. RESULTS: A comparative study of our method with the previously developed methods has been presented for the estimation of the missing values on six gene expression data sets. Among the three different input-vector coding schemes we tried, the orthogonal input coding scheme obtains the best estimation results with the minimum Normalized Root Mean Squared Error (NRMSE). The results also demonstrate that the SVR method has powerful estimation ability on different kinds of data sets with relatively small NRMSE. CONCLUSION: The SVR impute method shows better performance than, or at least comparable with, the previously developed methods in present research. The outstanding estimation ability of this impute method is partly due to the use of the most missing value information by incorporating orthogonal input coding scheme. In addition, the solid theoretical foundation of SVR method also helps in estimation of performance together with orthogonal input coding scheme. The promising estimation ability demonstrated in the results section suggests that the proposed approach provides a proper solution to the missing value estimation problem. The source code of the SVR method is available from http://202.38.78.189/downloads/svrimpute.html for non-commercial use.
Ao Li 0001, Huanqing Feng
BMC Bioinform.2
2006 PPSP: prediction of PK-specific phosphorylation site with Bayesian decision theory
abstract
BACKGROUND: As a reversible and dynamic post-translational modification (PTM) of proteins, phosphorylation plays essential regulatory roles in a broad spectrum of the biological processes. Although many studies have been contributed on the molecular mechanism of phosphorylation dynamics, the intrinsic feature of substrates specificity is still elusive and remains to be delineated. RESULTS: In this work, we present a novel, versatile and comprehensive program, PPSP (Prediction of PK-specific Phosphorylation site), deployed with approach of Bayesian decision theory (BDT). PPSP could predict the potential phosphorylation sites accurately for approximately 70 PK (Protein Kinase) groups. Compared with four existing tools Scansite, NetPhosK, KinasePhos and GPS, PPSP is more accurate and powerful than these tools. Moreover, PPSP also provides the prediction for many novel PKs, say, TRK, mTOR, SyK and MET/RON, etc. The accuracy of these novel PKs are also satisfying. CONCLUSION: Taken together, we propose that PPSP could be a potentially powerful tool for the experimentalists who are focusing on phosphorylation substrates with their PK-specific sites identification. Moreover, the BDT strategy could also be a ubiquitous approach for PTMs, such as sumoylation and ubiquitination, etc.
Yu Xue 0001, Ao Li 0001, Lirong Wang, Huanqing Feng, Xuebiao Yao
BMC Bioinform.2