VLDB 2026 Research / reviewers in the wild / expert
Jialiang Yang
dblp:84/3864
· DBLP profile ↗
30ranked-venue papers
3as first author
19since 2021 · last 2027
0000-0003-4689-8672ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | SMGC: Spatial multi-omics analysis with granular-ball contrastive learning framework
Xuejing Ma, Zijia Bai, Yajie Meng, Pan Zeng, Xianfang Tang, Feifei Cui, Peng Wang 0035, Jialiang Yang, Junlin Xu |
Expert Syst. Appl. | 10 |
| 2026 | EccoMamba: Enhanced Cross-hierarchical Continuity Orthogonal Mamba for Medical Image SegmentationabstractMedical image segmentation plays a crucial role in clinical diagnosis, lesion quantification, and preoperative planning. However, existing Mamba-based architectures, which rely on fixed-direction sequence modeling and flatten images into one-dimensional (1D) sequences, struggle to capture hierarchical anatomical features and spatial dependencies, thereby limiting their representational capacity for complex medical structures. To address these limitations, we propose EccoMamba (Enhanced Cross-hierarchical Continuity Orthogonal Mamba), a U-shaped encoder--decoder framework designed for medical image segmentation. In the encoder's downsampling path, we introduce a Hierarchical Aggregation Enhancement (HAE) module that integrates multi-scale convolutions with hierarchical attention mechanisms. The attention branch further incorporates cross-channel interactions, allowing the model to selectively enhance semantically relevant features while suppressing irrelevant background responses. For skip connections, we design a Structural Continuity Orthogonal (SCO) module to preserve spatial continuity by modeling cross-dimensional dependencies via orthogonal Axial Shifts (AS), thereby mitigating directional bias and improving anatomical consistency. Extensive experiments on four benchmark datasets---ISIC 2018, ISIC 2017, Synapse, and ACDC---show that EccoMamba consistently outperforms state-of-the-art methods in both segmentation accuracy and structural fidelity. Junlin Xu, Jincan Li, Feifei Cui, Jialiang Yang, Shuting Jin, Qiangguo Jin, Yajie Meng |
AAAI | 5 |
| 2026 | SpatialSyn: A synergistic graph framework for spatial domain identification in multi-omics
Pan Zeng, Runzhi Li, Yajie Meng, Feifei Cui, Xianfang Tang, Jialiang Yang, Junlin Xu |
Expert Syst. Appl. | 7 |
| 2026 | PDGCL-DTI: Parallel Dual-Channel Graph Contrastive Learning for Drug-Target Binding Prediction in Heterogeneous NetworksabstractPredicting drug-target interactions (DTI) is critical for advancing drug discovery. However, existing DTI approaches struggle with data imbalance and heterogeneous information. This study presents a novel framework called PDGCL-DTI, which leverages two graph contrastive learning frameworks in parallel to capture both local and global features from drug-target heterogeneous networks. First, PDGCL-DTI effectively handles data imbalance through the AdaL-GCL module, which dynamically adjusts the weights of minority class samples to mitigate the impact of the imbalance. Second, by combining local and global contrastive learning, it extracts features from both local node information and global structural information, improving its adaptability to complex heterogeneous networks. This dual strategy enables PDGCL-DTI to exhibit greater robustness and higher prediction accuracy when handling complex DTI data. Experimental results on the ChEMBL, DrugBank, and DAVIS datasets show that PDGCL-DTI outperforms existing DTI methods, achieving an average AUC of 0.958 and an average accuracy of 0.95 across the three datasets. Additionally, case studies demonstrate that PDGCL-DTI successfully predicts interactions between Enasidenib and GABA-AT, as well as Sorafenib and Caspase-3, underscoring its practical applicability in the visualization workflow on the ChEMBL dataset. Qihui Zheng, Xianfang Tang, Yajie Meng, Junlin Xu, Xueying Zeng 0001, Geng Tian, Jialiang Yang |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | MMsurv: a multimodal multi-instance multi-cancer survival prediction model integrating pathological images, clinical information, and sequencing dataabstractAccurate prediction of patient survival rates in cancer treatment is essential for effective therapeutic planning. Unfortunately, current models often underutilize the extensive multimodal data available, affecting confidence in predictions. This study presents MMSurv, an interpretable multimodal deep learning model to predict survival in different types of cancer. MMSurv integrates clinical information, sequencing data, and hematoxylin and eosin-stained whole-slide images (WSIs) to forecast patient survival. Specifically, we segment tumor regions from WSIs into image tiles and employ neural networks to encode each tile into one-dimensional feature vectors. We then optimize clinical features by applying word embedding techniques, inspired by natural language processing, to the clinical data. To better utilize the complementarity of multimodal data, this study proposes a novel fusion method, multimodal fusion method based on compact bilinear pooling and transformer, which integrates bilinear pooling with Transformer architecture. The fused features are then processed through a dual-layer multi-instance learning model to remove prognosis-irrelevant image patches and predict each patient's survival risk. Furthermore, we employ cell segmentation to investigate the cellular composition within the tiles that received high attention from the model, thereby enhancing its interpretive capacity. We evaluate our approach on six cancer types from The Cancer Genome Atlas. The results demonstrate that utilizing multimodal data leads to higher predictive accuracy compared to using single-modal image data, with an average C-index increase from 0.6750 to 0.7283. Additionally, we compare our proposed baseline model with state-of-the-art methods using the C-index and five-fold cross-validation approach, revealing a significant average improvement of nearly 10% in our model's performance. Shufang Shi, Lijing Liu, Yuhua Yao, Geng Tian, Peizhen Wang, Jialiang Yang |
Briefings Bioinform. | 9 |
| 2025 | Deep learning-based fusion of nuclear segmentation features for microsatellite instability and tumor mutational burden prediction in digestive tract cancers: a multicenter validation studyabstractMicrosatellite instability (MSI) and tumor mutational burden (TMB) are crucial biomarkers in gastric (GC) and colorectal cancer (CRC), yet their conventional sequencing-based detection is costly and time-consuming. Since only ~20% of patients are MSI-high or TMB-high and likely to benefit from immunotherapy, expensive genomic testing is often unjustified. This study developed a deep learning framework to predict MSI and TMB status directly from routinely available Hematoxylin and Eosin (H&E)-stained whole-slide images, leveraging fused nuclear segmentation features to improve accuracy. Using samples from TCGA (350 GC and 376 CRC for MSI; 400 GC and 387 CRC for TMB), image features were extracted with CLAM and nuclear features with Hover-Net. These features were combined via Multimodal Compact Bilinear Pooling and utilized in six distinct deep learning models. By fusing the nucleus segmentation features, the model increased area under the receiver operating characteristic curve (AUC) by 1%-3% and recall by 5%-11% in five-fold cross-validation, significantly outperforming models that relied solely on image features. External validation on a CRC dataset from the China-Japan Friendship hospital further validated the model's robustness, achieving an AUC of 0.81 and a recall of 0.80 for MSI prediction. Additionally, notable differences in cellular composition were observed across cancer types and clinical groups, emphasizing the pivotal role of cellular features in cancer development. These findings highlight the advantages of integrating H&E-stained image features with nuclear segmentation data and advanced deep learning techniques to improve predictive accuracy and reduce the cost of MSI/TMB testing, potentially advancing personalized cancer treatment strategies. Jiaying Han, Fengyuan Hu, Geng Tian, Dingrong Zhong, Jialiang Yang |
Briefings Bioinform. | 8 |
| 2025 | CAMIL: channel attention-based multiple instance learning for whole slide image classificationabstractMOTIVATION: The classification task based on whole-slide images (WSIs) is a classic problem in computational pathology. Multiple instance learning (MIL) provides a robust framework for analyzing whole slide images with slide-level labels at gigapixel resolution. However, existing MIL models typically focus on modeling the relationships between instances while neglecting the variability across the channel dimensions of instances, which prevents the model from fully capturing critical information in the channel dimension. RESULTS: To address this issue, we propose a plug-and-play module called Multi-scale Channel Attention Block (MCAB), which models the interdependencies between channels by leveraging local features with different receptive fields. By alternately stacking four layers of Transformer and MCAB, we designed a channel attention-based MIL model (CAMIL) capable of simultaneously modeling both inter-instance relationships and intra-channel dependencies. To verify the performance of the proposed CAMIL in classification tasks, several comprehensive experiments were conducted across three datasets: Camelyon16, TCGA-NSCLC, and TCGA-RCC. Empirical results demonstrate that, whether the feature extractor is pretrained on natural images or on WSIs, our CAMIL surpasses current state-of-the-art MIL models across multiple evaluation metrics. AVAILABILITY AND IMPLEMENTATION: All implementation code is available at https://github.com/maojy0914/CAMIL. Jinyang Mao, Junlin Xu, Xianfang Tang, Heaven Zhao, Geng Tian, Jialiang Yang |
Bioinform. | 7 |
| 2025 | CDPMF-DDA: contrastive deep probabilistic matrix factorization for drug-disease association predictionabstractThe process of new drug development is complex, whereas drug-disease association (DDA) prediction aims to identify new therapeutic uses for existing medications. However, existing graph contrastive learning approaches typically rely on single-view contrastive learning, which struggle to fully capture drug-disease relationships. Subsequently, we introduce a novel multi-view contrastive learning framework, named CDPMF-DDA, which enhances the model's ability to capture drug-disease associations by incorporating diverse information representations from different views. First, we decompose the original drug-disease association matrix into drug and disease feature matrices, which are then used to reconstruct the drug-disease association network, as well as the drug-drug and disease-disease similarity networks. This process effectively reduces noise in the data, establishing a reliable foundation for the networks produced. Next, we generate multiple contrastive views from both the original and generated networks. These views effectively capture hidden feature associations, significantly enhancing the model's ability to represent complex relationships. Extensive cross-validation experiments on three standard datasets show that CDPMF-DDA achieves an average AUC of 0.9475 and an AUPR of 0.5009, outperforming existing models. Additionally, case studies on Alzheimer's disease and epilepsy further validate the model's effectiveness, demonstrating its high accuracy and robustness in drug-disease association prediction. Based on a multi-view contrastive learning framework, CDPMF-DDA is capable of integrating multi-source information and effectively capturing complex drug-disease associations, making it a powerful tool for drug repositioning and the discovery of new therapeutic strategies. Xianfang Tang, Yawen Hou, Yajie Meng, Zhaojing Wang, Changcheng Lu, Juan Lv, Xinrong Hu, Junlin Xu, Jialiang Yang |
BMC Bioinform. | 9 |
| 2025 | Adaptive debiasing learning for drug repositioning
Yajie Meng, Xinrong Hu, Changcheng Lu, Xianfang Tang, Feifei Cui, Pan Zeng, Yuhua Yao, Jialiang Yang, Junlin Xu |
J. Biomed. Informatics | 9 |
| 2025 | Predicting drug-target interactions based on multivariate information fusion and graph contrast learning
Siying Yang, Ping-An He 0001, Pan Zeng, Yajie Meng, Feifei Cui, Yuhua Yao, Jialiang Yang, Junlin Xu |
J. Biomed. Informatics | 8 |
| 2025 | SWMA-UNet: Multi-Path Attention Network for Improved Medical Image SegmentationabstractIn recent years, deep learning achieves significant advancements in medical image segmentation. Research finds that integrating Transformers and CNNs effectively addresses the limitations of CNNs in managing long-distance dependencies and understanding global information.However, existing models typically employ a serial approach to combine Transformers and CNNs, which complicates the simultaneous processing of global and local information. To address this, our study proposes a parallel multi-path attention architecture, SWMA-UNET, that integrates Transformers and CNNs. This architecture deeply mines features through parallel strategies while capturing both local details and global context information, thereby enhancing the accuracy of medical image segmentation. Experimental results indicate that our method surpasses all previously reported methods in the literature on the Synapse, ACDC, ISIC 2018 and MoNuSeg datasets. Xianfang Tang, Jincan Li, Qianrui Liu, Chang Zhou 0007, Pan Zeng, Yajie Meng, Junlin Xu, Geng Tian, Jialiang Yang |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | Enhancing Drug Repositioning Through Local Interactive Learning With Bilinear Attention NetworksabstractDrug repositioning has emerged as a promising strategy for identifying new therapeutic applications for existing drugs. In this study, we present DRGBCN, a novel computational method that integrates heterogeneous information through a deep bilinear attention network to infer potential drugs for specific diseases. DRGBCN involves constructing a comprehensive drug-disease network by incorporating multiple similarity networks for drugs and diseases. Firstly, we introduce a layer attention mechanism to effectively learn the embeddings of graph convolutional layers from these networks. Subsequently, a bilinear attention network is constructed to capture pairwise local interactions between drugs and diseases. This combined approach enhances the accuracy and reliability of predictions. Finally, a multi-layer perceptron module is employed to evaluate potential drugs. Through extensive experiments on three publicly available datasets, DRGBCN demonstrates better performance over baseline methods in 10-fold cross-validation, achieving an average area under the receiver operating characteristic curve (AUROC) of 0.9399. Furthermore, case studies on bladder cancer and acute lymphoblastic leukemia confirm the practical application of DRGBCN in real-world drug repositioning scenarios. Importantly, our experimental results from the drug-disease network analysis reveal the successful clustering of similar drugs within the same community, providing valuable insights into drug-disease interactions. In conclusion, DRGBCN holds significant promise for uncovering new therapeutic applications of existing drugs, thereby contributing to the advancement of precision medicine. Xianfang Tang, Chang Zhou 0007, Changcheng Lu, Yajie Meng, Junlin Xu, Xinrong Hu, Geng Tian, Jialiang Yang |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | Drug repositioning based on weighted local information augmented graph neural networkabstractDrug repositioning, the strategy of redirecting existing drugs to new therapeutic purposes, is pivotal in accelerating drug discovery. While many studies have engaged in modeling complex drug-disease associations, they often overlook the relevance between different node embeddings. Consequently, we propose a novel weighted local information augmented graph neural network model, termed DRAGNN, for drug repositioning. Specifically, DRAGNN firstly incorporates a graph attention mechanism to dynamically allocate attention coefficients to drug and disease heterogeneous nodes, enhancing the effectiveness of target node information collection. To prevent excessive embedding of information in a limited vector space, we omit self-node information aggregation, thereby emphasizing valuable heterogeneous and homogeneous information. Additionally, average pooling in neighbor information aggregation is introduced to enhance local information while maintaining simplicity. A multi-layer perceptron is then employed to generate the final association predictions. The model's effectiveness for drug repositioning is supported by a 10-times 10-fold cross-validation on three benchmark datasets. Further validation is provided through analysis of the predicted associations using multiple authoritative data sources, molecular docking experiments and drug-disease network analysis, laying a solid foundation for future drug discovery. Yajie Meng, Junlin Xu, Changcheng Lu, Xianfang Tang, Ben-gong Zhang, Geng Tian, Jialiang Yang |
Briefings Bioinform. | 9 |
| 2024 | Drug repositioning based on tripartite cross-network embedding and graph convolutional network
Pan Zeng, Bofei Zhang, Aohang Liu, Yajie Meng, Xianfang Tang, Jialiang Yang, Junlin Xu |
Expert Syst. Appl. | 6 |
| 2022 | DGHNE: network enhancement-based method in identifying disease-causing genes through a heterogeneous biomedical networkabstractThe identification of disease-causing genes is critical for mechanistic understanding of disease etiology and clinical manipulation in disease prevention and treatment. Yet the existing approaches in tackling this question are inadequate in accuracy and efficiency, demanding computational methods with higher identification power. Here, we proposed a new method called DGHNE to identify disease-causing genes through a heterogeneous biomedical network empowered by network enhancement. First, a disease-disease association network was constructed by the cosine similarity scores between phenotype annotation vectors of diseases, and a new heterogeneous biomedical network was constructed by using disease-gene associations to connect the disease-disease network and gene-gene network. Then, the heterogeneous biomedical network was further enhanced by using network embedding based on the Gaussian random projection. Finally, network propagation was used to identify candidate genes in the enhanced network. We applied DGHNE together with five other methods into the most updated disease-gene association database termed DisGeNet. Compared with all other methods, DGHNE displayed the highest area under the receiver operating characteristic curve and the precision-recall curve, as well as the highest precision and recall, in both the global 5-fold cross-validation and predicting new disease-gene associations. We further performed DGHNE in identifying the candidate causal genes of Parkinson's disease and diabetes mellitus, and the genes connecting hyperglycemia and diabetes mellitus. In all cases, the predicted causing genes were enriched in disease-associated gene ontology terms and Kyoto Encyclopedia of Genes and Genomes pathways, and the gene-disease associations were highly evidenced by independent experimental studies. Binsheng He, Ju Xiang, Pingping Bing, Geng Tian, Cheng Guo 0010, Jialiang Yang |
Briefings Bioinform. | 9 |
| 2022 | A weighted bilinear neural collaborative filtering approach for drug repositioningabstractDrug repositioning is an efficient and promising strategy for traditional drug discovery and development. Many research efforts are focused on utilizing deep-learning approaches based on a heterogeneous network for modeling complex drug-disease associations. Similar to traditional latent factor models, which directly factorize drug-disease associations, they assume the neighbors are independent of each other in the network and thus tend to be ineffective to capture localized information. In this study, we propose a novel neighborhood and neighborhood interaction-based neural collaborative filtering approach (called DRWBNCF) to infer novel potential drugs for diseases. Specifically, we first construct three networks, including the known drug-disease association network, the drug-drug similarity and disease-disease similarity networks (using the nearest neighbors). To take the advantage of localized information in the three networks, we then design an integration component by proposing a new weighted bilinear graph convolution operation to integrate the information of the known drug-disease association, the drug's and disease's neighborhood and neighborhood interactions into a unified representation. Lastly, we introduce a prediction component, which utilizes the multi-layer perceptron optimized by the α-balanced focal loss function and graph regularization to model the complex drug-disease associations. Benchmarking comparisons on three datasets verified the effectiveness of DRWBNCF for drug repositioning. Importantly, the unknown drug-disease associations predicted by DRWBNCF were validated against clinical trials and three authoritative databases and we listed several new DRWBNCF-predicted potential drugs for breast cancer (e.g. valrubicin and teniposide) and small cell lung cancer (e.g. valrubicin and cytarabine). Yajie Meng, Changcheng Lu, Min Jin 0002, Junlin Xu, Xiangxiang Zeng, Jialiang Yang |
Briefings Bioinform. | 6 |
| 2022 | ICSDA: a multi-modal deep learning model to predict breast cancer recurrence and metastasis risk by integrating pathological, clinical and gene expression dataabstractBreast cancer patients often have recurrence and metastasis after surgery. Predicting the risk of recurrence and metastasis for a breast cancer patient is essential for the development of precision treatment. In this study, we proposed a novel multi-modal deep learning prediction model by integrating hematoxylin & eosin (H&E)-stained histopathological images, clinical information and gene expression data. Specifically, we segmented tumor regions in H&E into image blocks (256 × 256 pixels) and encoded each image block into a 1D feature vector using a deep neural network. Then, the attention module scored each area of the H&E-stained images and combined image features with clinical and gene expression data to predict the risk of recurrence and metastasis for each patient. To test the model, we downloaded all 196 breast cancer samples from the Cancer Genome Atlas with clinical, gene expression and H&E information simultaneously available. The samples were then divided into the training and testing sets with a ratio of 7: 3, in which the distributions of the samples were kept between the two datasets by hierarchical sampling. The multi-modal model achieved an area-under-the-curve value of 0.75 on the testing set better than those based solely on H&E image, sequencing data and clinical data, respectively. This study might have clinical significance in identifying high-risk breast cancer patients, who may benefit from postoperative adjuvant treatment. Yuhua Yao, Yaping Lv, Yuebin Liang, Shuxue Xi, Binbin Ji, Guanglu Zhang, Geng Tian, Xiyue Hu, Jialiang Yang |
Briefings Bioinform. | 13 |
| 2022 | Predicting colorectal cancer tumor mutational burden from histopathological images and clinical information using multi-modal deep learningabstractMOTIVATION: Tumor mutational burden (TMB) is an indicator of the efficacy and prognosis of immune checkpoint therapy in colorectal cancer (CRC). In general, patients with higher TMB values are more likely to benefit from immunotherapy. Though whole-exome sequencing is considered the gold standard for determining TMB, it is difficult to be applied in clinical practice due to its high cost. There are also a few DNA panel-based methods to estimate TMB; however, their detection cost is also high, and the associated wet-lab experiments usually take days, which emphasize the need for faster and cheaper alternatives. RESULTS: In this study, we propose a multi-modal deep learning model based on a residual network (ResNet) and multi-modal compact bilinear pooling to predict TMB status (i.e. TMB high (TMB_H) or TMB low(TMB_L)) directly from histopathological images and clinical data. We applied the model to CRC data from The Cancer Genome Atlas and compared it with four other popular methods, namely, ResNet18, ResNet50, VGG19 and AlexNet. We tested different TMB thresholds, namely, percentiles of 10%, 14.3%, 15%, 16.3%, 20%, 30% and 50%, to differentiate TMB_H and TMB_L.For the percentile of 14.3% (i.e. TMB value 20) and ResNet18, our model achieved an area under the receiver operating characteristic curve of 0.817 after 5-fold cross-validation, which was better than that of other compared models. In addition, we also found that TMB values were significantly associated with the tumor stage and N and M stages. Our study shows that deep learning models can predict TMB status from histopathological images and clinical information only, which is worth clinical application. Kaimei Huang, Binghu Lin, Yankun Liu, Jingwu Li, Geng Tian, Jialiang Yang |
Bioinform. | 7 |
| 2021 | CMF-Impute: an accurate imputation tool for single cell RNA-seq data
Junlin Xu, Wen Zhu, Jialiang Yang |
Bioinform. | 5 |
| 2020 | Identification of Human LncRNA-Disease Association by Fast Kernel Learning-Based Kronecker Regularized Least Squares
Wen Li 0009, Shulin Wang, Junlin Xu, Jialiang Yang |
ICIC (2) | 4 |
| 2020 | CMF-Impute: an accurate imputation tool for single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA-sequencing (scRNA-seq) technology provides a powerful tool for investigating cell heterogeneity and cell subpopulations by allowing the quantification of gene expression at single-cell level. However, scRNA-seq data analysis remains challenging because of various technical noises such as dropout events (i.e. excessive zero counts in the expression matrix). RESULTS: By taking consideration of the association among cells and genes, we propose a novel collaborative matrix factorization-based method called CMF-Impute to impute the dropout entries of a given scRNA-seq expression matrix. We test CMF-Impute and compare it with the other five state-of-the-art methods on six popular real scRNA-seq datasets of various sizes and three simulated datasets. For simulated datasets, CMF-Impute outperforms other methods in imputing the closest dropouts to the original expression values as evaluated by both the sum of squared error and Pearson correlation coefficient. For real datasets, CMF-Impute achieves the most accurate cell classification results in spite of the choice of different clustering methods like SC3 or T-SNE followed by K-means as evaluated by both adjusted rand index and normalized mutual information. Finally, we demonstrate that CMF-Impute is powerful in reconstructing cell-to-cell and gene-to-gene correlation, and in inferring cell lineage trajectories. AVAILABILITY AND IMPLEMENTATION: CMF-Impute is written as a Matlab package which is available at https://github.com/xujunlin123/CMFImpute.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Junlin Xu, Wen Zhu, Jialiang Yang |
Bioinform. | 5 |
| 2019 | Robust subspace clustering via symmetry constrained latent low rank representation with converted nuclear norm
Xian Fang, Zhixin Tie, Feiyang Song, Jialiang Yang |
Neurocomputing | 4 |
| 2019 | A new resource allocation strategy based on the relationship between subproblems for MOEA/D
Peng Wang 0035, Wen Zhu, Haihua Liu, Bo Liao 0002, Xiaohui Wei 0001, Siqi Ren, Jialiang Yang |
Inf. Sci. | 8 |
| 2017 | Matrix completion with side information and its applications in predicting the antigenicity of influenza virusesabstractMOTIVATION: Low-rank matrix completion has been demonstrated to be powerful in predicting antigenic distances among influenza viruses and vaccines from partially revealed hemagglutination inhibition table. Meanwhile, influenza hemagglutinin (HA) protein sequences are also effective in inferring antigenic distances. Thus, it is natural to integrate HA protein sequence information into low-rank matrix completion model to help infer influenza antigenicity, which is critical to influenza vaccine development. RESULTS: We have proposed a novel algorithm called biological matrix completion with side information (BMCSI), which first measures HA protein sequence similarities among influenza viruses (especially on epitopes) and then integrates the similarity information into a low-rank matrix completion model to predict influenza antigenicity. This algorithm exploits both the correlations among viruses and vaccines in serological tests and the power of HA sequence in predicting influenza antigenicity. We applied this model into H3N2 seasonal influenza virus data. Comparing to previous methods, we significantly reduced the prediction root-mean-square error in a 10-fold cross validation analysis. Based on the cartographies constructed from imputed data, we showed that the antigenic evolution of H3N2 seasonal influenza is generally S-shaped while the genetic evolution is half-circle shaped. We also showed that the Spearman correlation between genetic and antigenic distances (among antigenic clusters) is 0.83, demonstrating a globally high correspondence and some local discrepancies between influenza genetic and antigenic evolution. Finally, we showed that 4.4%±1.2% genetic variance (corresponding to 3.11 ± 1.08 antigenic distances) caused an antigenic drift event for H3N2 influenza viruses historically. AVAILABILITY AND IMPLEMENTATION: The software and data for this study are available at http://bi.sky.zstu.edu.cn/BMCSI/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xianhong Li, Yuhua Yao, Fayou Wang, Jiasheng Yang, Hailiang Sun 0001, Ping-An He 0001, Jialiang Yang |
Bioinform. | 12 |
| 2015 | Integrative random forest for gene regulatory network inferenceabstractMOTIVATION: Gene regulatory network (GRN) inference based on genomic data is one of the most actively pursued computational biological problems. Because different types of biological data usually provide complementary information regarding the underlying GRN, a model that integrates big data of diverse types is expected to increase both the power and accuracy of GRN inference. Towards this goal, we propose a novel algorithm named iRafNet: integrative random forest for gene regulatory network inference. RESULTS: iRafNet is a flexible, unified integrative framework that allows information from heterogeneous data, such as protein-protein interactions, transcription factor (TF)-DNA-binding, gene knock-down, to be jointly considered for GRN inference. Using test data from the DREAM4 and DREAM5 challenges, we demonstrate that iRafNet outperforms the original random forest based network inference algorithm (GENIE3), and is highly comparable to the community learning approach. We apply iRafNet to construct GRN in Saccharomyces cerevisiae and demonstrate that it improves the performance in predicting TF-target gene regulations and provides additional functional insights to the predicted gene regulations. AVAILABILITY AND IMPLEMENTATION: The R code of iRafNet implementation and a tutorial are available at: http://research.mssm.edu/tulab/software/irafnet.html Francesca Petralia, Pei Wang 0014, Jialiang Yang, Zhidong Tu |
Bioinform. | 3 |
| 2013 | BinAligner: a heuristic method to align biological networksabstractThe advances in high throughput omics technologies have made it possible to characterize molecular interactions within and across various species. Alignments and comparison of molecular networks across species will help detect orthologs and conserved functional modules and provide insights on the evolutionary relationships of the compared species. However, such analyses are not trivial due to the complexity of network and high computational cost. Here we develop a mixture of global and local algorithm, BinAligner, for network alignments. Based on the hypotheses that the similarity between two vertices across networks would be context dependent and that the information from the edges and the structures of subnetworks can be more informative than vertices alone, two scoring schema, 1-neighborhood subnetwork and graphlet, were introduced to derive the scoring matrices between networks, besides the commonly used scoring scheme from vertices. Then the alignment problem is formulated as an assignment problem, which is solved by the combinatorial optimization algorithm, such as the Hungarian method. The proposed algorithm was applied and validated in aligning the protein-protein interaction network of Kaposi's sarcoma associated herpesvirus (KSHV) and that of varicella zoster virus (VZV). Interestingly, we identified several putative functional orthologous proteins with similar functions but very low sequence similarity between the two viruses. For example, KSHV open reading frame 56 (ORF56) and VZV ORF55 are helicase-primase subunits with sequence identity 14.6%, and KSHV ORF75 and VZV ORF44 are tegument proteins with sequence identity 15.3%. These functional pairs can not be identified if one restricts the alignment into orthologous protein pairs. In addition, BinAligner identified a conserved pathway between two viruses, which consists of 7 orthologous protein pairs and these proteins are connected by conserved links. This pathway might be crucial for virus packing and infection. Jialiang Yang, Jun Li 0068, Stefan Grünewald, Xiu-Feng Wan |
BMC Bioinform. | 1 |
| 2012 | AntigenMap 3D: an online antigenic cartography resourceabstractSUMMARY: Antigenic cartography is a useful technique to visualize and minimize errors in immunological data by projecting antigens to 2D or 3D cartography. However, a 2D cartography may not be sufficient to capture the antigenic relationship from high-dimensional immunological data. AntigenMap 3D presents an online, interactive, and robust 3D antigenic cartography construction and visualization resource. AntigenMap 3D can be applied to identify antigenic variants and vaccine strain candidates for pathogens with rapid antigenic variations, such as influenza A virus. AVAILABILITY AND IMPLEMENTATION: http://sysbio.cvm.msstate.edu/AntigenMap3D J. Lamar Barnett, Jialiang Yang, Zhipeng Cai 0004, Tong Zhang 0001, Xiu-Feng Wan |
Bioinform. | 2 |
| 2011 | New methods to measure residues coevolution in proteinsabstractBACKGROUND: The covariation of two sites in a protein is often used as the degree of their coevolution. To quantify the covariation many methods have been developed and most of them are based on residues position-specific frequencies by using the mutual information (MI) model. RESULTS: In the paper, we proposed several new measures to incorporate new biological constraints in quantifying the covariation. The first measure is the mutual information with the amino acid background distribution (MIB), which incorporates the amino acid background distribution into the marginal distribution of the MI model. The modification is made to remove the effect of amino acid evolutionary pressure in measuring covariation. The second measure is the mutual information of residues physicochemical properties (MIP), which is used to measure the covariation of physicochemical properties of two sites. The third measure called MIBP is proposed by applying residues physicochemical properties into the MIB model. Moreover, scores of our new measures are applied to a robust indicator conn(k) in finding the covariation signal of each site. CONCLUSIONS: We find that incorporating amino acid background distribution is effective in removing the effect of evolutionary pressure of amino acids. Thus the MIB measure describes more biological background information for the coevolution of residues. Besides, our analysis also reveals that the covariation of physicochemical properties is a new aspect of coevolution information. Yongchao Dou, Jialiang Yang |
BMC Bioinform. | 3 |
| 2011 | Analysis on the reconstruction accuracy of the Fitch method for inferring ancestral statesabstractBACKGROUND: As one of the most widely used parsimony methods for ancestral reconstruction, the Fitch method minimizes the total number of hypothetical substitutions along all branches of a tree to explain the evolution of a character. Due to the extensive usage of this method, it has become a scientific endeavor in recent years to study the reconstruction accuracies of the Fitch method. However, most studies are restricted to 2-state evolutionary models and a study for higher-state models is needed since DNA sequences take the format of 4-state series and protein sequences even have 20 states. RESULTS: In this paper, the ambiguous and unambiguous reconstruction accuracy of the Fitch method are studied for N-state evolutionary models. Given an arbitrary phylogenetic tree, a recurrence system is first presented to calculate iteratively the two accuracies. As complete binary tree and comb-shaped tree are the two extremal evolutionary tree topologies according to balance, we focus on the reconstruction accuracies on these two topologies and analyze their asymptotic properties. Then, 1000 Yule trees with 1024 leaves are generated and analyzed to simulate real evolutionary scenarios. It is known that more taxa not necessarily increase the reconstruction accuracies under 2-state models. The result under N-state models is also tested. CONCLUSIONS: In a large tree with many leaves, the reconstruction accuracies of using all taxa are sometimes less than those of using a leaf subset under N-state models. For complete binary trees, there always exists an equilibrium interval [a, b] of conservation probability, in which the limiting ambiguous reconstruction accuracy equals to the probability of randomly picking a state. The value b decreases with the increase of the number of states, and it seems to converge. When the conservation probability is greater than b, the reconstruction accuracies of the Fitch method increase rapidly. The reconstruction accuracies on 1000 simulated Yule trees also exhibit similar behaviors. For comb-shaped trees, the limiting reconstruction accuracies of using all taxa are always less than or equal to those of using the nearest root-to-leaf path when the conservation probability is not less than 1/N. As a result, more taxa are suggested for ancestral reconstruction when the tree topology is balanced and the sequences are highly similar, and a few taxa close to the root are recommended otherwise. Jialiang Yang, Jun Li 0068, Liuhuan Dong, Stefan Grünewald |
BMC Bioinform. | 1 |
| 2008 | Run Probability of High-Order Seed Patterns and its Applications to Finding Good Transition Seeds
Jialiang Yang, Louxin Zhang |
APBC | 1 |