Jiani Ma

dblp:69/5199 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-8126-0431ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 6 first-author · 9 since 2021
YearPublicationVenuePosition
2026 scDBImpute: Dual-Branch Imputation for Single-Cell RNA-Seq Data Dropouts
abstract
Single-cell RNA sequencing (scRNA-seq) enables a comprehensive analysis of the expression patterns of individual cells in tissue with the resolution of a single cell. However, "dropouts" will lead to an excess of zeros in the scRNA-seq data due to technical constraints, which could hinder further analysis. Consequently, imputing the dropout values becomes particularly critical in assisting with the recovery of biological information. Herein, we propose a dual-branch imputation method for scRNA-seq data, which helps to impute the drops in scRNA-seq data. Unlike previous methods which assume a preconceived structure guiding to impute the dropouts, we believe that there are linear and non-linear associations that help structure the data, thus, both linear and non-linear pipelines are combined for dropout imputation. The evaluation and comprehensive downstream results on both simulated and real datasets show that our method outperforms the state-of-the-art methods for recovery of gene expression, cell clustering, differential expression analysis, and pseudo-time trajectory analysis tasks.
Lin Zhang 0015, Feng Wang 0064, Jiani Ma, Hui Liu 0024
IEEE Trans. Comput. Biol. Bioinform.3
2025 pMHChat, characterizing the interactions between major histocompatibility complex class II molecules and peptides with large language models and deep hypergraph learning
abstract
Characterizing the binding interactions between major histocompatibility complex (MHC) class II molecules and peptides is crucial for studying the immune system, offering potential applications for neoantigen design, vaccine development, and personalized immunotherapy. Motivated by this profound meaning, we developed a model that integrates large language models (LLMs) and deep hypergraph learning for predicting MHC class II-peptide binding reactivity, affinity, and residue contact profiling. pMHChat takes MHC pseudo-sequences and peptide sequences as inputs and processes them through four stages: LLMs fine-tune stage, feature encoding and map fusion stage, task-specific prediction stage, and downstream analysis stage. pMHChat distinguishes itself in capturing contextually relevant and high-order spatial interactions of the peptide-MHC (pMHC) complex. Specifically, in a five-fold cross-validation experiment, pMHChat achieves superior performance, with a mean area under the receiver operating characteristic curve of 0.8744 and an area under the precision-recall curve of 0.8390 in the binding reactivity task, as well as a mean Pearson correlation coefficient of 0.7311 in the binding affinity prediction task. Furthermore, pMHChat also demonstrates the best performance in both the leave-one-molecule-out setting and independent evaluation. Notably, pMHChat can provide residue contact profiling, showing its potential application in recognizing critical binding patterns of the pMHC complex. Our findings highlight pMHChat's capacity to advance both predictive accuracy and detailed insights into the MHC-peptide binding process. We anticipate that pMHChat will serve as a powerful tool for elucidating MHC-peptide interactions, with promising applications in immunological research and therapeutic development.
Jiani Ma, Zhikang Wang, Cen Tong, Lin Zhang 0015, Hui Liu 0024
Briefings Bioinform.1
2025 PathoRM: Computational inference of pathogenic RNA methylation sites by incorporating multi-view features
abstract
Identifying pathogenic RNA methylation sites with a reasonable biological explanation has important implications for the treatment of diseases. Due to the limitations of in vitro experiments in identifying pathogenic RNA methylation sites, there is a growing need for computational workflows to enable accurate inference. Here, motivated by this profound meaning, we developed PathoRM, a biologically informed deep learning model, to infer associations between RNA methylation sites and diseases. PathoRM could provide convincing pathogenic RNA methylation sites and unravel the enigma of pathology in the epi-transcriptomic layer. PathoRM fuses RNA methylation host sequences and pathogenic descriptions as inputs, and subsequently employs large language models, multi-view learning algorithm, graph neural networks, an adversarial training approach, and "guilty-by-association"-derived negative sampling approach. PathoRM distils the semantically enriched feature embeddings, leading to more accurate and robust prediction performance across the metrics and datasets. Notably, incorporated with attention mechanism, PathoRM bestows itself biological interpretability through illuminating the dark matters in the host sequences of RNA methylation sites. This work is expected to assist in the discovery of pathogenic RNA methylation sites and conserved motifs, contributing to the advancement of genome research. Codes and pre-trained model are accessible at https://github.com/jianiM/PathoRM.
Hui Liu 0024, Jiani Ma, Xianjun Ma, Lin Zhang 0015
PLoS Comput. Biol.2
2024 'Bingo' - a large language model- and graph neural network-based workflow for the prediction of essential genes from protein data
abstract
The identification and characterization of essential genes are central to our understanding of the core biological functions in eukaryotic organisms, and has important implications for the treatment of diseases caused by, for example, cancers and pathogens. Given the major constraints in testing the functions of genes of many organisms in the laboratory, due to the absence of in vitro cultures and/or gene perturbation assays for most metazoan species, there has been a need to develop in silico tools for the accurate prediction or inference of essential genes to underpin systems biological investigations. Major advances in machine learning approaches provide unprecedented opportunities to overcome these limitations and accelerate the discovery of essential genes on a genome-wide scale. Here, we developed and evaluated a large language model- and graph neural network (LLM-GNN)-based approach, called 'Bingo', to predict essential protein-coding genes in the metazoan model organisms Caenorhabditis elegans and Drosophila melanogaster as well as in Mus musculus and Homo sapiens (a HepG2 cell line) by integrating LLM and GNNs with adversarial training. Bingo predicts essential genes under two 'zero-shot' scenarios with transfer learning, showing promise to compensate for a lack of high-quality genomic and proteomic data for non-model organisms. In addition, the attention mechanisms and GNNExplainer were employed to manifest the functional sites and structural domain with most contribution to essentiality. In conclusion, Bingo provides the prospect of being able to accurately infer the essential genes of little- or under-studied organisms of interest, and provides a biological explanation for gene essentiality.
Jiani Ma, Jiangning Song, Neil D. Young, Bill C. H. Chang, Pasi K. Korhonen, Tulio L. Campos, Hui Liu 0024, Robin B. Gasser
Briefings Bioinform.1
2024 Dual-stream multi-dependency graph neural network enables precise cancer survival analysis
abstract
Histopathology image-based survival prediction aims to provide a precise assessment of cancer prognosis and can inform personalized treatment decision-making in order to improve patient outcomes. However, existing methods cannot automatically model the complex correlations between numerous morphologically diverse patches in each whole slide image (WSI), thereby preventing them from achieving a more profound understanding and inference of the patient status. To address this, here we propose a novel deep learning framework, termed dual-stream multi-dependency graph neural network (DM-GNN), to enable precise cancer patient survival analysis. Specifically, DM-GNN is structured with the feature updating and global analysis branches to better model each WSI as two graphs based on morphological affinity and global co-activating dependencies. As these two dependencies depict each WSI from distinct but complementary perspectives, the two designed branches of DM-GNN can jointly achieve the multi-view modeling of complex correlations between the patches. Moreover, DM-GNN is also capable of boosting the utilization of dependency information during graph construction by introducing the affinity-guided attention recalibration module as the readout function. This novel module offers increased robustness against feature perturbation, thereby ensuring more reliable and stable predictions. Extensive benchmarking experiments on five TCGA datasets demonstrate that DM-GNN outperforms other state-of-the-art methods and offers interpretable prediction insights based on the morphological depiction of high-attention patches. Overall, DM-GNN represents a powerful and auxiliary tool for personalized cancer prognosis from histopathology images and has great potential to assist clinicians in making personalized treatment decisions and improving patient outcomes.
Zhikang Wang, Jiani Ma, Chris Bain, Seiya Imoto, Pietro Liò, Hongmin Cai, Hao Chen 0011, Jiangning Song
Medical Image Anal.2
2023 MULGA, a unified multi-view graph autoencoder-based approach for identifying drug-protein interaction and drug repositioning
abstract
MOTIVATION: Identifying drug-protein interactions (DPIs) is a critical step in drug repositioning, which allows reuse of approved drugs that may be effective for treating a different disease and thereby alleviates the challenges of new drug development. Despite the fact that a great variety of computational approaches for DPI prediction have been proposed, key challenges, such as extendable and unbiased similarity calculation, heterogeneous information utilization, and reliable negative sample selection, remain to be addressed. RESULTS: To address these issues, we propose a novel, unified multi-view graph autoencoder framework, termed MULGA, for both DPI and drug repositioning predictions. MULGA is featured by: (i) a multi-view learning technique to effectively learn authentic drug affinity and target affinity matrices; (ii) a graph autoencoder to infer missing DPI interactions; and (iii) a new "guilty-by-association"-based negative sampling approach for selecting highly reliable non-DPIs. Benchmark experiments demonstrate that MULGA outperforms state-of-the-art methods in DPI prediction and the ablation studies verify the effectiveness of each proposed component. Importantly, we highlight the top drugs shortlisted by MULGA that target the spike glycoprotein of severe acute respiratory syndrome coronavirus 2 (SAR-CoV-2), offering additional insights into and potentially useful treatment option for COVID-19. Together with the availability of datasets and source codes, we envision that MULGA can be explored as a useful tool for DPI prediction and drug repositioning. AVAILABILITY AND IMPLEMENTATION: MULGA is publicly available for academic purposes at https://github.com/jianiM/MULGA/.
Jiani Ma, Chen Li 0021, Zhikang Wang, Shanshan Li 0008, Yuming Guo 0001, Lin Zhang 0015, Hui Liu 0024, Xin Gao 0001, Jiangning Song
Bioinform.1
2022 LRTCLS: low-rank tensor completion with Laplacian smoothing regularization for unveiling the post-transcriptional machinery of N6-methylation (m6A)-mediated diseases
abstract
Recently, N6-methylation (m6A) has recently become a hot topic due to its key role in disease pathogenesis. Identifying disease-related m6A sites aids in the understanding of the molecular mechanisms and biosynthetic pathways underlying m6A-mediated diseases. Existing methods treat it primarily as a binary classification issue, focusing solely on whether an m6A-disease association exists or not. Although they achieved good results, they all shared one common flaw: they ignored the post-transcriptional regulation events during disease pathogenesis, which makes biological interpretation unsatisfactory. Thus, accurate and explainable computational models are required to unveil the post-transcriptional regulation mechanisms of disease pathogenesis mediated by m6A modification, rather than simply inferring whether the m6A sites cause disease or not. Emerging laboratory experiments have revealed the interactions between m6A and other post-transcriptional regulation events, such as circular RNA (circRNA) targeting, microRNA (miRNA) targeting, RNA-binding protein binding and alternative splicing events, etc., present a diverse landscape during tumorigenesis. Based on these findings, we proposed a low-rank tensor completion-based method to infer disease-related m6A sites from a biological standpoint, which can further aid in specifying the post-transcriptional machinery of disease pathogenesis. It is so exciting that our biological analysis results show that Coronavirus disease 2019 may play a role in an m6A- and miRNA-dependent manner in inducing non-small cell lung cancer.
Jiani Ma, Hui Liu 0024, Yumeng Mao, Lin Zhang 0015
Briefings Bioinform.1
2022 BRPCA: Bounded Robust Principal Component Analysis to Incorporate Similarity Network for N7-Methylguanosine(m7G) Site-Disease Association Prediction
abstract
Recent studies have revealed that N7-methylguanosine(m7G) plays a pivotal role in various biological processes and disease pathogenesis. To date, transctriptome-wide m7G modification sites have been identified by high-throughput sequencing approaches, and some related information has been recorded in a few biological databases. However, the mechanism of site action in disease remains uncharted. Wet experiments can help identify true m7G sites with high confidence, but it is time-consuming to find the true ones in such a large number of sites, which will also cost too much. Thus, computational methods are emergently needed to predict the associations between m7G sites and various diseases, thus help to uncover potential active sites for specific diseases. In this article, we proposed a bounded robust principal component analysis (BRPCA) method to predict unknown m7G-disease association based on similarity information. Importantly, BRPCA tolerates the noise and redundancy existing in association and similarity information. Moreover, a suitable bounded constraint is incorporated into BRPCA to ensure that the predicted association scores locate in a meaningful interval. The extensive experiments demonstrate the rationality and superiority of the BRPCA.
Jiani Ma, Lin Zhang 0015, Hui Liu 0024
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 m7GDisAI: N7-methylguanosine (m7G) sites and diseases associations inference based on heterogeneous network
abstract
Abstract Background Recent studies have confirmed that N7-methylguanosine (m7G) modification plays an important role in regulating various biological processes and has associations with multiple diseases. Wet-lab experiments are cost and time ineffective for the identification of disease-associated m7G sites. To date, tens of thousands of m7G sites have been identified by high-throughput sequencing approaches and the information is publicly available in bioinformatics databases, which can be leveraged to predict potential disease-associated m7G sites using a computational perspective. Thus, computational methods for m7G-disease association prediction are urgently needed, but none are currently available at present. Results To fill this gap, we collected association information between m7G sites and diseases, genomic information of m7G sites, and phenotypic information of diseases from different databases to build an m7G-disease association dataset. To infer potential disease-associated m7G sites, we then proposed a heterogeneous network-based model, m7G Sites and Diseases Associations Inference (m7GDisAI) model. m7GDisAI predicts the potential disease-associated m7G sites by applying a matrix decomposition method on heterogeneous networks which integrate comprehensive similarity information of m7G sites and diseases. To evaluate the prediction performance, 10 runs of tenfold cross validation were first conducted, and m7GDisAI got the highest AUC of 0.740(± 0.0024). Then global and local leave-one-out cross validation (LOOCV) experiments were implemented to evaluate the model’s accuracy in global and local situations respectively. AUC of 0.769 was achieved in global LOOCV, while 0.635 in local LOOCV. A case study was finally conducted to identify the most promising ovarian cancer-related m7G sites for further functional analysis. Gene Ontology (GO) enrichment analysis was performed to explore the complex associations between host gene of m7G sites and GO terms. The results showed that m7GDisAI identified disease-associated m7G sites and their host genes are consistently related to the pathogenesis of ovarian cancer, which may provide some clues for pathogenesis of diseases. Conclusion The m7GDisAI web server can be accessed at http://180.208.58.66/m7GDisAI/ , which provides a user-friendly interface to query disease associated m7G. The list of top 20 m7G sites predicted to be associted with 177 diseases can be achieved. Furthermore, detailed information about specific m7G sites and diseases are also shown.
Jiani Ma, Lin Zhang 0015, Chenxuan Zang, Hui Liu 0024
BMC Bioinform.1