EDBT 2026 Demo / reviewers in the wild / expert
Lin Zhang 0015
dblp:37/1629-15
· DBLP profile ↗
34ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-7404-5231ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly semi-supervised cardiac MRI segmentation with frequency-domain pseudo label dynamic mixed supervision and partial-dice constraints
Wenzong Li, Hui Liu 0024, Jingcui Qin, Hongdang Zheng, Lin Zhang 0015 |
Neural Networks | 5 |
| 2026 | scDBImpute: Dual-Branch Imputation for Single-Cell RNA-Seq Data DropoutsabstractSingle-cell RNA sequencing (scRNA-seq) enables a comprehensive analysis of the expression patterns of individual cells in tissue with the resolution of a single cell. However, "dropouts" will lead to an excess of zeros in the scRNA-seq data due to technical constraints, which could hinder further analysis. Consequently, imputing the dropout values becomes particularly critical in assisting with the recovery of biological information. Herein, we propose a dual-branch imputation method for scRNA-seq data, which helps to impute the drops in scRNA-seq data. Unlike previous methods which assume a preconceived structure guiding to impute the dropouts, we believe that there are linear and non-linear associations that help structure the data, thus, both linear and non-linear pipelines are combined for dropout imputation. The evaluation and comprehensive downstream results on both simulated and real datasets show that our method outperforms the state-of-the-art methods for recovery of gene expression, cell clustering, differential expression analysis, and pseudo-time trajectory analysis tasks. Lin Zhang 0015, Feng Wang 0064, Jiani Ma, Hui Liu 0024 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2026 | HTCNet: Hierarchical Point-Graph Tooth Point Cloud Completion With Image AssistanceabstractWith the growing prominence of digital dentistry, high-quality and cost-effective three-dimensional (3D) intra-oral scanned (IOS) tooth data has become essential in various dental applications. However, limited by sensor resolution and occlusion, the acquired tooth point clouds often suffer from sparsity and incompleteness, resulting in missing regions and insufficient 3D detail. In this paper, we propose a Hierarchical Tooth Completion Network(HTCNet), a novel image-assistance framework to integrate geometry and structure learning from 2D images and 3D point clouds for 3D tooth completion. It employs a dual-stream-based hierarchical feature extraction, utilizing a 2D stream for extracting image features and a 3D stream for processing point clouds. Additionally, we introduce 3DTeethSegX for evaluating image-assistance tooth completion to address the lack of consideration for the comprehensive situation in the current dataset. Extensive experiments on two tooth completion datasets demonstrate the superiority and robustness of HTCNet, showcasing its potential to generate high-resolution 3D tooth data from low-resolution, cost-effective sensors with the assistance of 2D images. To the best of our knowledge, it is the first to use image-assistance techniques to significantly improve the accuracy and effectiveness of 3D tooth completion. Source code will be available at: https://github.com/labiip/HTCNet. Fucheng Niu, Hui Liu 0024, Mengqi Liang, Zekuan Yu, Lin Zhang 0015 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Boosting Contrastive Learning for Myopia Prediction: Which Data Augmentation Works Best?abstractMyopia is an increasingly prevalent global health issue, underscoring the need for accurate and early-stage prediction methods. Given the inherent scarcity of labeled medical data, contrastive learning offers a promising solution for myopia prediction. However, it still remains unexplored. In this paper, we focus on investigating the impact of data augmentation strategies on contrastive learning for fundus-based myopia prediction and propose CImPred, a contrastive learning-based framework for myopia prediction using retinal fundus images. To support this task, we construct a fundus image dataset, RFMPred. We systematically evaluate 36 combinations of common augmentation techniques and demonstrates that the combination of color distortion and random rotation yields the best performance. CImPred achieves an RMSE of 1.630, PCC of 0.672, and MAE of 1.321 on the RFMPred dataset. We also provide preliminary insights into fundus regions associated with early-stage myopia, which may support early intervention and risk reduction efforts. The code will be made publicly available upon publication. Fucheng Niu, Liao Ya, Hui Liu 0024, Lin Zhang 0015 |
BIBM | 5 |
| 2025 | pMHChat, characterizing the interactions between major histocompatibility complex class II molecules and peptides with large language models and deep hypergraph learningabstractCharacterizing the binding interactions between major histocompatibility complex (MHC) class II molecules and peptides is crucial for studying the immune system, offering potential applications for neoantigen design, vaccine development, and personalized immunotherapy. Motivated by this profound meaning, we developed a model that integrates large language models (LLMs) and deep hypergraph learning for predicting MHC class II-peptide binding reactivity, affinity, and residue contact profiling. pMHChat takes MHC pseudo-sequences and peptide sequences as inputs and processes them through four stages: LLMs fine-tune stage, feature encoding and map fusion stage, task-specific prediction stage, and downstream analysis stage. pMHChat distinguishes itself in capturing contextually relevant and high-order spatial interactions of the peptide-MHC (pMHC) complex. Specifically, in a five-fold cross-validation experiment, pMHChat achieves superior performance, with a mean area under the receiver operating characteristic curve of 0.8744 and an area under the precision-recall curve of 0.8390 in the binding reactivity task, as well as a mean Pearson correlation coefficient of 0.7311 in the binding affinity prediction task. Furthermore, pMHChat also demonstrates the best performance in both the leave-one-molecule-out setting and independent evaluation. Notably, pMHChat can provide residue contact profiling, showing its potential application in recognizing critical binding patterns of the pMHC complex. Our findings highlight pMHChat's capacity to advance both predictive accuracy and detailed insights into the MHC-peptide binding process. We anticipate that pMHChat will serve as a powerful tool for elucidating MHC-peptide interactions, with promising applications in immunological research and therapeutic development. Jiani Ma, Zhikang Wang, Cen Tong, Lin Zhang 0015, Hui Liu 0024 |
Briefings Bioinform. | 5 |
| 2025 | MCLCBA: multi-view contrastive learning network for RNA methylation site predictionabstractAbstract Background RNA methylation (RM) regulates gene expression regulation, RNA stability, and protein translation. Accurate prediction of RM modification sites is essential for understanding their biological functions. However, existing wet-lab detection techniques face challenges including operational complexity and high costs. Deep learning (DL) methods have been applied to this task. However, existing methods show performance degradation with smaller training datasets. For instance, the Bidirectional Gated Recurrent Unit (BGRU) demonstrates substantial performance degradation. Contrastive Learning Network (CNN) can extract local pattern features but learns overly specific patterns with sample-limited data, resulting in poor feature generalization. Bidirectional Long Short-Term Memory (BiLSTM) excels at modeling long-range dependencies but cannot sufficiently learn gating mechanism parameters to capture effective sequence representations with limited samples. Transformer processes sequences in parallel and captures global dependencies through self-attention, but its quadratic computational complexity and large parameter count make it prone to overfitting on small datasets. Current DL methods show reduced performance when training data is limited. Results This study proposes a Multi-view Contrastive Learning with CNN-BiLSTM-Attention (MCLCBA) framework for RM modification site prediction. The multi-view approach comprises a primary view and auxiliary view, where the primary view utilizes DNA Bidirectional Encoder Representations from Transformers (DNABERT) to extract sequence contextual features, and the auxiliary view employs Chaos Game Representation (CGR) to extract structural features. Feature extraction includes four components: data augmentation, multi-view encoders, projection heads, and contrastive loss functions. By implementing dual differential data augmentation strategies and constructing multi-view network architectures for feature processing and fusion, the model learns discriminative feature representations invariant to data augmentation through maximizing positive sample similarity while minimizing negative sample similarity. This effectively addresses sample-limited feature learning scenarios. Experimental results on the sample-limited m 7 G dataset demonstrate that MCLCBA achieves AUROC and AUPRC of 85.64% and 86.94%, respectively, improving upon existing methods by 5–6% in both metrics. Conclusions Through multi-view contrastive learning, MCLCBA provides an approach for RM sites under sample-limited scenarios. Yanjing Sun, Zhaoyang Liu 0002, Lin Zhang 0015 |
BMC Bioinform. | 5 |
| 2025 | PathoRM: Computational inference of pathogenic RNA methylation sites by incorporating multi-view featuresabstractIdentifying pathogenic RNA methylation sites with a reasonable biological explanation has important implications for the treatment of diseases. Due to the limitations of in vitro experiments in identifying pathogenic RNA methylation sites, there is a growing need for computational workflows to enable accurate inference. Here, motivated by this profound meaning, we developed PathoRM, a biologically informed deep learning model, to infer associations between RNA methylation sites and diseases. PathoRM could provide convincing pathogenic RNA methylation sites and unravel the enigma of pathology in the epi-transcriptomic layer. PathoRM fuses RNA methylation host sequences and pathogenic descriptions as inputs, and subsequently employs large language models, multi-view learning algorithm, graph neural networks, an adversarial training approach, and "guilty-by-association"-derived negative sampling approach. PathoRM distils the semantically enriched feature embeddings, leading to more accurate and robust prediction performance across the metrics and datasets. Notably, incorporated with attention mechanism, PathoRM bestows itself biological interpretability through illuminating the dark matters in the host sequences of RNA methylation sites. This work is expected to assist in the discovery of pathogenic RNA methylation sites and conserved motifs, contributing to the advancement of genome research. Codes and pre-trained model are accessible at https://github.com/jianiM/PathoRM. Hui Liu 0024, Jiani Ma, Xianjun Ma, Lin Zhang 0015 |
PLoS Comput. Biol. | 4 |
| 2025 | ESNet: End-to-End Chromosome Instance Segmentation Method Based on Edge Supervised NetworkabstractChromosomal numerical and structural abnormalities are common in newborn defects. Karyotype analysis, which can detect abnormalities based on chromosome microscope images, has emerged as one of the gold standards for the diagnosis of such diseases. Individual chromosomes are segmented, numbered and identified throughout this procedure. Thus, precise chromosomal segmentation, as the key of karyotyping, affects the success of subsequent missions. Most segmentation methods cannot identify a chromosome in case of overlapping since there is little distinction between inter and intra chromosomal classes. The loss and redundancy of the segmentation mask may be to blame for that. An end-to-end framework named Edge Supervised Network (ESNet) based on Mask RCNN is proposed in this paper. To incorporate edge prior knowledge, an edge supervised branch and feature fusion module are designed to help recognize individual chromosomes from clusters, and a spatial attention module to capture more contextual information for better edge identification. Balance loss weight is also proposed for edge loss in the training phase. Experimental results reveal that ESNet achieves better segmentation performance in comparison with other competing methods, which can be a potential baseline network for end-to-end chromosome instance segmentation. Hui Liu 0024, Xiuyu Li, Zhaolin Lu, Lin Zhang 0015 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | FDDSeg: Unleashing the Power of Scribble Annotation for Cardiac MRI Images Through Feature Decomposition DistillationabstractCardiovascular diseases can be diagnosed with computer assistance when using the magnetic resonance imaging (MRI) image that is produced by the MRI sensor. Deep learning-based scribbling MRI image segmentation has demonstrated impressive results recently. However, the majority of current approaches possess an excessive number of model parameters and do not completely utilize scribbling annotations. To develop a feature decomposition distillation deep learning method, named FDDSeg, for scribble-supervised cardiac MRI image segmentation. Public ACDC and MSCMR cardiac MRI datasets were used to evaluate the segmentation performance of FDDSeg. FDDSeg adopts a scribble annotation reuse policy to help provide accurate boundaries, and the intermediate features are split class region and class-free region by using the pseudo labels to further improve feature learning. Effective distillation knowledge is then captured by feature decomposition. FDDSeg was compared with 7 state-of-the-art methods, MAAG, ShapePU, CycleMix, Dual-Branch, ZscribbleSeg, Perturbation Dual-Branch as well as ScribbleVC on both ACDC and MSCMR datasets. FDDSeg is shown to perform the best in DSC(89.05% and 88.75%), JC(80.30% and 79.78%) as well as HD95(5.76% and 4.44%) metrics with only 2.01 M of parameters. FDDSeg methods can segment cardiac MRI images more precise with only scribble annotations at lower computation cost, which may help increase the efficiency of quantitative analysis of cardiac. Lin Zhang 0015, Wenzong Li, Kaiyue Bi, Hui Liu 0024 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | MSCAN: multi-scale self- and cross-attention network for RNA methylation site predictionabstractAbstract Background Epi-transcriptome regulation through post-transcriptional RNA modifications is essential for all RNA types. Precise recognition of RNA modifications is critical for understanding their functions and regulatory mechanisms. However, wet experimental methods are often costly and time-consuming, limiting their wide range of applications. Therefore, recent research has focused on developing computational methods, particularly deep learning (DL). Bidirectional long short-term memory (BiLSTM), convolutional neural network (CNN), and the transformer have demonstrated achievements in modification site prediction. However, BiLSTM cannot achieve parallel computation, leading to a long training time, CNN cannot learn the dependencies of the long distance of the sequence, and the Transformer lacks information interaction with sequences at different scales. This insight underscores the necessity for continued research and development in natural language processing (NLP) and DL to devise an enhanced prediction framework that can effectively address the challenges presented. Results This study presents a multi-scale self- and cross-attention network (MSCAN) to identify the RNA methylation site using an NLP and DL way. Experiment results on twelve RNA modification sites (m6A, m1A, m5C, m5U, m6Am, m7G, Ψ, I, Am, Cm, Gm, and Um) reveal that the area under the receiver operating characteristic of MSCAN obtains respectively 98.34%, 85.41%, 97.29%, 96.74%, 99.04%, 79.94%, 76.22%, 65.69%, 92.92%, 92.03%, 95.77%, 89.66%, which is better than the state-of-the-art prediction model. This indicates that the model has strong generalization capabilities. Furthermore, MSCAN reveals a strong association among different types of RNA modifications from an experimental perspective. A user-friendly web server for predicting twelve widely occurring human RNA modification sites (m6A, m1A, m5C, m5U, m6Am, m7G, Ψ, I, Am, Cm, Gm, and Um) is available at http://47.242.23.141/MSCAN/index.php . Conclusions A predictor framework has been developed through binary classification to predict RNA methylation sites. Wenliang Zeng, Yanjing Sun, Lin Zhang 0015 |
BMC Bioinform. | 6 |
| 2024 | DBDNMF: A Dual Branch Deep Neural Matrix Factorization method for drug response predictionabstractAnti-cancer response of cell lines to drugs is in urgent need for individualized precision medical decision-making in the era of precision medicine. Measurements with wet-experiments is time-consuming and expensive and it is almost impossible for wide ranges of application. The design of computational models that can precisely predict the responses between drugs and cell lines could provide a credible reference for further research. Existing methods of response prediction based on matrix factorization or neural networks have revealed that both linear or nonlinear latent characteristics are applicable and effective for the precise prediction of drug responses. However, the majority of them consider only linear or nonlinear relationships for drug response prediction. Herein, we propose a Dual Branch Deep Neural Matrix Factorization (DBDNMF) method to address the above-mentioned issues. DBDNMF learns the latent representation of drugs and cell lines through flexible inputs and reconstructs the partially observed matrix through a series of hidden neural network layers. Experimental results on the datasets of Cancer Cell Line Encyclopedia (CCLE) and Genomics of Drug Sensitivity in Cancer (GDSC) show that the accuracy of drug prediction exceeds state-of-the-art drug response prediction algorithms, demonstrating its reliability and stability. The hierarchical clustering results show that drugs with similar response levels tend to target similar signaling pathway, and cell lines coming from the same tissue subtype tend to share the same pattern of response, which are consistent with previously published studies. Hui Liu 0024, Feng Wang 0064, Chaoju Gong, Lin Zhang 0015 |
PLoS Comput. Biol. | 7 |
| 2023 | MULGA, a unified multi-view graph autoencoder-based approach for identifying drug-protein interaction and drug repositioningabstractMOTIVATION: Identifying drug-protein interactions (DPIs) is a critical step in drug repositioning, which allows reuse of approved drugs that may be effective for treating a different disease and thereby alleviates the challenges of new drug development. Despite the fact that a great variety of computational approaches for DPI prediction have been proposed, key challenges, such as extendable and unbiased similarity calculation, heterogeneous information utilization, and reliable negative sample selection, remain to be addressed. RESULTS: To address these issues, we propose a novel, unified multi-view graph autoencoder framework, termed MULGA, for both DPI and drug repositioning predictions. MULGA is featured by: (i) a multi-view learning technique to effectively learn authentic drug affinity and target affinity matrices; (ii) a graph autoencoder to infer missing DPI interactions; and (iii) a new "guilty-by-association"-based negative sampling approach for selecting highly reliable non-DPIs. Benchmark experiments demonstrate that MULGA outperforms state-of-the-art methods in DPI prediction and the ablation studies verify the effectiveness of each proposed component. Importantly, we highlight the top drugs shortlisted by MULGA that target the spike glycoprotein of severe acute respiratory syndrome coronavirus 2 (SAR-CoV-2), offering additional insights into and potentially useful treatment option for COVID-19. Together with the availability of datasets and source codes, we envision that MULGA can be explored as a useful tool for DPI prediction and drug repositioning. AVAILABILITY AND IMPLEMENTATION: MULGA is publicly available for academic purposes at https://github.com/jianiM/MULGA/. Jiani Ma, Chen Li 0021, Zhikang Wang, Shanshan Li 0008, Yuming Guo 0001, Lin Zhang 0015, Hui Liu 0024, Xin Gao 0001, Jiangning Song |
Bioinform. | 7 |
| 2023 | FGFICA: Independent Component Analysis of Fusion Genomic Features for Mining Epi-Transcriptome Profiling DataabstractA. Shutao Chen, Lin Zhang 0015, Xiangzhi Chen 0003, Hui Liu 0024 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | LRTCLS: low-rank tensor completion with Laplacian smoothing regularization for unveiling the post-transcriptional machinery of N6-methylation (m6A)-mediated diseasesabstractRecently, N6-methylation (m6A) has recently become a hot topic due to its key role in disease pathogenesis. Identifying disease-related m6A sites aids in the understanding of the molecular mechanisms and biosynthetic pathways underlying m6A-mediated diseases. Existing methods treat it primarily as a binary classification issue, focusing solely on whether an m6A-disease association exists or not. Although they achieved good results, they all shared one common flaw: they ignored the post-transcriptional regulation events during disease pathogenesis, which makes biological interpretation unsatisfactory. Thus, accurate and explainable computational models are required to unveil the post-transcriptional regulation mechanisms of disease pathogenesis mediated by m6A modification, rather than simply inferring whether the m6A sites cause disease or not. Emerging laboratory experiments have revealed the interactions between m6A and other post-transcriptional regulation events, such as circular RNA (circRNA) targeting, microRNA (miRNA) targeting, RNA-binding protein binding and alternative splicing events, etc., present a diverse landscape during tumorigenesis. Based on these findings, we proposed a low-rank tensor completion-based method to infer disease-related m6A sites from a biological standpoint, which can further aid in specifying the post-transcriptional machinery of disease pathogenesis. It is so exciting that our biological analysis results show that Coronavirus disease 2019 may play a role in an m6A- and miRNA-dependent manner in inducing non-small cell lung cancer. Jiani Ma, Hui Liu 0024, Yumeng Mao, Lin Zhang 0015 |
Briefings Bioinform. | 4 |
| 2022 | EMDLP: Ensemble multiscale deep learning model for RNA methylation site predictionabstractAbstract Background Recent research recommends that epi-transcriptome regulation through post-transcriptional RNA modifications is essential for all sorts of RNA. Exact identification of RNA modification is vital for understanding their purposes and regulatory mechanisms. However, traditional experimental methods of identifying RNA modification sites are relatively complicated, time-consuming, and laborious. Machine learning approaches have been applied in the procedures of RNA sequence features extraction and classification in a computational way, which may supplement experimental approaches more efficiently. Recently, convolutional neural network (CNN) and long short-term memory (LSTM) have been demonstrated achievements in modification site prediction on account of their powerful functions in representation learning. However, CNN can learn the local response from the spatial data but cannot learn sequential correlations. And LSTM is specialized for sequential modeling and can access both the contextual representation but lacks spatial data extraction compared with CNN. There is strong motivation to construct a prediction framework using natural language processing (NLP), deep learning (DL) for these reasons. Results This study presents an ensemble multiscale deep learning predictor (EMDLP) to identify RNA methylation sites in an NLP and DL way. It organically combines the dilated convolution and Bidirectional LSTM (BiLSTM), which helps to take better advantage of the local and global information for site prediction. The first step of EMDLP is to represent the RNA sequences in an NLP way. Thus, three encodings, e.g., RNA word embedding, One-hot encoding, and RGloVe, which is an improved learning method of word vector representation based on GloVe, are adopted to decipher sites from the viewpoints of the local and global information. Then, a dilated convolutional Bidirectional LSTM network (DCB) model is constructed with the dilated convolutional neural network (DCNN) followed by BiLSTM to extract potential contributing features for methylation site prediction. Finally, these three encoding methods are integrated by a soft vote to obtain better predictive performance. Experiment results on m1A and m6A reveal that the area under the receiver operating characteristic(AUROC) of EMDLP obtains respectively 95.56%, 85.24%, and outperforms the state-of-the-art models. To maximize user convenience, a user-friendly webserver for EMDLP was publicly available at http://www.labiip.net/EMDLP/index.php ( http://47.104.130.81/EMDLP/index.php ). Conclusions We developed a predictor for m1A and m6A methylation sites. Hui Liu 0024, Gangshen Li, Lin Zhang 0015, Yanjing Sun |
BMC Bioinform. | 5 |
| 2022 | FBCwPlaid: A Functional Biclustering Analysis of Epi-Transcriptome Profiling Data Via a Weighted Plaid ModelabstractRecent studies have shown that in-depth studies on epi-transcriptomic patterns of N6-methyladenosine (m6A) may help understand its complex functions and co-regulatory mechanisms. Since most biclustering algorithms are developed in scenarios of gene expression analysis, which does not share the same characteristics with m6A methylation profile, we propose a weighted Plaid biclustering model (FBCwPlaid) based on the Lagrange multiplier method to discover the potential functional patterns. Each pattern is achieved by minimizing approximation error between FBCwPlaid predicted value and real data. To address the issue that site expression level determines methylation level confidence, it uses RNA expression levels of each site as weights to make lower expressed sites less confident. FBCwPlaid also allows overlapping biclusters, indicating some sites may participate in multiple biological functions. FBCwPlaid was then applied on MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions, each of which represented a stimulus to a particular cell line or environment. Finally, three patterns were discovered, and further pathway analysis and enzyme specificity test showed that sites involved in each pattern are highly relevant to m6A methyltransferases. Further detailed analyses showed that some patterns are condition-specific, indicating that some specific sites’ methylation profiles may occur in specific cell lines or conditions. Shutao Chen, Lin Zhang 0015, Jia Meng 0001, Hui Liu 0024 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | BRPCA: Bounded Robust Principal Component Analysis to Incorporate Similarity Network for N7-Methylguanosine(m7G) Site-Disease Association PredictionabstractRecent studies have revealed that N7-methylguanosine(m7G) plays a pivotal role in various biological processes and disease pathogenesis. To date, transctriptome-wide m7G modification sites have been identified by high-throughput sequencing approaches, and some related information has been recorded in a few biological databases. However, the mechanism of site action in disease remains uncharted. Wet experiments can help identify true m7G sites with high confidence, but it is time-consuming to find the true ones in such a large number of sites, which will also cost too much. Thus, computational methods are emergently needed to predict the associations between m7G sites and various diseases, thus help to uncover potential active sites for specific diseases. In this article, we proposed a bounded robust principal component analysis (BRPCA) method to predict unknown m7G-disease association based on similarity information. Importantly, BRPCA tolerates the noise and redundancy existing in association and similarity information. Moreover, a suitable bounded constraint is incorporated into BRPCA to ensure that the predicted association scores locate in a meaningful interval. The extensive experiments demonstrate the rationality and superiority of the BRPCA. Jiani Ma, Lin Zhang 0015, Hui Liu 0024 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | BDBB: A Novel Beta-Distribution-Based Biclustering Algorithm for Revealing Local Co-Methylation Patterns in Epi-Transcriptome Profiling DataabstractN6-methyladenosine (m6A) has been shown to play crucial roles in RNA metabolism, physiology, and pathological processes. However, the specific regulatory mechanisms of most methylation sites remain uncharted due to the complexity of life processes. Biological experimental methods are costly to solve this problem, and computational methods are relatively lacking. The discovery of local co-methylation patterns (LCPs) of m6A epi-transcriptome data can benefit to solve the above problems. Based on this, we propose a novel biclustering algorithm based on the beta distribution (BDBB), which realizes the mining of LCPs of m6A epi-transcriptome data. BDBB employs the Gibbs sampling method to complete parameter estimation. In the process of modeling, LCPs are recognized as sharp beta distributions compared to the background distribution. Simulation study showed BDBB can extract all the three actual LCPs implanted in the background data and the overlap conditions between them with considerable accuracy (almost close to 100%). On MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions from 10 human cell lines, BDBB unveiled two LCPs, and Gene Ontology (GO) enrichment analysis showed that they were enriched in histone modification and embryo development, etc. important biological processes respectively. The GOE_Score scoring indicated that the biclustering results of BDBB in the m6A epi-transcriptome data are more biologically meaningful than the results of other biclustering algorithms. Zhaoyang Liu 0002, Yuteng Xiao, Hongsheng Yin 0001, Shutao Chen, Kaijian Xia, Lin Zhang 0015 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Lysosome activation in peripheral blood mononuclear cells and prognostic significance of circulating LC3B in COVID-19abstractCoronavirus disease 2019 (COVID-19) has spread rapidly worldwide, causing significant mortality. There is a mechanistic relationship between intracellular coronavirus replication and deregulated autophagosome-lysosome system. We performed transcriptome analysis of peripheral blood mononuclear cells (PBMCs) from COVID-19 patients and identified the aberrant upregulation of genes in the lysosome pathway. We further determined the capability of two circulating markers, namely microtubule-associated proteins 1A/1B light chain 3B (LC3B) and (p62/SQSTM1) p62, both of which depend on lysosome for degradation, in predicting the emergence of moderate-to-severe disease in COVID-19 patients requiring hospitalization for supplemental oxygen therapy. Logistic regression analyses showed that LC3B was associated with moderate-to-severe COVID-19, independent of age, sex and clinical risk score. A decrease in LC3B concentration <5.5 ng/ml increased the risk of oxygen and ventilatory requirement (adjusted odds ratio: 4.6; 95% CI: 1.1-22.0; P = 0.04). Serum concentrations of p62 in the moderate-to-severe group were significantly lower in patients aged 50 or below. In conclusion, lysosome function is deregulated in PBMCs isolated from COVID-19 patients, and the related biomarker LC3B may serve as a novel tool for stratifying patients with moderate-to-severe COVID-19 from those with asymptomatic or mild disease. COVID-19 patients with a decrease in LC3B concentration <5.5 ng/ml will require early hospital admission for supplemental oxygen therapy and other respiratory support. Shisong Fang, Lin Zhang 0015, Yingzhi Liu, Wenye Xu, Weihua Wu, Ziheng Huang 0005, Hui Liu 0024, Renli Zhang, Jun Yu 0009, Francis Ka-Leung Chan, Siew Chien Ng, Sunny Hei Wong, Maggie Haitian Wang, Tony Gin, Gavin Matthew Joynt, David Shu Cheong Hui, Tiejian Feng, William Ka Kei Wu, Matthew Tak Vai Chan, Xuan Zou, Junjie Xia |
Briefings Bioinform. | 2 |
| 2021 | Multi-omic analysis suggests tumor suppressor genes evolved specific promoter features to optimize cancer resistanceabstractTumor suppressor genes (TSGs) exhibit distinct evolutionary features. We speculated that TSG promoters could have evolved specific features that facilitate their tumor-suppressing functions. We found that the promoter CpG dinucleotide frequencies of TSGs are significantly higher than that of non-cancer genes across vertebrate genomes, and positively correlated with gene expression across tissue types. The promoter CpG dinucleotide frequencies of all genes gradually increase with gene age, for which young TSGs have been subject to a stronger evolutionary pressure. Transcription-related features, namely chromatin accessibility, methylation and ZNF263-, SP1-, E2F4- and SP2-binding elements, are associated with gene expression. Moreover, higher promoter CpG dinucleotide frequencies and chromatin accessibility are positively associated with the ability of TSGs to resist downregulation during tumorigenesis. These results were successfully validated with independent datasets. In conclusion, TSGs evolved specific promoter features that optimized cancer resistance through achieving high expression in normal tissues and resistance to downregulation during tumorigenesis. Xiansong Wang, Yingzhi Liu, Ziheng Huang 0005, Xiaoxu Hu, Hung Chan, Yidan Zou, Idy H. T. Ho, Alfred S. L. Cheng, Ka F. To, Maggie Haitian Wang, Sunny Hei Wong, Jun Yu 0009, Tony Gin, Qingpeng Zhang, Jianxiong Shen, Lin Zhang 0015, Matthew Tak Vai Chan, William Ka Kei Wu |
Briefings Bioinform. | 22 |
| 2021 | Bioinformatic analysis of SMN1-ACE/ACE2 interactions hinted at a potential protective effect of spinal muscular atrophy against COVID-19-induced lung injuryabstractPatients with spinal muscular atrophy (SMA) are susceptible to the respiratory infections and might be at a heightened risk of poor clinical outcomes upon contracting coronavirus disease 2019 (COVID-19). In the face of the COVID-19 pandemic, the potential associations of SMA with the susceptibility to and prognostication of COVID-19 need to be clarified. We documented an SMA case who contracted COVID-19 but only developed mild-to-moderate clinical and radiological manifestations of pneumonia, which were relieved by a combined antiviral and supportive treatment. We then reviewed a cohort of patients with SMA who had been living in the Hubei province since November 2019, among which the only 1 out of 56 was diagnosed with COVID-19 (1.79%, 1/56). Bioinformatic analysis was carried out to delineate the potential genetic crosstalk between SMN1 (mutation of which leads to SMA) and COVID-19/lung injury-associated pathways. Protein-protein interaction analysis by STRING suggested that loss-of-function of SMN1 might modulate COVID-19 pathogenesis through CFTR, CXCL8, TNF and ACE. Expression quantitative trait loci analysis also revealed a link between SMN1 and ACE2, despite low-confidence protein-protein interactions as suggested by STRING. This bioinformatic analysis could give hint on why SMA might not necessarily lead to poor outcomes in patients with COVID-19. Xingye Li, Jianxiong Shen, Haining Tan, Tianhua Rong, Youxi Lin, Erwei Feng, Zhengguang Chen, Lin Zhang 0015, Matthew Tak Vai Chan, William Ka Kei Wu |
Briefings Bioinform. | 11 |
| 2021 | m7GDisAI: N7-methylguanosine (m7G) sites and diseases associations inference based on heterogeneous networkabstractAbstract Background Recent studies have confirmed that N7-methylguanosine (m7G) modification plays an important role in regulating various biological processes and has associations with multiple diseases. Wet-lab experiments are cost and time ineffective for the identification of disease-associated m7G sites. To date, tens of thousands of m7G sites have been identified by high-throughput sequencing approaches and the information is publicly available in bioinformatics databases, which can be leveraged to predict potential disease-associated m7G sites using a computational perspective. Thus, computational methods for m7G-disease association prediction are urgently needed, but none are currently available at present. Results To fill this gap, we collected association information between m7G sites and diseases, genomic information of m7G sites, and phenotypic information of diseases from different databases to build an m7G-disease association dataset. To infer potential disease-associated m7G sites, we then proposed a heterogeneous network-based model, m7G Sites and Diseases Associations Inference (m7GDisAI) model. m7GDisAI predicts the potential disease-associated m7G sites by applying a matrix decomposition method on heterogeneous networks which integrate comprehensive similarity information of m7G sites and diseases. To evaluate the prediction performance, 10 runs of tenfold cross validation were first conducted, and m7GDisAI got the highest AUC of 0.740(± 0.0024). Then global and local leave-one-out cross validation (LOOCV) experiments were implemented to evaluate the model’s accuracy in global and local situations respectively. AUC of 0.769 was achieved in global LOOCV, while 0.635 in local LOOCV. A case study was finally conducted to identify the most promising ovarian cancer-related m7G sites for further functional analysis. Gene Ontology (GO) enrichment analysis was performed to explore the complex associations between host gene of m7G sites and GO terms. The results showed that m7GDisAI identified disease-associated m7G sites and their host genes are consistently related to the pathogenesis of ovarian cancer, which may provide some clues for pathogenesis of diseases. Conclusion The m7GDisAI web server can be accessed at http://180.208.58.66/m7GDisAI/ , which provides a user-friendly interface to query disease associated m7G. The list of top 20 m7G sites predicted to be associted with 177 diseases can be achieved. Furthermore, detailed information about specific m7G sites and diseases are also shown. Jiani Ma, Lin Zhang 0015, Chenxuan Zang, Hui Liu 0024 |
BMC Bioinform. | 2 |
| 2021 | EDLm6APred: ensemble deep learning approach for mRNA m6A site predictionabstractAbstract Background As a common and abundant RNA methylation modification, N6-methyladenosine (m6A) is widely spread in various species' transcriptomes, and it is closely related to the occurrence and development of various life processes and diseases. Thus, accurate identification of m6A methylation sites has become a hot topic. Most biological methods rely on high-throughput sequencing technology, which places great demands on the sequencing library preparation and data analysis. Thus, various machine learning methods have been proposed to extract various types of features based on sequences, then occupied conventional classifiers, such as SVM, RF, etc., for m6A methylation site identification. However, the identification performance relies heavily on the extracted features, which still need to be improved. Results This paper mainly studies feature extraction and classification of m6A methylation sites in a natural language processing way, which manages to organically integrate the feature extraction and classification simultaneously, with consideration of upstream and downstream information of m6A sites. One-hot, RNA word embedding, and Word2vec are adopted to depict sites from the perspectives of the base as well as its upstream and downstream sequence. The BiLSTM model, a well-known sequence model, was then constructed to discriminate the sequences with potential m6A sites. Since the above-mentioned three feature extraction methods focus on different perspectives of m6A sites, an ensemble deep learning predictor (EDLm6APred) was finally constructed for m6A site prediction. Experimental results on human and mouse data sets show that EDLm6APred outperforms the other single ones, indicating that base, upstream, and downstream information are all essential for m6A site detection. Compared with the existing m6A methylation site prediction models without genomic features, EDLm6APred obtains 86.6% of the area under receiver operating curve on the human data sets, indicating the effectiveness of sequential modeling on RNA. To maximize user convenience, a webserver was developed as an implementation of EDLm6APred and made publicly available at www.xjtlu.edu.cn/biologicalsciences/EDLm6APred . Conclusions Our proposed EDLm6APred method is a reliable predictor for m6A methylation sites. Lin Zhang 0015, Gangshen Li, Xiuyu Li, Shutao Chen, Hui Liu 0024 |
BMC Bioinform. | 1 |
| 2020 | REW-ISA: unveiling local functional blocks in epi-transcriptome profiling data via an RNA expression-weighted iterative signature algorithmabstractAbstract Background Recent studies have shown that N6-methyladenosine (m6A) plays a critical role in numbers of biological processes and complex human diseases. However, the regulatory mechanisms of most methylation sites remain uncharted. Thus, in-depth study of the epi-transcriptomic patterns of m6A may provide insights into its complex functional and regulatory mechanisms. Results Due to the high economic and time cost of wet experimental methods, revealing methylation patterns through computational models has become a more preferable way, and drawn more and more attention. Considering the theoretical basics and applications of conventional clustering methods, an RNA Expression Weighted Iterative Signature Algorithm (REW-ISA) is proposed to find potential local functional blocks (LFBs) based on MeRIP-Seq data, where sites are hyper-methylated or hypo-methylated simultaneously across the specific conditions. REW-ISA adopts RNA expression levels of each site as weights to make sites of lower expression level less significant. It starts from random sets of sites, then follows iterative search strategies by thresholds of rows and columns to find the LFBs in m6A methylation profile. Its application on MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions unveiled 6 LFBs, which achieve higher enrichment scores than ISA. Pathway analysis and enzyme specificity test showed that sites remained in LFBs are highly relevant to the m6A methyltransferase, such as METTL3, METTL14, WTAP and KIAA1429. Further detailed analyses for each LFB even showed that some LFBs are condition-specific, indicating that methylation profiles of some specific sites may be condition relevant. Conclusions REW-ISA finds potential local functional patterns presented in m6A profiles, where sites are co-methylated under specific conditions. Lin Zhang 0015, Shutao Chen, Jia Meng 0001, Hui Liu 0024 |
BMC Bioinform. | 1 |
| 2019 | RNA methylation and diseases: experimental results, databases, Web servers and computational modelsabstractRibonucleic acid (RNA) methylation is a type of posttranscriptional modifications occurring in all kingdoms of life. It is strongly related to important biological process, thus making it linked to a number of human diseases. Owing to the development of high-throughput sequencing technology, plenty of achievement had been obtained in RNA methylation research recently. Meanwhile, various computational models have been developed to analyze and mining increasing RNA methylation data. In this review, we first made a brief introduction about eight types of most popular RNA methylation, the biological functions of RNA methylation, the relationship between RNA methylation and disease and five important RNA methylation-related diseases. The research of RNA methylation is based on sequencing data processing, and effective bioinformatics techniques can benefit better understanding of RNA methylation. We further introduced seven publicly available RNA methylation-related databases, and some important publicly available RNA-methylation-related Web servers and software for RNA methylation site identification, differential analysis and so on. Furthermore, we provided detailed analysis of the state-of-the-art computational models used in these Web servers and software. We also analyzed the limitations of these models and discussed the future directions of developing computational models for RNA methylation research. Xing Chen 0001, Ya-Zhou Sun, Hui Liu 0024, Lin Zhang 0015, Jianqiang Li 0001, Jia Meng 0001 |
Briefings Bioinform. | 4 |
| 2019 | m6Acomet: large-scale functional prediction of individual m6A RNA methylation sites from an RNA co-methylation networkabstractOver one hundred different types of post-transcriptional RNA modifications have been identified in human. Researchers discovered that RNA modifications can regulate various biological processes, and RNA methylation, especially N6-methyladenosine, has become one of the most researched topics in epigenetics. To date, the study of epitranscriptome layer gene regulation is mostly focused on the function of mediator proteins of RNA methylation, i.e., the readers, writers and erasers. There is limited investigation of the functional relevance of individual m6A RNA methylation site. To address this, we annotated human m6A sites in large-scale based on the guilt-by-association principle from an RNA co-methylation network. It is constructed based on public human MeRIP-Seq datasets profiling the m6A epitranscriptome under 32 independent experimental conditions. By systematically examining the network characteristics obtained from the RNA methylation profiles, a total of 339,158 putative gene ontology functions associated with 1446 human m6A sites were identified. These are biological functions that may be regulated at epitranscriptome layer via reversible m6A RNA methylation. The results were further validated on a soft benchmark by comparing to a random predictor. An online web server m6Acomet was constructed to support direct query for the predicted biological functions of m6A sites as well as the sites exhibiting co-methylated patterns at the epitranscriptome layer. The m6Acomet web server is freely available at: www.xjtlu.edu.cn/biologicalsciences/m6acomet . Kunqi Chen, Jionglong Su, Hui Liu 0024, Lin Zhang 0015, Jia Meng 0001 |
BMC Bioinform. | 7 |
| 2018 | trumpet: transcriptome-guided quality assessment of m6A-seq dataabstractMethylated RNA immunoprecipitation sequencing (MeRIP-seq or m 6 A-seq) has been extensively used for profiling transcriptome-wide distribution of RNA N6-Methyl-Adnosine methylation. However, due to the intrinsic properties of RNA molecules and the intricate procedures of this technique, m 6 A-seq data often suffer from various flaws. A convenient and comprehensive tool is needed to assess the quality of m 6 A-seq data to ensure that they are suitable for subsequent analysis. From a technical perspective, m 6 A-seq can be considered as a combination of ChIP-seq and RNA-seq; hence, by effectively combing the data quality assessment metrics of the two techniques, we developed the trumpet R package for evaluation of m 6 A-seq data quality. The trumpet package takes the aligned BAM files from m 6 A-seq data together with the transcriptome information as the inputs to generate a quality assessment report in the HTML format. The trumpet R package makes a valuable tool for assessing the data quality of m 6 A-seq, and it is also applicable to other fragmented RNA immunoprecipitation sequencing techniques, including m 1 A-seq, CeU-Seq, Ψ-seq, etc. Shaowu Zhang 0001, Lin Zhang 0015, Jia Meng 0001 |
BMC Bioinform. | 3 |
| 2018 | MeTDiff: A Novel Differential RNA Methylation Analysis for MeRIP-Seq DataabstractN6-Methyladenosine (m6A) transcriptome methylation is an exciting new research area that just captures the attention of research community. We present in this paper, MeTDiff, a novel computational tool for predicting differential m6A methylation sites from Methylated RNA immunoprecipitation sequencing (MeRIP-Seq) data. Compared with the existing algorithm exomePeak, the advantages of MeTDiff are that it explicitly models the reads variation in data and also devices a more power likelihood ratio test for differential methylation site prediction. Comprehensive evaluation of MeTDiff's performance using both simulated and real datasets showed that MeTDiff is much more robust and achieved much higher sensitivity and specificity over exomePeak. Lin Zhang 0015, Jia Meng 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Cancer Progression Prediction Using Gene Interaction Regularized Elastic NetabstractDifferent types of genomic aberration may simultaneously contribute to tumorigenesis. To obtain a more accurate prognostic assessment to guide therapeutic regimen choice for cancer patients, the heterogeneous multi-omics data should be integrated harmoniously, which can often be difficult. For this purpose, we propose a Gene Interaction Regularized Elastic Net (GIREN) model that predicts clinical outcome by integrating multiple data types. GIREN conveniently embraces both gene measurements and gene-gene interaction information under an elastic net formulation, enforcing structure sparsity, and the "grouping effect" in solution to select the discriminate features with prognostic value. An iterative gradient descent algorithm is also developed to solve the model with regularized optimization. GIREN was applied to human ovarian cancer and breast cancer datasets obtained from The Cancer Genome Atlas, respectively. Result shows that, the proposed GIREN algorithm obtained more accurate and robust performance over competing algorithms (LASSO, Elastic Net, and Semi-supervised PCA, with or without average pathway expression features) in predicting cancer progression on both two datasets in terms of median area under curve (AUC) and interquartile range (IQR), suggesting a promising direction for more effective integration of gene measurement and gene interaction information. Lin Zhang 0015, Hui Liu 0024, Yufei Huang 0001, Xuesong Wang 0001, Yidong Chen 0002, Jia Meng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | Sketching the distribution of transcriptomic features on RNA transcripts with Travis coordinatesabstractBiological features, such as, genes, transcription factor binding sites, SNPs, etc., are usually denoted with genome-based coordinates as the genomic features. While genome-based representation is usually very effective, it can be tedious to examine the distribution of RNA-related genomic features on RNA transcripts with existing tools due to the conversion and comparison between genome-based coordinates to RNA-based coordinates. We developed here an open source R package Travis for sketching the transcriptomic view of genomic features so as to facilitate the analysis of RNA-related but genome-based coordinates. Internally, Travis package extracts the coordinates relative to the landmarks of transcripts, with which the distribution of RNA-related genomic features can then be conveniently analyzed. We demonstrated the usage of Travis package in analyzing post-transcriptional RNA modifications (5-MethylCytosine and N6-MethylAdenosine) derived from high-throughput sequencing approaches (MeRIP-Seq and RNA BS-Seq). The Travis R package is now publicly available from GitHub: https://github.com/lzcyzm/Travis. Lin Zhang 0015, Hui Liu 0024, Shaowu Zhang 0001, Yufei Huang 0001, Jia Meng 0001 |
BIBM | 3 |
| 2013 | Unveiling the dynamics in RNA epigenetic regulationsabstractDespite the prevalent studies of DNA/Chromatin related epigenetics, such as, histone modifications and DNA methylation, RNA epigenetics did not receive deserved attention due to the lack of high throughput approach for profiling epitranscriptome. Recently, a new affinity-based sequencing approach MeRIPseq was developed and applied to survey the global mRNA N6-methyladenosine (m6A) in mammalian cells. As a marriage of ChIPseq and RNAseq, MeRIPseq has the potential to study, for the first time, the transcriptome-wide distribution of different types of post-transcriptional RNA modifications. Yet, this technology introduced new computational challenges that have not been adequately addressed. We have previously developed a MATLAB-based package ‘exomePeak’ for detection of RNA methylation sites from MeRIPseq data. Here, we extend the features of exomePeak by including a novel computational framework that enables differential analysis to unveil the dynamics in RNA epigenetic regulations. The novel differential analysis monitors the percentage of modified RNA molecules among the total transcribed RNAs, which directly reflects the impact of RNA epigenetic regulations. In contrast, current available software packages developed for sequencing-based differential analysis such as DESeq or edgeR monitors the changes in the absolute amount of molecules, and, if applied to MeRIPseq data, might be dominated by transcriptional gene differential expression. The algorithm is implemented as an R-package ‘exomePeak’ and freely available. It takes directly the aligned BAM files as input, statistically supports biological replicates, corrects PCR artifacts, and outputs exome-based results in BED format, which is compatible with all major genome browsers for convenient visualization and manipulation. Examples are also provided to depict how exomePeak R-package is integrated with exiting tools for MeRIPseq based peak calling and differential analysis. Particularly, the rationales behind each processing step as well as the specific method used, the best practice, and possible alternative strategies are briefly discussed. The algorithm was applied to the human HepG2 cell MeRIPseq data sets and detects more than 16000 RNA m6A sites, many of which are differentially methylated under ultraviolet radiation. The challenges and potentials of MeRIPseq in epitranscriptome studies are discussed in the end. Jia Meng 0001, Hui Liu 0024, Lin Zhang 0015, Shaowu Zhang 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001 |
BIBM | 4 |
| 2013 | Integration of gene expression, genome wide DNA methylation, and gene networks for clinical outcome prediction in ovarian cancerabstractIntegrative clinical outcome prediction model called gene interaction regularized elastic net (GIREN) method is proposed in this paper. GIREN combines gene expression, methylation profiles, and gene interaction networks in order to reveal genomic and epigenomic features that bear important prognostic value. With GIREN, gene expression and DNA methylation profiles are first jointly analyzed in a linear regression model, and additional gene interaction network is simultaneously integrated as a regularizing penalty that follow an elastic net formulation. Such regularization also enforce sparsity in the solution so that features with prognostic values are automatically selected. To solve the regularized optimization, an iterative gradient descent algorithm is also developed. We applied GIREN to a set of 87 human ovarian cancer samples, which underwent a rigorous sample selection. The predicted outcome was used to group patients into high-risk vs. low-risk. Validation showed that GIREN outperformed other competing algorithms including SuperPCA. Lin Zhang 0015, Hui Liu 0024, Jia Meng 0001, Xuesong Wang 0001, Yidong Chen 0002, Yufei Huang 0001 |
BIBM | 1 |
| 2010 | SysMicrO: A Novel Systems Approach for miRNA Target Prediction
Hui Liu 0024, Lin Zhang 0015, Qilong Sun, Yidong Chen 0002, Yufei Huang 0001 |
ICIC (2) | 2 |
| 2010 | miRNA Target Prediction Method Based on the Combination of Multiple Algorithms
Lin Zhang 0015, Hui Liu 0024, Dong Yue 0002, Yufei Huang 0001 |
ICIC (1) | 1 |