EDBT 2026 Demo / reviewers in the wild / expert
Zhixiang Lin
dblp:215/7767
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0001-9301-947XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoBep: a region-oriented deep learning model for B-cell epitope predictionabstractMOTIVATION: Accurate in silico identification of B-cell epitope residues is crucial for antibody design and structure-guided vaccine development. Although recent protein language models and structure-aware methods can capture spatial information of tertiary structure when generating residue embeddings, most existing epitope predictors use these embeddings to perform classification for individual residues one by one, without enforcing spatial continuity for reported epitope residues. Such methods often result in biologically implausible predictions because B-cell epitope residues always cluster together on the antigen surface. RESULTS: We present RoBep, a region-oriented B-cell epitope predictor that explicitly models the spatial clustering of epitope residues. RoBep introduces a novel region constraint mechanism and combines the advanced protein language model ESM-Cambrian with an equivariant graph neural network. Our method outperforms existing structure-based methods on the benchmark dataset, demonstrating improvements of 26%, 45%, 13%, and 43% in F1, Matthews correlation coefficient, area under the precision-recall curve, and AUROC0.1, respectively. In addition to residue-level predictions, RoBep can also provide antibody-antigen binding regions. Importantly, the predicted epitope residues are ensured to be spatially compact, enhancing biological plausibility and practical relevance for immunotherapeutic design. AVAILABILITY AND IMPLEMENTATION: A user-friendly website for using RoBep is provided at https://huggingface.co/spaces/NielTT/RoBep. All datasets, source code used in this work, and implementation instructions of the website are publicly available at https://github.com/YitaoXU/RoBep. Guanyun Wei, Jingying Zhou, Yuanhua Huang, Weichuan Yu, Zhixiang Lin, Xiaodan Fan |
Bioinform. | 6 |
| 2025 | STIFT: spatiotemporal transcriptomics integration through spatially informed multi-timepoint bridgingabstractRecent advances in spatial transcriptomics have highlighted the need for integrating spatiotemporal transcriptomics data, defined as spatially resolved gene expression profiles captured across sequential time points in developmental or regenerative processes. We present STIFT (SpatioTemporal Integration Framework for Transcriptomics) specifically designed for integrating spatiotemporal transcriptomics data. STIFT is a three-component framework combining developmental spatiotemporal optimal transport, spatiotemporal graph construction, and a graph attention autoencoder informed by temporal triplet learning. STIFT integrates large-scale 2D or 3D spatiotemporal transcriptomics data, enabling batch effect removal, spatial domain identification, trajectory inference and exploration of developmental dynamics. Applied to axolotl brain regeneration, mouse embryonic development, and 3D planarian regeneration datasets, STIFT removes batch effects and achieves clear spatial domain identification while preserving temporal developmental patterns and biological variations across hundreds of thousands of spots, demonstrating its effectiveness and specificity in integrating spatiotemporal transcriptomics data. Muyang Ge, Jishuai Miao, Xiaocheng Zhou, Zhixiang Lin |
Briefings Bioinform. | 5 |
| 2025 | InterVelo: a mutually enhancing model for estimating pseudotime and RNA velocity in multi-omic single-cell dataabstractMOTIVATION: RNA velocity has become a powerful tool for uncovering transcriptional dynamics in snapshot single-cell data. However, current RNA velocity approaches often assume constant transcriptional rates and treat genes independently with gene-specific times, which may introduce biases and deviate from biological realities. Here, we present InterVelo, a novel deep learning framework that simultaneously learns cellular pseudotime and RNA velocity. RESULTS: InterVelo leverages an unsupervised cellular time to guide RNA velocity estimation, while the estimated RNA velocity in turn refines the direction of pseudotime. By benchmarking InterVelo against existing methods on both simulated and real datasets, we demonstrate its superior performance in recovering pseudotime and RNA velocity. InterVelo yields more precise velocity estimations in terms of both direction and magnitude, with outstanding robustness across diverse scenarios. Furthermore, it successfully identifies driver genes and enables reliable gene activity enrichment analysis. The flexible architecture of InterVelo also allows for the integration of multi-omic data, enhancing its applicability to complex biological systems. AVAILABILITY AND IMPLEMENTATION: InterVelo is implemented using Python, and the code is available on GitHub https://github.com/yurouwang-rosie/InterVelo and has been archived with a DOI https://doi.org/10.5281/zenodo.16158798 for reproducibility. Yurou Wang, Zhixiang Lin, Tao Wang 0067 |
Bioinform. | 2 |
| 2024 | SGCAST: symmetric graph convolutional auto-encoder for scalable and accurate study of spatial transcriptomicsabstractRecent advances in spatial transcriptomics (ST) have enabled comprehensive profiling of gene expression with spatial information in the context of the tissue microenvironment. However, with the improvements in the resolution and scale of ST data, deciphering spatial domains precisely while ensuring efficiency and scalability is still challenging. Here, we develop SGCAST, an efficient auto-encoder framework to identify spatial domains. SGCAST adopts a symmetric graph convolutional auto-encoder to learn aggregated latent embeddings via integrating the gene expression similarity and the proximity of the spatial spots. This framework in SGCAST enables a mini-batch training strategy, which makes SGCAST memory-efficient and scalable to high-resolution spatial transcriptomic data with a large number of spots. SGCAST improves the overall accuracy of spatial domain identification on benchmarking data. We also validated the performance of SGCAST on ST datasets at various scales across multiple platforms. Our study illustrates the superior capacity of SGCAST on analyzing spatial transcriptomic data. Jinzhao Li, Zhixiang Lin |
Briefings Bioinform. | 3 |
| 2024 | scICML: Information-Theoretic Co-Clustering-Based Multi-View Learning for the Integrative Analysis of Single-Cell Multi-Omics DataabstractModern high-throughput sequencing technologies have enabled us to profile multiple molecular modalities from the same single cell, providing unprecedented opportunities to assay cellular heterogeneity from multiple biological layers. However, the datasets generated from these technologies tend to have high level of noise and are highly sparse, bringing challenges to data analysis. In this paper, we develop a novel information-theoretic co-clustering-based multi-view learning (scICML) method for multi-omics single-cell data integration. scICML utilizes co-clusterings to aggregate similar features for each view of data and uncover the common clustering pattern for cells. In addition, scICML automatically matches the clusters of the linked features across different data types for considering the biological dependency structure across different types of genomic features. Our experiments on four real-world datasets demonstrate that scICML improves the overall clustering performance and provides biological insights into the data analysis of peripheral blood mononuclear cells. Pengcheng Zeng, Zhixiang Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | RefTM: reference-guided topic modeling of single-cell chromatin accessibility dataabstractSingle-cell analysis is a valuable approach for dissecting the cellular heterogeneity, and single-cell chromatin accessibility sequencing (scCAS) can profile the epigenetic landscapes for thousands of individual cells. It is challenging to analyze scCAS data, because of its high dimensionality and a higher degree of sparsity compared with scRNA-seq data. Topic modeling in single-cell data analysis can lead to robust identification of the cell types and it can provide insight into the regulatory mechanisms. Reference-guided approach may facilitate the analysis of scCAS data by utilizing the information in existing datasets. We present RefTM (Reference-guided Topic Modeling of single-cell chromatin accessibility data), which not only utilizes the information in existing bulk chromatin accessibility and annotated scCAS data, but also takes advantage of topic models for single-cell data analysis. RefTM simultaneously models: (1) the shared biological variation among reference data and the target scCAS data; (2) the unique biological variation in scCAS data; (3) other variations from known covariates in scCAS data. Shengquan Chen, Zhixiang Lin |
Briefings Bioinform. | 3 |
| 2023 | stVAE deconvolves cell-type composition in large-scale cellular resolution spatial transcriptomicsabstractMOTIVATION: Recent rapid developments in spatial transcriptomic techniques at cellular resolution have gained increasing attention. However, the unique characteristics of large-scale cellular resolution spatial transcriptomic datasets, such as the limited number of transcripts captured per spot and the vast number of spots, pose significant challenges to current cell-type deconvolution methods. RESULTS: In this study, we introduce stVAE, a method based on the variational autoencoder framework to deconvolve the cell-type composition of cellular resolution spatial transcriptomic datasets. To assess the performance of stVAE, we apply it to five datasets across three different biological tissues. In the Stereo-seq and Slide-seqV2 datasets of the mouse brain, stVAE accurately reconstructs the laminar structure of the pyramidal cell layers in the cortex, which are mainly organized by the subtypes of telencephalon projecting excitatory neurons. In the Stereo-seq dataset of the E12.5 mouse embryo, stVAE resolves the complex spatial patterns of osteoblast subtypes, which are supported by their marker genes. In Stereo-seq and Pixel-seq datasets of the mouse olfactory bulb, stVAE accurately delineates the spatial distributions of known cell types. In summary, stVAE can accurately identify spatial patterns of cell types and their relative proportions across spots for cellular resolution spatial transcriptomic data. It is instrumental in understanding the heterogeneity of cell populations and their interactions within tissues. AVAILABILITY AND IMPLEMENTATION: stVAE is available in GitHub (https://github.com/lichen2018/stVAE) and Figshare (https://figshare.com/articles/software/stVAE/23254538). Ting-Fung Chan, Can Yang 0002, Zhixiang Lin |
Bioinform. | 4 |
| 2023 | scAWMV: an adaptively weighted multi-view learning framework for the integrative analysis of parallel scRNA-seq and scATAC-seq dataabstractMOTIVATION: Technological advances have enabled us to profile single-cell multi-omics data from the same cells, providing us with an unprecedented opportunity to understand the cellular phenotype and links to its genotype. The available protocols and multi-omics datasets [including parallel single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) data profiled from the same cell] are growing increasingly. However, such data are highly sparse and tend to have high level of noise, making data analysis challenging. The methods that integrate the multi-omics data can potentially improve the capacity of revealing the cellular heterogeneity. RESULTS: We propose an adaptively weighted multi-view learning (scAWMV) method for the integrative analysis of parallel scRNA-seq and scATAC-seq data profiled from the same cell. scAWMV considers both the difference in importance across different modalities in multi-omics data and the biological connection of the features in the scRNA-seq and scATAC-seq data. It generates biologically meaningful low-dimensional representations for the transcriptomic and epigenomic profiles via unsupervised learning. Application to four real datasets demonstrates that our framework scAWMV is an efficient method to dissect cellular heterogeneity for single-cell multi-omics data. AVAILABILITY AND IMPLEMENTATION: The software and datasets are available at https://github.com/pengchengzeng/scAWMV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pengcheng Zeng, Zhixiang Lin |
Bioinform. | 3 |
| 2022 | JSNMF enables effective and accurate integrative analysis of single-cell multiomics dataabstractThe single-cell multiomics technologies provide an unprecedented opportunity to study the cellular heterogeneity from different layers of transcriptional regulation. However, the datasets generated from these technologies tend to have high levels of noise, making data analysis challenging. Here, we propose jointly semi-orthogonal nonnegative matrix factorization (JSNMF), which is a versatile toolkit for the integrative analysis of transcriptomic and epigenomic data profiled from the same cell. JSNMF enables data visualization and clustering of the cells and also facilitates downstream analysis, including the characterization of markers and functional pathway enrichment analysis. The core of JSNMF is an unsupervised method based on JSNMF, where it assumes different latent variables for the two molecular modalities, and integrates the information of transcriptomic and epigenomic data with consensus graph fusion, which better tackles the distinct characteristics and levels of noise across different molecular modalities in single-cell multiomics data. We applied JSNMF to single-cell multiomics datasets from different tissues and different technologies. The results demonstrate the superior performance of JSNMF in clustering and data visualization of the cells. JSNMF also allows joint analysis of multiple single-cell multiomics experiments and single-cell multiomics data with more than two modalities profiled on the same cell. JSNMF also provides rich biological insight on the markers, cell-type-specific region-gene associations and the functions of the identified cell subpopulation. Zexuan Sun, Pengcheng Zeng, Zhixiang Lin |
Briefings Bioinform. | 5 |
| 2022 | FIRM: Flexible integration of single-cell RNA-sequencing data for large-scale multi-tissue cell atlas datasetsabstractSingle-cell RNA-sequencing (scRNA-seq) is being used extensively to measure the mRNA expression of individual cells from deconstructed tissues, organs and even entire organisms to generate cell atlas references, leading to discoveries of novel cell types and deeper insight into biological trajectories. These massive datasets are usually collected from many samples using different scRNA-seq technology platforms, including the popular SMART-Seq2 (SS2) and 10X platforms. Inherent heterogeneities between platforms, tissues and other batch effects make scRNA-seq data difficult to compare and integrate, especially in large-scale cell atlas efforts; yet, accurate integration is essential for gaining deeper insights into cell biology. We present FIRM, a re-scaling algorithm which accounts for the effects of cell type compositions, and achieve accurate integration of scRNA-seq datasets across multiple tissue types, platforms and experimental batches. Compared with existing state-of-the-art integration methods, FIRM provides accurate mixing of shared cell type identities and superior preservation of original structure without overcorrection, generating robust integrated datasets for downstream exploration and analysis. FIRM is also a facile way to transfer cell type labels and annotations from one dataset to another, making it a reliable and versatile tool for scRNA-seq analysis, especially for cell atlas data integration. Jingsi Ming, Zhixiang Lin, Can Yang 0002, Angela Ruohao Wu |
Briefings Bioinform. | 2 |
| 2021 | scAMACE: model-based approach to the joint analysis of single-cell data on chromatin accessibility, gene expression and methylationabstractMOTIVATION: The advancement in technologies and the growth of available single-cell datasets motivate integrative analysis of multiple single-cell genomic datasets. Integrative analysis of multimodal single-cell datasets combines complementary information offered by single-omic datasets and can offer deeper insights on complex biological process. Clustering methods that identify the unknown cell types are among the first few steps in the analysis of single-cell datasets, and they are important for downstream analysis built upon the identified cell types. RESULTS: We propose scAMACE for the integrative analysis and clustering of single-cell data on chromatin accessibility, gene expression and methylation. We demonstrate that cell types are better identified and characterized through analyzing the three data types jointly. We develop an efficient Expectation-Maximization algorithm to perform statistical inference, and evaluate our methods on both simulation study and real data applications. We also provide the GPU implementation of scAMACE, making it scalable to large datasets. AVAILABILITY AND IMPLEMENTATION: The software and datasets are available at https://github.com/cuhklinlab/scAMACE_py (python implementation) and https://github.com/cuhklinlab/scAMACE (R implementation). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiaxuan Wangwu, Zexuan Sun, Zhixiang Lin |
Bioinform. | 3 |
| 2021 | coupleCoC+: An information-theoretic co-clustering-based transfer learning framework for the integrative analysis of single-cell genomic dataabstractTechnological advances have enabled us to profile multiple molecular layers at unprecedented single-cell resolution and the available datasets from multiple samples or domains are growing. These datasets, including scRNA-seq data, scATAC-seq data and sc-methylation data, usually have different powers in identifying the unknown cell types through clustering. So, methods that integrate multiple datasets can potentially lead to a better clustering performance. Here we propose coupleCoC+ for the integrative analysis of single-cell genomic data. coupleCoC+ is a transfer learning method based on the information-theoretic co-clustering framework. In coupleCoC+, we utilize the information in one dataset, the source data, to facilitate the analysis of another dataset, the target data. coupleCoC+ uses the linked features in the two datasets for effective knowledge transfer, and it also uses the information of the features in the target data that are unlinked with the source data. In addition, coupleCoC+ matches similar cell types across the source data and the target data. By applying coupleCoC+ to the integrative clustering of mouse cortex scATAC-seq data and scRNA-seq data, mouse and human scRNA-seq data, mouse cortex sc-methylation and scRNA-seq data, and human blood dendritic cells scRNA-seq data from two batches, we demonstrate that coupleCoC+ improves the overall clustering performance and matches the cell subpopulations across multimodal single-cell genomic datasets. coupleCoC+ has fast convergence and it is computationally efficient. The software is available at https://github.com/cuhklinlab/coupleCoC_plus. Pengcheng Zeng, Zhixiang Lin |
PLoS Comput. Biol. | 2 |