Lei Wei 0009

dblp:99/4916-9 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0002-1546-6458ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 14 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cross-dataset annotation harmonization for cell-type hierarchy construction
abstract
MOTIVATION: Single-cell transcriptomic datasets annotate cell types with diverse schemes and varying resolution. This poses challenges in building unified hierarchical cell-type structures and hinders integration of large-scale datasets. To address this, several computational methods have been developed to harmonize cell type annotations across datasets and build data-driven hierarchies of cell types. RESULTS: Here, we benchmarked three state-of-the-art methods: scHPL, treeArches, and CellHint. We evaluated these methods across five simulated scenarios and five real-world scenarios across cell types and organs. To assess harmonization results, we designed three metrics, Annotation Harmonization F1-score (AH-F1), Tree Edit Distance Similarity and Parent-Children Branches Similarity, comparing the constructed cell-type hierarchies and the knowledge-based ones. Based on the benchmarking results, we found that methods performed well in simulated scenarios but still have room for improvement in complex real-world data. Thus, we developed OTHarmonizer, a tool based on partial optimal transport (OT) for cell-type harmonization and hierarchy construction. OTHarmonizer excels in accurately capturing equivalent and hierarchical relationships between cell types, offering a more effective approach for the cell-type hierarchy construction across datasets. AVAILABILITY AND IMPLEMENTATION: The simulated and real-world datasets in the benchmark are available on https://figshare.com/articles/dataset/OTHarmonizer/28243205. The source codes for the benchmark and OTHarmonizer are available online on GitHub at https://github.com/Duck-Boss/OTHarmonizer.
Tianhong Zhou, Yingtao Zhu, Jinmeng Jia, Xuegong Zhang, Lei Wei 0009
Bioinform.6
2025 DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
abstract
Vision–language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly reason like human clinicians or merely imitate superficial patterns. To address this gap, we propose DrVD-Bench, the first multimodal benchmark for clinical visual reasoning. DrVD-Bench consists of three modules: Visual Evidence Comprehension, Reasoning Trajectory Assessment, and Report Generation Evaluation, comprising a total of 7,789 image–question pairs. Our benchmark covers 20 task types, 17 diagnostic categories, and five imaging modalities—CT, MRI, ultrasound, radiography, and pathology. DrVD-Bench is explicitly structured to reflect the clinical reasoning workflow from modality recognition to lesion identification and diagnosis. We benchmark 19 VLMs, including general-purpose and medical-specific, open-source and proprietary models, and observe that performance drops sharply as reasoning complexity increases. While some models begin to exhibit traces of human-like reasoning, they often still rely on shortcut correlations rather than grounded visual understanding. DrVD-Bench offers a rigorous and structured evaluation framework to guide the development of clinically trustworthy VLMs.
Tianhong Zhou, Yingtao Zhu, Chuxi Xiao, Haiyang Bian, Lei Wei 0009, Xuegong Zhang
NeurIPS6
2025 Computational methods and data resources for predicting tumor neoantigens
abstract
Neoantigens are tumor-specific antigens presented exclusively by cancer cells. These antigens are recognized as nonself by the host immune system, thereby eliciting an antitumor T-cell response. This response is significantly enhanced through neoantigen-based immunotherapies, such as personalized cancer vaccines. The repertoire of neoantigens is unique to each cancer patient, necessitating neoantigen prediction for designing patient-specific immunotherapies. This review presents the computational methods and data resources used for neoantigen prediction, as well as the prediction-associated challenges. Neoantigen prediction typically uses human leukocyte antigen typing, RNA-seq transcript quantification, somatic variant calling, peptide-major histocompatibility complex (pMHC) presentation prediction, and pMHC recognition prediction as the main computational steps. The immunoinformatics tools used for these steps and for the overall prediction of neoantigens are systematically summarized and detailed in this review.
Xiaofei Zhao 0005, Lei Wei 0009, Xuegong Zhang
Briefings Bioinform.2
2025 uHAF: a unified hierarchical annotation framework for cell type standardization and harmonization
abstract
SUMMARY: In single-cell transcriptomics, inconsistent cell type annotations due to varied naming conventions and hierarchical granularity impede data integration, machine learning applications, and meaningful evaluations. To address this challenge, we developed the unified Hierarchical Annotation Framework (uHAF), which includes organ-specific hierarchical cell type trees (uHAF-T) and a mapping tool (uHAF-Agent) based on large language models. uHAF-T provides standardized hierarchical references for 38 organs, allowing for consistent label unification and analysis at different levels of granularity. uHAF-Agent leverages GPT-4 to accurately map diverse and informal cell type labels onto uHAF-T nodes, streamlining the harmonization process. By simplifying label unification, uHAF enhances data integration, supports machine learning applications, and enables biologically meaningful evaluations of annotation methods. Our framework serves as an essential resource for standardizing cell type annotations and fostering collaborative refinement in the single-cell research community. AVAILABILITY AND IMPLEMENTATION: uHAF is publicly available at: https://uhaf.unifiedcellatlas.org and https://github.com/SuperBianC/uhaf.
Haiyang Bian, Yinxin Chen, Lei Wei 0009, Xuegong Zhang
Bioinform.3
2024 scMulan: A Multitask Generative Pre-Trained Language Model for Single-Cell Analysis
Haiyang Bian, Xiaomin Dong, Chen Li 0001, Minsheng Hao, Jinyi Hu, Maosong Sun 0001, Lei Wei 0009, Xuegong Zhang
RECOMB9
2024 Discovering and Overcoming the Bias in Neoantigen Identification by Unified Machine Learning Models
Ziting Zhang, Wenxu Wu, Lei Wei 0009, Xiaowo Wang
RECOMB3
2024 scDecouple: decoupling cellular response from infected proportion bias in scCRISPR-seq
abstract
Single-cell clustered regularly interspaced short palindromic repeats-sequencing (scCRISPR-seq) is an emerging high-throughput CRISPR screening technology where the true cellular response to perturbation is coupled with infected proportion bias of guide RNAs (gRNAs) across different cell clusters. The mixing of these effects introduces noise into scCRISPR-seq data analysis and thus obstacles to relevant studies. We developed scDecouple to decouple true cellular response of perturbation from the influence of infected proportion bias. scDecouple first models the distribution of gene expression profiles in perturbed cells and then iteratively finds the maximum likelihood of cell cluster proportions as well as the cellular response for each gRNA. We demonstrated its performance in a series of simulation experiments. By applying scDecouple to real scCRISPR-seq data, we found that scDecouple enhances the identification of biologically perturbation-related genes. scDecouple can benefit scCRISPR-seq data analysis, especially in the case of heterogeneous samples or complex gRNA libraries.
Qiuchen Meng, Lei Wei 0009, Joshua W. K. Ho, Yinqing Li, Xuegong Zhang
Briefings Bioinform.2
2024 Benchmarking multi-omics integration algorithms across single-cell RNA and ATAC data
abstract
Recent advancements in single-cell sequencing technologies have generated extensive omics data in various modalities and revolutionized cell research, especially in the single-cell RNA and ATAC data. The joint analysis across scRNA-seq data and scATAC-seq data has paved the way to comprehending the cellular heterogeneity and complex cellular regulatory networks. Multi-omics integration is gaining attention as an important step in joint analysis, and the number of computational tools in this field is growing rapidly. In this paper, we benchmarked 12 multi-omics integration methods on three integration tasks via qualitative visualization and quantitative metrics, considering six main aspects that matter in multi-omics data analysis. Overall, we found that different methods have their own advantages on different aspects, while some methods outperformed other methods in most aspects. We therefore provided guidelines for selecting appropriate methods for specific scenarios and tasks to help obtain meaningful insights from multi-omics data integration.
Chuxi Xiao, Qiuchen Meng, Lei Wei 0009, Xuegong Zhang
Briefings Bioinform.4
2024 scDiffusion: conditional generation of high-quality single-cell data using diffusion model
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) data are important for studying the laws of life at single-cell level. However, it is still challenging to obtain enough high-quality scRNA-seq data. To mitigate the limited availability of data, generative models have been proposed to computationally generate synthetic scRNA-seq data. Nevertheless, the data generated with current models are not very realistic yet, especially when we need to generate data with controlled conditions. In the meantime, diffusion models have shown their power in generating data with high fidelity, providing a new opportunity for scRNA-seq generation. RESULTS: In this study, we developed scDiffusion, a generative model combining the diffusion model and foundation model to generate high-quality scRNA-seq data with controlled conditions. We designed multiple classifiers to guide the diffusion process simultaneously, enabling scDiffusion to generate data under multiple condition combinations. We also proposed a new control strategy called Gradient Interpolation. This strategy allows the model to generate continuous trajectories of cell development from a given cell state. Experiments showed that scDiffusion could generate single-cell gene expression data closely resembling real scRNA-seq data. Also, scDiffusion can conditionally produce data on specific cell types including rare cell types. Furthermore, we could use the multiple-condition generation of scDiffusion to generate cell type that was out of the training data. Leveraging the Gradient Interpolation strategy, we generated a continuous developmental trajectory of mouse embryonic cells. These experiments demonstrate that scDiffusion is a powerful tool for augmenting the real scRNA-seq data and can provide insights into cell fate research. AVAILABILITY AND IMPLEMENTATION: scDiffusion is openly available at the GitHub repository https://github.com/EperLuo/scDiffusion or Zenodo https://zenodo.org/doi/10.5281/zenodo.13268742.
Erpai Luo, Minsheng Hao, Lei Wei 0009, Xuegong Zhang
Bioinform.3
2023 Decoding functional cell-cell communication events by multi-view graph learning on spatial transcriptomics
abstract
Cell-cell communication events (CEs) are mediated by multiple ligand-receptor (LR) pairs. Usually only a particular subset of CEs directly works for a specific downstream response in a particular microenvironment. We name them as functional communication events (FCEs) of the target responses. Decoding FCE-target gene relations is: important for understanding the mechanisms of many biological processes, but has been intractable due to the mixing of multiple factors and the lack of direct observations. We developed a method HoloNet for decoding FCEs using spatial transcriptomic data by integrating LR pairs, cell-type spatial distribution and downstream gene expression into a deep learning model. We modeled CEs as a multi-view network, developed an attention-based graph learning method to train the model for generating target gene expression with the CE networks, and decoded the FCEs for specific downstream genes by interpreting trained models. We applied HoloNet on three Visium datasets of breast cancer and liver cancer. The results detangled the multiple factors of FCEs by revealing how LR signals and cell types affect specific biological processes, and specified FCE-induced effects in each single cell. We conducted simulation experiments and showed that HoloNet is more reliable on LR prioritization in comparison with existing methods. HoloNet is a powerful tool to illustrate cell-cell communication landscapes and reveal vital FCEs that shape cellular phenotypes. HoloNet is available as a Python package at https://github.com/lhc17/HoloNet.
Haochen Li 0003, Tianxing Ma, Minsheng Hao, Wenbo Guo 0010, Jin Gu, Xuegong Zhang, Lei Wei 0009
Briefings Bioinform.7
2022 Evaluating methylation of human ribosomal DNA at each CpG site reveals its utility for cancer detection using cell-free DNA
abstract
Ribosomal deoxyribonucleic acid (DNA) (rDNA) repeats are tandemly located on five acrocentric chromosomes with up to hundreds of copies in the human genome. DNA methylation, the most well-studied epigenetic mechanism, has been characterized for most genomic regions across various biological contexts. However, rDNA methylation patterns remain largely unexplored due to the repetitive structure. In this study, we designed a specific mapping strategy to investigate rDNA methylation patterns at each CpG site across various physiological and pathological processes. We found that CpG sites on rDNA could be categorized into two types. One is within or adjacent to transcribed regions; the other is distal to transcribed regions. The former shows highly variable methylation levels across samples, while the latter shows stable high methylation levels in normal tissues but severe hypomethylation in tumors. We further showed that rDNA methylation profiles in plasma cell-free DNA could be used as a biomarker for cancer detection. It shows good performances on public datasets, including colorectal cancer [area under the curve (AUC) = 0.85], lung cancer (AUC = 0.84), hepatocellular carcinoma (AUC = 0.91) and in-house generated hepatocellular carcinoma dataset (AUC = 0.96) even at low genome coverage (<1×). Taken together, these findings broaden our understanding of rDNA regulation and suggest the potential utility of rDNA methylation features as disease biomarkers.
Xianglin Zhang, Bixi Zhong, Lei Wei 0009, Jiaqi Li 0025, Wei Zhang 0241, Huan Fang 0003, Yanda Li, Yinying Lu, Xiaowo Wang
Briefings Bioinform.4
2022 ARIC: accurate and robust inference of cell type proportions from bulk gene expression or DNA methylation data
abstract
Quantifying cell proportions, especially for rare cell types in some scenarios, is of great value in tracking signals associated with certain phenotypes or diseases. Although some methods have been proposed to infer cell proportions from multicomponent bulk data, they are substantially less effective for estimating the proportions of rare cell types which are highly sensitive to feature outliers and collinearity. Here we proposed a new deconvolution algorithm named ARIC to estimate cell type proportions from gene expression or DNA methylation data. ARIC employs a novel two-step marker selection strategy, including collinear feature elimination based on the component-wise condition number and adaptive removal of outlier markers. This strategy can systematically obtain effective markers for weighted $\upsilon$-support vector regression to ensure a robust and precise rare proportion prediction. We showed that ARIC can accurately estimate fractions in both DNA methylation and gene expression data from different experiments. We further applied ARIC to the survival prediction of ovarian cancer and the condition monitoring of chronic kidney disease, and the results demonstrate the high accuracy and robustness as well as clinical potentials of ARIC. Taken together, ARIC is a promising tool to solve the deconvolution problem of bulk data where rare components are of vital importance.
Wei Zhang 0241, Rong Qiao, Bixi Zhong, Xianglin Zhang, Jin Gu, Xuegong Zhang, Lei Wei 0009, Xiaowo Wang
Briefings Bioinform.8
2021 DISMIR: Deep learning-based noninvasive cancer detection by integrating DNA sequence and methylation information of individual cell-free DNA reads
abstract
Detecting cancer signals in cell-free DNA (cfDNA) high-throughput sequencing data is emerging as a novel noninvasive cancer detection method. Due to the high cost of sequencing, it is crucial to make robust and precise predictions with low-depth cfDNA sequencing data. Here we propose a novel approach named DISMIR, which can provide ultrasensitive and robust cancer detection by integrating DNA sequence and methylation information in plasma cfDNA whole-genome bisulfite sequencing (WGBS) data. DISMIR introduces a new feature termed as 'switching region' to define cancer-specific differentially methylated regions, which can enrich the cancer-related signal at read-resolution. DISMIR applies a deep learning model to predict the source of every single read based on its DNA sequence and methylation state and then predicts the risk that the plasma donor is suffering from cancer. DISMIR exhibited high accuracy and robustness on hepatocellular carcinoma detection by plasma cfDNA WGBS data even at ultralow sequencing depths. Further analysis showed that DISMIR tends to be insensitive to alterations of single CpG sites' methylation states, which suggests DISMIR could resist to technical noise of WGBS. All these results showed DISMIR with the potential to be a precise and robust method for low-cost early cancer detection.
Jiaqi Li 0025, Lei Wei 0009, Xianglin Zhang, Wei Zhang 0241, Bixi Zhong, Hairong Lv, Xiaowo Wang
Briefings Bioinform.2
2021 CellTracker: an automated toolbox for single-cell segmentation and tracking of time-lapse microscopy images
abstract
SUMMARY: Recent advances of long-term time-lapse microscopy have made it easy for researchers to quantify cell behavior and molecular dynamics at single-cell resolution. However, the lack of easy-to-use software tools optimized for customized research is still a major challenge for quantitatively understanding biological processes through microscopy images. Here, we present CellTracker, a highly integrated graphical user interface software, for automated cell segmentation and tracking of time-lapse microscopy images. It covers essential steps in image analysis including project management, image pre-processing, cell segmentation, cell tracking, manually correction and statistical analysis such as the quantification of cell size and fluorescence intensity, etc. Furthermore, CellTracker provides an annotation tool and supports model training from scratch, thus proposing a flexible and scalable solution for customized dataset analysis. AVAILABILITY AND IMPLEMENTATION: CellTracker is an open-source software under the GPL-3.0 license. It is implemented in Python and provides an easy-to-use graphical user interface. The source code, instruction manual and demos can be found at https://github.com/WangLabTHU/CellTracker. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tao Hu 0020, Shixiong Xu, Lei Wei 0009, Xuegong Zhang, Xiaowo Wang
Bioinform.3
2021 cfDNApipe: a comprehensive quality control and analysis pipeline for cell-free DNA high-throughput sequencing data
abstract
MOTIVATION: Cell-free DNA (cfDNA) is gaining substantial attention from both biological and clinical fields as a promising marker for liquid biopsy. Many aspects of disease-related features have been discovered from cfDNA high-throughput sequencing (HTS) data. However, there is still a lack of integrative and systematic tools for cfDNA HTS data analysis and quality control (QC). RESULTS: Here, we propose cfDNApipe, an easy-to-use and systematic python package for cfDNA whole-genome sequencing (WGS) and whole-genome bisulfite sequencing (WGBS) data analysis. It covers the entire analysis pipeline for the cfDNA data, including raw sequencing data processing, QC and sophisticated statistical analysis such as detecting copy number variations (CNVs), differentially methylated regions and DNA fragment size alterations. cfDNApipe provides one-command-line-execution pipelines and flexible application programming interfaces for customized analysis. AVAILABILITY AND IMPLEMENTATION: https://xwanglabthu.github.io/cfDNApipe/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wei Zhang 0241, Lei Wei 0009, Bixi Zhong, Jiaqi Li 0025, Shuying He, Juhong Liu, Hairong Lv, Xiaowo Wang
Bioinform.2