Yuansong Zeng

dblp:188/1294 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
16since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2027 SPINE: Context-guided iterative protein inference from spatial transcriptomes
abstract
Spatially resolved protein abundance is critical for understanding cellular function, yet remains difficult to measure at scale, motivating its inference from transcriptomic data as an important intelligent prediction problem in spatial multi-omics. However, existing methods are primarily developed in the single-cell setting and formulate this task as independent point-wise prediction, ignoring the structured spatial dependencies induced by local cellular interactions and tissue architecture. Here, we present SPINE, a neighborhood-guided, flow-matching-inspired iterative refinement framework for protein inference from spatial transcriptomes. Rather than treating protein prediction as independent point-wise regression, SPINE reformulates this task as a conditional refinement problem, in which an initial protein prior is iteratively updated toward the target protein profile under transcriptomic and spatial-neighborhood guidance. To incorporate tissue-context information, SPINE combines neighborhood-aware spatial modeling with expression-based graph structure, while an auxiliary reconstruction branch helps preserve transcriptomic semantics during cross-modal inference. Across paired spatial RNA–protein datasets, SPINE achieves superior overall prediction performance compared with representative RNA-to-protein and multimodal baselines, and better preserves protein-informed biological structure in downstream spatial analyses.
Xingming Lu, Jiangshan Xu, Jiajin Wang, Nuodi Fan, Sijie Wan, Yunqing Fu, Yuansong Zeng
Expert Syst. Appl.8
2025 Cellular-Resolution Reconstruction of Spatial Transcriptomics from Histology Images via Foundation Models and KANs
abstract
Spatial transcriptomics (ST) enables spatially resolved profiling of gene expression within tissues and is commonly accompanied by paired histology images, such as high-resolution H&E-stained sections. Given that histological morphology often correlates with underlying gene expression patterns, the availability of paired images offers a promising opportunity for enhancing spatial resolution. Here we present Hist2Sr, a novel framework that reconstructs cellular-resolution spatial transcriptomics from histology images. Hist2SR first extracts fine-grained morphological representations for individual cells using a foundation model pre-trained on large-scale pathological image datasets. To bridge the scale gap between cell-level features and spot-level gene expression measurements, we adopt a multiple instance learning (MIL) strategy: cell-wise gene expressions within each spot region are aggregated and supervised using the corresponding spot-level transcriptomic profiles. This allows the model to learn informative cell-level features that align with spatial gene expression signals. Finally, a Kolmogorov-Arnold Network (KAN)-based predictor is employed to capture complex nonlinear mappings from cell morphology to gene expression. Benchmark experiments indicate that Hist2SR achieves state-of-the-art performance across various human tissue types.
Ruipeng Huang, Leiming Fang, Wenbing Li, Yuansong Zeng, Yuedong Yang
BIBM4
2025 Drug Sensitivity Inference Through Foundation Model-Driven Contrastive Integration of Bulk and Single-Cell Omics
abstract
Tumor heterogeneity hinders drug response prediction: bulk RNA-seq obscures cell-level resistance, while scRNAseq lacks annotated drug response data. Existing methods transferring knowledge from bulk to single-cell often underuse pretrained models and ignore inter-cell relationships, limiting heterogeneity modeling. We propose COIN (Contrastive Learning-driven Omics Integration Network), integrating labeled bulk RNA-seq with unlabeled scRNA-seq for single-cell drug sensitivity prediction. COIN leverages CellFM, a foundation model pretrained on 100 M cells, and uses a shared feature extractor with contrastive learning to align shared patterns and capture micro-heterogeneity, while an auxiliary reconstruction loss ensures robust representations. Trained solely on bulk data, COIN predicts single-cell sensitivity and outperforms existing methods, with 9% average AUC and 7% AUPR improvements. COIN effectively overcomes the limitations of bulk data for heterogeneity modeling and eliminates the need for costly singlecell drug screening annotations, advancing personalized cancer therapy.
Ningyuan Shangguan, Yuansong Zeng, Wenbing Li, Yuedong Yang
BIBM2
2025 DANet: spatial gene expression prediction from H&E histology images through dynamic alignment
abstract
Predicting spatial gene expression from Hematoxylin and Eosin histology images offers a promising approach to significantly reduce the time and cost associated with gene expression sequencing, thereby facilitating a deeper understanding of tissue architecture and disease mechanisms. Achieving accurate gene expression prediction requires the extraction of highly refined features from pathological images; however, existing methods often struggle to effectively capture fine-grained local details and model gene-gene correlations. Moreover, in bimodal contrastive learning, dynamically and efficiently aligning heterogeneous modalities remains a critical challenge. To address these issues, we propose a novel method for predicting gene expression. First, we introduce a dense connective structure that enables efficient feature reuse, thereby enhancing the capturing and mining of local refinement features. Second, we leverage the state space models to uncover underlying patterns and capture dependencies within 1D gene expression data, enabling more accurate modeling of gene-gene correlations. Furthermore, we design the Residual Kolmogorov-Arnold Network (RKAN) that uses a learnable activation function to dynamically adjust bimodal mappings based on input characteristics. Through continuous parameter updates during contrastive training, RKAN progressively refines the alignment between modalities. Extensive experiments conducted on two publicly available datasets, GSE240429 and HER2+, demonstrate the effectiveness of our approach and its significant improvements over existing methods. Source codes are available at https://github.com/202324131016T/DANet.
Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yuansong Zeng
Briefings Bioinform.5
2024 Dual Interaction and Kernel-Diverse Network for Accurate Drug-Target Binding Affinity Prediction
abstract
Drug-target binding affinities measure the binding strength between drugs and targets, serving as the foundation for computer-aided drug design. The existing methods based on neural network rarely incorporate knowledge from biochemistry, which limits their performance. In this paper, we propose a dual interaction and kernel-diverse network for drug-target binding affinity prediction, termed DTANet, drawing inspiration from the biochemical properties of drugs and targets. Specifically, since functional groups and peptide chains are pivotal in determining the properties of drugs and their targets, and their lengths vary significantly. Drawing inspiration from this, we propose a Kernel-diverse Two-stream Network (KTN) and a Cross-Scale Interaction Module (CSIM) to extract the features of drugs and targets. The binding site of drug and target is crucial for precise affinity prediction. Inspired by this, we introduce the Drug-Target Interaction Module (DTIM) to explore the relationships between drugs and targets to focus on binding sites. Experiments have been conducted on two popular benchmarks: KIBA and Davis. Our DTANet obtains superior performance compared to existing methods on both datasets. Source codes are available at https://github.com/202324131016T/DTANet.
Jin Xie 0005, Jing Nie 0001, Yuansong Zeng
BIBM5
2024 Mamba-DTA: Drug-Target Binding Affinity Prediction with State Space Model
abstract
Biochemical methods for measuring drug-target binding are costly and slow, while deep learning offers a crucial solution. Deep learning methods for predicting drug-target binding affinity have achieved remarkable success, but existing approaches face challenges: adjacent elements in a three-dimensional structure may be far apart in a one-dimensional representation, and long sequences of drug and target data often contain lots of noise. In this paper, we introduce MambaDTA, a novel architecture for drug-target affinity prediction based on the State Space Model (SSM). Mamba-DTA utilizes SSM to model the drug molecules and target molecules and extract more discriminative spatial structural features efficiently and stably. Additionally, we design Interaction-based Selective Filtering (ISF) module to model drug-target interactions and filter out redundant information. The experimental results on two publicly available datasets, namely Davis and KIBA, demonstrate the effectiveness and superiority of our Mamba-DTA. Specifically, Mamba-DTA achieves a relative gain of 13.3% in terms of MAE on the Davis dataset. Source codes are available at https://github.com/202324131016T/Mamba-DTA.
Jin Xie 0005, Jing Nie 0001, Xiaohong Zhang 0002, Yuansong Zeng
BIBM5
2024 Accurately Deciphering Novel Cell Type in Spatially Resolved Single-Cell Data Through Optimal Transport
Mai Luo, Yuansong Zeng, Jianing Chen 0004, Ningyuan Shangguan, Yuedong Yang
ISBRA (2)2
2024 Detecting novel cell type in single-cell chromatin accessibility data via open-set domain adaptation
abstract
Recent advances in single-cell technologies enable the rapid growth of multi-omics data. Cell type annotation is one common task in analyzing single-cell data. It is a challenge that some cell types in the testing set are not present in the training set (i.e. unknown cell types). Most scATAC-seq cell type annotation methods generally assign each cell in the testing set to one known type in the training set but neglect unknown cell types. Here, we present OVAAnno, an automatic cell types annotation method which utilizes open-set domain adaptation to detect unknown cell types in scATAC-seq data. Comprehensive experiments show that OVAAnno successfully identifies known and unknown cell types. Further experiments demonstrate that OVAAnno also performs well on scRNA-seq data. Our codes are available online at https://github.com/lisaber/OVAAnno/tree/master.
Yuefan Lin, Zixiang Pan, Yuansong Zeng, Yuedong Yang, Zhiming Dai
Briefings Bioinform.3
2023 Identifying spatial domain by adapting transcriptomics with histology through contrastive learning
abstract
Recent advances in spatial transcriptomics have enabled measurements of gene expression at cell/spot resolution meanwhile retaining both the spatial information and the histology images of the tissues. Accurately identifying the spatial domains of spots is a vital step for various downstream tasks in spatial transcriptomics analysis. To remove noises in gene expression, several methods have been developed to combine histopathological images for data analysis of spatial transcriptomics. However, these methods either use the image only for the spatial relations for spots, or individually learn the embeddings of the gene expression and image without fully coupling the information. Here, we propose a novel method ConGI to accurately exploit spatial domains by adapting gene expression with histopathological images through contrastive learning. Specifically, we designed three contrastive loss functions within and between two modalities (the gene expression and image data) to learn the common representations. The learned representations are then used to cluster the spatial domains on both tumor and normal spatial transcriptomics datasets. ConGI was shown to outperform existing methods for the spatial domain identification. In addition, the learned representations have also been shown powerful for various downstream tasks, including trajectory inference, clustering, and visualization.
Yuansong Zeng, Mai Luo, Jianing Chen 0004, Zixiang Pan, Yutong Lu, Weijiang Yu, Yuedong Yang
Briefings Bioinform.1
2023 Identifying B-cell epitopes using AlphaFold2 predicted structures and pretrained language model
abstract
MOTIVATION: Identifying the B-cell epitopes is an essential step for guiding rational vaccine development and immunotherapies. Since experimental approaches are expensive and time-consuming, many computational methods have been designed to assist B-cell epitope prediction. However, existing sequence-based methods have limited performance since they only use contextual features of the sequential neighbors while neglecting structural information. RESULTS: Based on the recent breakthrough of AlphaFold2 in protein structure prediction, we propose GraphBepi, a novel graph-based model for accurate B-cell epitope prediction. For one protein, the predicted structure from AlphaFold2 is used to construct the protein graph, where the nodes/residues are encoded by ESM-2 learning representations. The graph is input into the edge-enhanced deep graph neural network (EGNN) to capture the spatial information in the predicted 3D structures. In parallel, a bidirectional long short-term memory neural networks (BiLSTM) are employed to capture long-range dependencies in the sequence. The learned low-dimensional representations by EGNN and BiLSTM are then combined into a multilayer perceptron for predicting B-cell epitopes. Through comprehensive tests on the curated epitope dataset, GraphBepi was shown to outperform the state-of-the-art methods by more than 5.5% and 44.0% in terms of AUC and AUPR, respectively. A web server is freely available at http://bio-web1.nscc-gz.cn/app/graphbepi. AVAILABILITY AND IMPLEMENTATION: The datasets, pre-computed features, source codes, and the trained model are available at https://github.com/biomed-AI/GraphBepi.
Yuansong Zeng, Zhuoyi Wei, Qianmu Yuan, Weijiang Yu, Yutong Lu, Jianzhao Gao, Yuedong Yang
Bioinform.1
2022 A Meta-learning based Graph-Hierarchical Clustering Method for Single Cell RNA-Seq Data
abstract
Single cell sequencing techniques enable researchers view complex bio-tissues from a more precise perspective to identify cell types. However, more and more recent works have been done to find more detailed subtypes within already known cell types. Here, we present MeHi-SCC, a method which utilized meta-learning protocol and brought in multi scRNA-seq datasets’ information in order to assist graph-based hierarchical sub-clustering process. In result, MeHi-SCC outperformed current-prevailing scRNA clustering methods and successfully identified cell subtypes in two large scale cell atlas. Our codes and datasets are available online at https://github.com/biomed-AI/MeHi-SCC
Zixiang Pan, Yuansong Zeng, Yuefan Lin, Weijiang Yu, Haokun Zhang, Yuedong Yang
BIBM2
2022 SCdenoise: a reference-based scRNA-seq denoising method using semi-supervised learning
abstract
scRNA-seq is a promising technology to perform unbiased, high-throughput, and high-resolution transcriptome analysis at single-cell resolution. The raw data usually suffers from noise and low quality, such as dropout events, which hinder downstream analysis. Thus, it is essential to improve the quality of single-cell data. Although many methods have been developed for denoising scRNA-seq data, the existing methods mainly focus on finding the relationship within the data itself without fully utilizing other datasets with annotated cell labels. Here, we proposed SCdenoise, a semi-supervised denoising method, to denoise unlabeled target data based on annotated cells in the reference datasets, which could utilize biological characteristics hidden in the high-quality reference datasets. Extensive downstream analyses showed that our method outperformed state-of-the-art methods on both simulated and real datasets for single-cell data analyses, including gene expression recovery, differential analysis, and clustering analysis. The source code is available at https://github.com/zhongfqi/SCdenoise-.
Fengqi Zhong, Yuansong Zeng, Yuedong Yang
BIBM2
2022 A robust and scalable graph neural network for accurate single-cell classification
abstract
Single-cell RNA sequencing (scRNA-seq) techniques provide high-resolution data on cellular heterogeneity in diverse tissues, and a critical step for the data analysis is cell type identification. Traditional methods usually cluster the cells and manually identify cell clusters through marker genes, which is time-consuming and subjective. With the launch of several large-scale single-cell projects, millions of sequenced cells have been annotated and it is promising to transfer labels from the annotated datasets to newly generated datasets. One powerful way for the transferring is to learn cell relations through the graph neural network (GNN), but traditional GNNs are difficult to process millions of cells due to the expensive costs of the message-passing procedure at each training epoch. Here, we have developed a robust and scalable GNN-based method for accurate single-cell classification (GraphCS), where the graph is constructed to connect similar cells within and between labelled and unlabeled scRNA-seq datasets for propagation of shared information. To overcome the slow information propagation of GNN at each training epoch, the diffused information is pre-calculated via the approximate Generalized PageRank algorithm, enabling sublinear complexity over cell numbers. Compared with existing methods, GraphCS demonstrates better performance on simulated, cross-platform, cross-species and cross-omics scRNA-seq datasets. More importantly, our model provides a high speed and scalability on large datasets, and can achieve superior performance for 1 million cells within 50 min.
Yuansong Zeng, Zhuoyi Wei, Zixiang Pan, Yutong Lu, Yuedong Yang
Briefings Bioinform.1
2022 Spatial transcriptomics prediction from histology jointly through Transformer and graph neural networks
abstract
The rapid development of spatial transcriptomics allows the measurement of RNA abundance at a high spatial resolution, making it possible to simultaneously profile gene expression, spatial locations of cells or spots, and the corresponding hematoxylin and eosin-stained histology images. It turns promising to predict gene expression from histology images that are relatively easy and cheap to obtain. For this purpose, several methods are devised, but they have not fully captured the internal relations of the 2D vision features or spatial dependency between spots. Here, we developed Hist2ST, a deep learning-based model to predict RNA-seq expression from histology images. Around each sequenced spot, the corresponding histology image is cropped into an image patch and fed into a convolutional module to extract 2D vision features. Meanwhile, the spatial relations with the whole image and neighbored patches are captured through Transformer and graph neural network modules, respectively. These learned features are then used to predict the gene expression by following the zero-inflated negative binomial distribution. To alleviate the impact by the small spatial transcriptomics data, a self-distillation mechanism is employed for efficient learning of the model. By comprehensive tests on cancer and normal datasets, Hist2ST was shown to outperform existing methods in terms of both gene expression prediction and spatial region identification. Further pathway analyses indicated that our model could reserve biological information. Thus, Hist2ST enables generating spatial transcriptomics data from histology images for elucidating molecular signatures of tissues.
Yuansong Zeng, Zhuoyi Wei, Weijiang Yu, Yuchen Yuan, Bingling Li, Zhonghui Tang, Yutong Lu, Yuedong Yang
Briefings Bioinform.1
2022 A parameter-free deep embedded clustering method for single-cell RNA-seq data
abstract
Clustering analysis is widely used in single-cell ribonucleic acid (RNA)-sequencing (scRNA-seq) data to discover cell heterogeneity and cell states. While many clustering methods have been developed for scRNA-seq analysis, most of these methods require to provide the number of clusters. However, it is not easy to know the exact number of cell types in advance, and experienced determination is not always reliable. Here, we have developed ADClust, an automatic deep embedding clustering method for scRNA-seq data, which can accurately cluster cells without requiring a predefined number of clusters. Specifically, ADClust first obtains low-dimensional representation through pre-trained autoencoder and uses the representations to cluster cells into initial micro-clusters. The clusters are then compared in between by a statistical test, and similar micro-clusters are merged into larger clusters. According to the clustering, cell representations are updated so that each cell will be pulled toward centers of its assigned cluster and similar clusters, while cells are separated to keep distances between clusters. This is accomplished through jointly optimizing the carefully designed clustering and autoencoder loss functions. This merging process continues until convergence. ADClust was tested on 11 real scRNA-seq datasets and was shown to outperform existing methods in terms of both clustering performance and the accuracy on the number of the determined clusters. More importantly, our model provides high speed and scalability for large datasets.
Yuansong Zeng, Zhuoyi Wei, Fengqi Zhong, Zixiang Pan, Yutong Lu, Yuedong Yang
Briefings Bioinform.1
2021 scAdapt: virtual adversarial domain adaptation network for single cell RNA-seq data classification across platforms and species
abstract
In single cell analyses, cell types are conventionally identified based on expressions of known marker genes, whose identifications are time-consuming and irreproducible. To solve this issue, many supervised approaches have been developed to identify cell types based on the rapid accumulation of public datasets. However, these approaches are sensitive to batch effects or biological variations since the data distributions are different in cross-platforms or species predictions. In this study, we developed scAdapt, a virtual adversarial domain adaptation network, to transfer cell labels between datasets with batch effects. scAdapt used both the labeled source and unlabeled target data to train an enhanced classifier and aligned the labeled source centroids and pseudo-labeled target centroids to generate a joint embedding. The scAdapt was demonstrated to outperform existing methods for classification in simulated, cross-platforms, cross-species, spatial transcriptomic and COVID-19 immune datasets. Further quantitative evaluations and visualizations for the aligned embeddings confirm the superiority in cell mixing and the ability to preserve discriminative cluster structure present in the original datasets.
Yuansong Zeng, Huiying Zhao, Yuedong Yang
Briefings Bioinform.3
2020 Accurately Clustering Single-cell RNA-seq data by Capturing Structural Relations between Cells through Graph Convolutional Network
abstract
Recent advances in single-cell RNA sequencing (scRNA-seq) technologies provide a great opportunity to study gene expression at cellular resolution, and the scRNA-seq data has been routinely conducted to unfold cell heterogeneity and diversity. A critical step for the scRNA-seq analyses is to cluster the same type of cells, and many methods have been developed for cell clustering. However, existing clustering methods are limited to extract the representations from expression data of individual cells, while ignoring the high-order structural relations between cells. Here, we proposed a new method (GraphSCC) to cluster cells based on scRNA-seq data by accounting structural relations between cells through a graph convolutional network. The representation learned from the graph convolutional network, together with another representation output from a denoising autoencoder network, are optimized by a dual self-supervised module for better cell clustering. Extensive experiments indicate that GraphSCC model outperforms state-of-the-art methods in various evaluation metrics on both simulated and real datasets.
Yuansong Zeng, Jiahua Rao, Yutong Lu, Yuedong Yang
BIBM1
2018 Efficient wear leveling for inodes of file systems on persistent memories
abstract
Existing persistent memory file systems achieve high-performance file accesses by exploiting advanced characteristics of persistent memories (PMs), such as PCM. However, they ignore the limited endurance of PMs. Particularly, the frequently updated inodes are stored on fixed locations throughout their lifetime, which can easily damage PM with common file operations. To address such issues, we propose a new mechanism, Virtualized Inode (VInode), for the wear leveling of inodes of persistent memory file systems. In VInode, we develop an algorithm called Pages as Communicating Vessels (PCV) to efficiently find and migrate the heavily written inodes. We implement VInode in SIMFS, a typical persistent memory file system. Experiments are conducted with well-known benchmarks. Compared with original SIMFS, experimental results show that VInode can reduce the maximum value and standard deviation of the write counts of pages to 1800x and 6200x lower, respectively.
Xianzhang Chen, Edwin H.-M. Sha, Yuansong Zeng, Chaoshu Yang, Weiwen Jiang, Qingfeng Zhuge
DATE3
2016 The design of an efficient swap mechanism for hybrid DRAM-NVM systems
abstract
Non-Volatile Memory (NVM) is becoming an attractive candidate to be the swap area in embedded systems for its near-DRAM speed, low energy consumption, high density, and byte-addressability. Swapping data from DRAM out to NVM, however, can cause large performance/energy penalty and deplete the lifetime of NVM. Traditional swap mechanisms may need to be re-studied. Even through there are several swap mechanisms proposed for the hybrid DRAM-NVM systems, most of them have limited performance without considering the data access features of applications.
Xianzhang Chen, Edwin H.-M. Sha, Weiwen Jiang, Qingfeng Zhuge, Junxi Chen, Jiejie Qin, Yuansong Zeng
EMSOFT7