Shuilin Jin

dblp:134/8402 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-2318-432XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 8 since 2021
YearPublicationVenuePosition
2026 Accurate and efficient HiChIP interaction detection by modeling restriction enzyme cut site density as biological signal
abstract
HiChIP enables high-resolution chromatin interaction mapping, but existing methods generally overlook restriction enzyme (RE) cut site density or treat it as a technical bias requiring normalization or removal, discarding chromatin accessibility information that distinguishes functional regulatory elements. Here we introduce sintHiChIP to address this methodological gap. sintHiChIP explicitly models RE cut site density as a biological signal and integrates Gaussian kernel smoothing with distance-dependent statistics, allowing detection of chromatin loops while capturing local regulatory heterogeneity. Moreover, the algorithm employs adaptive probability distributions to resolve inherent data overdispersion and sparsity dynamically. Validation against independent datasets and comparison with existing methods demonstrate that sintHiChIP reliably recovers canonical chromatin loops, exhibiting distinct superiority in regulatory H3K27ac environments and comparable accuracy in structural cohesin contexts. Notably, sintHiChIP achieves exceptional precision in predicting CRISPRi experiments and reveals highly coherent cell-type-specific genetic regulatory networks. Executing efficiently on standard workstations, our method delivers a promising analytical framework for functional 3D genomic studies.
Weiyue Ding, Quanhong Liu, Yiyuan Guo, Chiping Zhang, Shuilin Jin
Briefings Bioinform.6
2025 scBCN: deep learning-based batch correction network for integration of heterogeneous single-cell data
abstract
With the continuous application of single-cell data, effectively correcting batch effects and accurately identifying cell types has emerged as a critical challenge in biomedical research. However, existing methods often struggle to disentangle technical effects from genuine biological variation, limiting their performance on heterogeneous datasets. Here, we introduce single-cell Batch Correction Network (scBCN), an integration framework that combines robust inter-batch similar cluster identification with a deep residual neural network to correct batch effects while preserving biological variability. To evaluate the performance of scBCN, we conduct benchmarking experiments on various simulated and real datasets, demonstrating its superiority in both batch correction and biological variation conservation. Furthermore, scBCN shows its applicability in cross-species and cross-omics data integration, underscoring its potential for uncovering and characterizing cell type-specific gene expression patterns.
Yang Zhou 0038, Xingzhi Wang, Harry Qin, Shuilin Jin
Briefings Bioinform.5
2025 reDA: differential abundance testing on scATAC-seq data using random walk with restart
abstract
SUMMARY: Identifying cell states associated with disease progression or experimental perturbations from single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) data is critical for unraveling disease pathogenesis. However, the high dimensionality, extreme sparsity, and nearly binary nature of scATAC-seq data pose significant challenges. Here, we present reDA, a cluster-free computational framework that performs differential abundance testing based on the random walk with restart. Through comprehensive experiments on simulated and real datasets, reDA outperforms six baseline methods, demonstrating superior accuracy, computational efficiency, and the ability to capture disease-specific molecular signatures. AVAILABILITY AND IMPLEMENTATION: The reDA along with detailed documentation is freely available at https://github.com/Jinsl-lab/reDA. It can be seamlessly integrated into existing scATAC-seq analysis workflows.
Jiao Hua, Lu Ba, Tianyun He, Boran Yang, Shuilin Jin
Bioinform.7
2025 scGT: integration algorithm for single-cell RNA-seq and ATAC-seq based on graph transformer
abstract
MOTIVATION: Multi-omics analysis of individual cells offers remarkable opportunities for exploring the dynamics and relationships of gene regulatory states across large atlas data. However, the current integration algorithms have limited performance, largely due to ignoring the impact of correlation features within the dataset on the discrepancies between omics. RESULTS: In this study, we propose scGT, a model based on Graph Transformer for single-cell RNA-seq and ATAC-seq data, which leverages the robust graph structures strengthened by correlation features present in each raw dataset to harmonize representations of multi-omics data, enabling the integration of multi-omics and effective label transfer. We compare scGT with other state-of-the-art methods on paired and unpaired datasets. The results show that scGT accomplishes more accurate label transfer and is capable of integrating datasets with millions of cells. Meanwhile, scGT achieves better performance for preserving biological variation during integration. AVAILABILITY AND IMPLEMENTATION: The source code and data used in this article can be found at https://github.com/Jinsl-lab/scGT.
Yunjing Qi, Yulong Kan, Shuilin Jin
Bioinform.4
2025 Integration of unpaired single cell omics data by deep transfer graph convolutional network
abstract
The rapid advance of large-scale atlas-level single cell RNA sequences and single-cell chromatin accessibility data provide extraordinary avenues to broad and deep insight into complex biological mechanism. Leveraging the datasets and transfering labels from scRNA-seq to scATAC-seq will empower the exploration of single-cell omics data. However, the current label transfer methods have limited performance, largely due to the lower capable of preserving fine-grained cell populations and intrinsic or extrinsic heterogeneity between datasets. Here, we present a robust deep transfer model based graph convolutional network, scTGCN, which achieves versatile performance in preserving biological variation, while achieving integration hundreds of thousands cells in minutes with low memory consumption. We show that scTGCN is powerful to the integration of mouse atlas data and multimodal data generated from APSA-seq and CITE-seq. Thus, scTGCN shows high label transfer accuracy and effectively knowledge transfer across different modalities.
Yulong Kan, Yunjing Qi, Zhongxiao Zhang, Xikeng Liang, Shuilin Jin
PLoS Comput. Biol.6
2025 Spatial domain identification method based on multi-view graph convolutional network and contrastive learning
abstract
Spatial transcriptomics is a rapidly developing field of single-cell genomics that quantitatively measures gene expression while providing spatial information within tissues. A key challenge in spatial transcriptomics is identifying spatially structured domains, which involves analyzing transcriptomic data to find clusters of cells with similar expression patterns and their spatial distribution. To address these challenges, we propose a novel deep-learning method called DMGCN for domain identification. The process begins with preprocessing that constructs two types of graphs: a spatial graph based on Euclidean distance and a feature graph based on Cosine distance. These graphs represent spatial positions and gene expressions, respectively. The embeddings of both graphs are generated using a multi-view graph convolutional encoder with an attention mechanism, enabling separate and co-convolution of the graphs, as well as corrupted feature convolution for contrastive learning. Finally, a fully connected network (FCN) decoder is employed to generate domain labels and reconstruct gene expressions for downstream analysis. Experimental results demonstrate that DMGCN consistently outperforms state-of-the-art methods in various tasks, including spatial clustering, trajectory inference, and gene expression broadcasting.
Xikeng Liang, Shutong Xiao, Lu Ba, Yuhui Feng, Zhicheng Ma, Fatima T. Adilova, Shuilin Jin
PLoS Comput. Biol.8
2022 Correction: SDImpute: A statistical block imputation method based on cell-level and gene-level information for dropouts in single-cell RNA-seq data
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1009118.].
Zicen Zhao, Shuilin Jin
PLoS Comput. Biol.4
2021 SDImpute: A statistical block imputation method based on cell-level and gene-level information for dropouts in single-cell RNA-seq data
abstract
The single-cell RNA sequencing (scRNA-seq) technologies obtain gene expression at single-cell resolution and provide a tool for exploring cell heterogeneity and cell types. As the low amount of extracted mRNA copies per cell, scRNA-seq data exhibit a large number of dropouts, which hinders the downstream analysis of the scRNA-seq data. We propose a statistical method, SDImpute (Single-cell RNA-seq Dropout Imputation), to implement block imputation for dropout events in scRNA-seq data. SDImpute automatically identifies the dropout events based on the gene expression levels and the variations of gene expression across similar cells and similar genes, and it implements block imputation for dropouts by utilizing gene expression unaffected by dropouts from similar cells. In the experiments, the results of the simulated datasets and real datasets suggest that SDImpute is an effective tool to recover the data and preserve the heterogeneity of gene expression across cells. Compared with the state-of-the-art imputation methods, SDImpute improves the accuracy of the downstream analysis including clustering, visualization, and differential expression analysis.
Zicen Zhao, Shuilin Jin
PLoS Comput. Biol.4
2020 NDRindex: a method for the quality assessment of single-cell RNA-Seq preprocessing data
abstract
BACKGROUND: Single-cell RNA sequencing can be used to fairly determine cell types, which is beneficial to the medical field, especially the many recent studies on COVID-19. Generally, single-cell RNA data analysis pipelines include data normalization, size reduction, and unsupervised clustering. However, different normalization and size reduction methods will significantly affect the results of clustering and cell type enrichment analysis. Choices of preprocessing paths is crucial in scRNA-Seq data mining, because a proper preprocessing path can extract more important information from complex raw data and lead to more accurate clustering results. RESULTS: We proposed a method called NDRindex (Normalization and Dimensionality Reduction index) to evaluate data quality of outcomes of normalization and dimensionality reduction methods. The method includes a function to calculate the degree of data aggregation, which is the key to measuring data quality before clustering. For the five single-cell RNA sequence datasets we tested, the results proved the efficacy and accuracy of our index. CONCLUSIONS: This method we introduce focuses on filling the blanks in the selection of preprocessing paths, and the result proves its effectiveness and accuracy. Our research provides useful indicators for the evaluation of RNA-Seq data.
Ruiyu Xiao, Guoshan Lu, Wanqian Guo, Shuilin Jin
BMC Bioinform.4
2020 ERDS-Exome: A Hybrid Approach for Copy Number Variant Detection from Whole-Exome Sequencing Data
abstract
Copy number variants (CNVs) play important roles in human disease and evolution. With the rapid development of next-generation sequencing technologies, many tools have been developed for inferring CNVs based on whole-exome sequencing (WES) data. However, as a result of the sparse distribution of exons in the genome, the limitations of the WES technique, and the nature of high-level signal noises in WES data, the efficacy of these variants remains less than desirable. Thus, there is need for the development of an effective tool to achieve a considerable power in WES CNVs discovery. In the present study, we describe a novel method, Estimation by Read Depth (RD) with Single-nucleotide variants from exome sequencing data (ERDS-exome). ERDS-exome employs a hybrid normalization approach to normalize WES data and to incorporate RD and single-nucleotide variation information together as a hybrid signal into a paired hidden Markov model to infer CNVs from WES data. Based on systematic evaluations of real data from the 1000 Genomes Project using other state-of-the-art tools, we observed that ERDS-exome demonstrates higher sensitivity and provides comparable or even better specificity than other tools. ERDS-exome is publicly available at: https://erds-exome.github.io.
Renjie Tan, Jixuan Wang, Liran Juan, Qing Zhan, Shuilin Jin, Qinghua Jiang
IEEE ACM Trans. Comput. Biol. Bioinform.9
2019 NDRindex: A method for the quality assessment of single-cell RNA-Seq preprocessing data
abstract
Background: Single-cell RNA sequencing can be used to determine cell types in an unbiased way. Normally, the analysis pipeline of single-cell RNA data includes data n ormalization, dimension reduction and unsupervised clustering. However, different normalization and dimension reduction methods will influence the results of clustering and cell type enrichment analysis significantly. Choices of preprocessing paths is crucial in scRNA-Seq data mining because an appropriate preprocessing path can extract more important information from complex raw data and lead to a more accurate clustering result. Results: We propose a method called NDRindex(Normalization and Dimensionality Reduction index) to evaluate single-cell RNA-seq data quality. The method includes a function that calculates the degree of aggregation of data, which is the key to benchmarking data quality before clustering. For five single-cell RNA sequencing data sets we tested, the result shows the effectiveness and the accuracy of our index. Conclusions: This method we introduce focuses on filling the blanks in the selection of preprocessing paths and the result proves its effectiveness and accuracy. Our study provides a useful indicator for RNA-Seq data assessment.
Ruiyu Xiao, Guoshan Lu, Shuilin Jin
BIBM3
2019 BIN1 rs744373 variant shows different association with Alzheimer's disease in Caucasian and Asian populations
abstract
Abstract Background The association between BIN1 rs744373 variant and Alzheimer’s disease (AD) had been identified by genome-wide association studies (GWASs) as well as candidate gene studies in Caucasian populations. But in East Asian populations, both positive and negative results had been identified by association studies. Considering the smaller sample sizes of the studies in East Asian, we believe that the results did not have enough statistical power. Results We conducted a meta-analysis with 71,168 samples (22,395 AD cases and 48,773 controls, from 37 studies of 19 articles). Based on the additive model, we observed significant genetic heterogeneities in pooled populations as well as Caucasians and East Asians. We identified a significant association between rs744373 polymorphism with AD in pooled populations (P = 5 × 10− 07, odds ratio (OR) = 1.12, and 95% confidence interval (CI) 1.07–1.17) and in Caucasian populations (P = 3.38 × 10− 08, OR = 1.16, 95% CI 1.10–1.22). But in the East Asian populations, the association was not identified (P = 0.393, OR = 1.057, and 95% CI 0.95–1.15). Besides, the regression analysis suggested no significant publication bias. The results for sensitivity analysis as well as meta-analysis under the dominant model and recessive model remained consistent, which demonstrated the reliability of our finding. Conclusions The large-scale meta-analysis highlighted the significant association between rs744373 polymorphism and AD risk in Caucasian populations but not in the East Asian populations.
Zhifa Han, Wenyang Zhou, Jian Zong, Yang Hu 0008, Shuilin Jin, Qinghua Jiang
BMC Bioinform.9
2019 ProbPFP: a multiple sequence alignment algorithm combining hidden Markov model optimized by particle swarm optimization with partition function
abstract
BACKGROUND: During procedures for conducting multiple sequence alignment, that is so essential to use the substitution score of pairwise alignment. To compute adaptive scores for alignment, researchers usually use Hidden Markov Model or probabilistic consistency methods such as partition function. Recent studies show that optimizing the parameters for hidden Markov model, as well as integrating hidden Markov model with partition function can raise the accuracy of alignment. The combination of partition function and optimized HMM, which could further improve the alignment's accuracy, however, was ignored by these researches. RESULTS: A novel algorithm for MSA called ProbPFP is presented in this paper. It intergrate optimized HMM by particle swarm with partition function. The algorithm of PSO was applied to optimize HMM's parameters. After that, the posterior probability obtained by the HMM was combined with the one obtained by partition function, and thus to calculate an integrated substitution score for alignment. In order to evaluate the effectiveness of ProbPFP, we compared it with 13 outstanding or classic MSA methods. The results demonstrate that the alignments obtained by ProbPFP got the maximum mean TC scores and mean SP scores on these two benchmark datasets: SABmark and OXBench, and it got the second highest mean TC scores and mean SP scores on the benchmark dataset BAliBASE. ProbPFP is also compared with 4 other outstanding methods, by reconstructing the phylogenetic trees for six protein families extracted from the database TreeFam, based on the alignments obtained by these 5 methods. The result indicates that the reference trees are closer to the phylogenetic trees reconstructed from the alignments obtained by ProbPFP than the other methods. CONCLUSIONS: We propose a new multiple sequence alignment method combining optimized HMM and partition function in this paper. The performance validates this method could make a great improvement of the alignment's accuracy.
Qing Zhan, Shuilin Jin, Renjie Tan, Qinghua Jiang
BMC Bioinform.3
2018 ProbPFP: A Multiple Sequence Alignment Algorithm Combining Partition Function and Hidden Markov Model with Particle Swarm Optimization
Qing Zhan, Shuilin Jin, Renjie Tan, Qinghua Jiang
BIBM3
2018 BIN1 rs744373 Variant Is Significantly Associated with Alzheimer's Disease in Caucasian but Not East Asian Populations
Zhifa Han, Wenyang Zhou, Jian Zong, Yang Hu 0008, Shuilin Jin, Qinghua Jiang
ICIC (1)9
2018 EWAS: epigenome-wide association study software 2.0
abstract
Motivation: With the development of biotechnology, DNA methylation data showed exponential growth. Epigenome-wide association study (EWAS) provide a systematic approach to uncovering epigenetic variants underlying common diseases/phenotypes. But the EWAS software has lagged behind compared with genome-wide association study (GWAS). To meet the requirements of users, we developed a convenient and useful software, EWAS2.0. Results: EWAS2.0 can analyze EWAS data and identify the association between epigenetic variations and disease/phenotype. On the basis of EWAS1.0, we have added more distinctive features. EWAS2.0 software was developed based on our 'population epigenetic framework' and can perform: (i) epigenome-wide single marker association study; (ii) epigenome-wide methylation haplotype (meplotype) association study and (iii) epigenome-wide association meta-analysis. Users can use EWAS2.0 to execute chi-square test, t-test, linear regression analysis, logistic regression analysis, identify the association between epi-alleles, identify the methylation disequilibrium (MD) blocks, calculate the MD coefficient, the frequency of meplotype and Pearson's correlation coefficients and carry out meta-analysis and so on. Finally, we expect EWAS2.0 to become a popular software and be widely used in epigenome-wide associated studies in the future. Availability and implementation: The EWAS software is freely available at http://www.ewas.org.cn or http://www.bioapp.org/ewas.
Linna Zhao, Simeng Hu, Xiuling Song, Hongchao Lv, Qinghua Jiang, Guiyou Liu, Shuilin Jin, Mingzhi Liao, Rennan Feng, Fanwu Kong, Liangde Xu, Yongshuai Jiang
Bioinform.12
2017 DTWscore: differential expression and cell clustering analysis for time-series single-cell RNA-seq data
abstract
BACKGROUND: The development of single-cell RNA sequencing has enabled profound discoveries in biology, ranging from the dissection of the composition of complex tissues to the identification of novel cell types and dynamics in some specialized cellular environments. However, the large-scale generation of single-cell RNA-seq (scRNA-seq) data collected at multiple time points remains a challenge to effective measurement gene expression patterns in transcriptome analysis. RESULTS: We present an algorithm based on the Dynamic Time Warping score (DTWscore) combined with time-series data, that enables the detection of gene expression changes across scRNA-seq samples and recovery of potential cell types from complex mixtures of multiple cell types. CONCLUSIONS: The DTWscore successfully classify cells of different types with the most highly variable genes from time-series scRNA-seq data. The study was confined to methods that are implemented and available within the R framework. Sample datasets and R packages are available at https://github.com/xiaoxiaoxier/DTWscore .
Shuilin Jin, Guiyou Liu, Xiurui Zhang, Deliang Wu, Yang Hu 0008, Chiping Zhang, Qinghua Jiang, Yadong Wang 0001
BMC Bioinform.2
2016 ERDS-pe: A paired hidden Markov model for copy number variant detection from whole-exome sequencing data
abstract
Detecting copy number variants (CNVs) is an essential part in variant calling process. Here, we describe a novel method ERDS-pe to detect CNVs from whole-exome sequencing (WES) data. ERDS-pe first employs principal component analysis to normalize WES data. Then, ERDS-pe incorporates read depth signal and single-nucleotide variation information together as a hybrid signal into a paired hidden Markov model to infer CNVs from WES data. Experimental results on real human WES data show that ERDS-pe demonstrates higher sensitivity and provides comparable or even better specificity than other tools. ERDS-pe is publicly available at: https://github.com/microtan0902/erds-pe.
Renjie Tan, Jixuan Wang, Guoqiang Wan, Zhijie Han 0002, Wenyang Zhou, Shuilin Jin, Qinghua Jiang, Yadong Wang 0001
BIBM9