VLDB 2026 Research / reviewers in the wild / expert
Jinting Guan
dblp:159/8353
· DBLP profile ↗
13ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-9433-5677ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MISF: Multimodal Data Integration Through Adaptive Similarity Learning and Matrix FactorizationabstractSingle-cell multimodal data can simultaneously provide cellular features at different levels, such as gene expression, chromatin accessibility, and spatial location. The integration of multimodal data can efficiently utilize the information from various views, thereby enhancing the reliability and accuracy of cellular research. The current integration methods mainly focus on obtaining the representation of cells but neglect the representation of genes, not beneficial to cell type-specific gene module analyses. Besides, some integration algorithms only can integrate multi-omics data and cannot be applied to spatial transcriptome data for integrating transcriptomic data and spatial location. To this end, we propose MISF, a Multimodal data Integration algorithm based on adaptive Similarity network learning and matrix Factorization. MISF integrates multimodal data and learns the lower-dimensional representations of cells and genes. We validate the feasibility of MISF on multiple single-cell multi-omics data and spatial transcriptome data and compare it with the existing multi-omics data integration methods as well as spatial transcriptome data analysis algorithms. The results demonstrate that MISF can effectively integrate multimodal data, localize different types of cell clusters, and outline cellular spatial distribution pattern. Furthermore, MISF facilitates cell clustering and cell type-specific gene module analyses, providing new insights for the study of cellular heterogeneity. Fengfan Zhou, Xinqi Chen, Yusheng Jiang, Jinting Guan |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | SpatialDSSC: Estimating Cell Type Abundance and Expression Profile from Spatial Transcriptomic Data
Chenqi Wang, Jinting Guan |
ICIC (26) | 3 |
| 2025 | scTsI: an effective two-stage imputation method for single-cell RNA-seq dataabstractSingle-cell RNA-seq facilitates the understanding of cell types and states and the revealing of the cellular heterogeneity in developmental processes and disease mechanisms. However, the dropout events in single-cell RNA-seq data, in which genes are not detected due to technical noise or limited sequencing depth, seriously affect downstream analyses. Imputation is an effective way to relieve the impact of dropout events. However, the current methods may introduce new noise or modify the high expression values in the imputation process and their performance may be lower than expected when dealing with data with a high dropout rate, facing with different types of data, and aiming at various downstream analyses. We propose a two-stage imputation algorithm, scTsI, for single-cell RNA-seq data. In the first stage, scTsI imputes the zero values using the information of neighboring cells and genes. In the second stage, scTsI transforms the expression matrix into a vector, performs row transformation, and adjusts the imputed values through ridge regression and leveraging bulk RNA-seq data as a constraint. scTsI ensures that the original highly expressed values are unchanged, avoids introducing new noise, and allows sparse matrix input to accelerate imputation. We conduct experiments on a variety of simulated and real data with different dropout rates and compare scTsI with the commonly used imputation methods. The results show that scTsI can restore gene expression and maintain cell-cell similarity across different data dimensions and dropout rates. scTsI can also improve the performance of data visualization, clustering, and cell trajectory inference. Weining Li, Jinting Guan |
Briefings Bioinform. | 3 |
| 2024 | UFGOT: Unbalanced Filter Graph Alignment with Optimal Transport for Cancer Subtyping Based on Multi-omics Data
Yusheng Jiang, Jinting Guan |
ISBRA (1) | 3 |
| 2024 | scINRB: single-cell gene expression imputation with network regularization and bulk RNA-seq dataabstractSingle-cell RNA sequencing (scRNA-seq) facilitates the study of cell type heterogeneity and the construction of cell atlas. However, due to its limitations, many genes may be detected to have zero expressions, i.e. dropout events, leading to bias in downstream analyses and hindering the identification and characterization of cell types and cell functions. Although many imputation methods have been developed, their performances are generally lower than expected across different kinds and dimensions of data and application scenarios. Therefore, developing an accurate and robust single-cell gene expression data imputation method is still essential. Considering to maintain the original cell-cell and gene-gene correlations and leverage bulk RNA sequencing (bulk RNA-seq) data information, we propose scINRB, a single-cell gene expression imputation method with network regularization and bulk RNA-seq data. scINRB adopts network-regularized non-negative matrix factorization to ensure that the imputed data maintains the cell-cell and gene-gene similarities and also approaches the gene average expression calculated from bulk RNA-seq data. To evaluate the performance, we test scINRB on simulated and experimental datasets and compare it with other commonly used imputation methods. The results show that scINRB recovers gene expression accurately even in the case of high dropout rates and dimensions, preserves cell-cell and gene-gene similarities and improves various downstream analyses including visualization, clustering and trajectory inference. Jinting Guan |
Briefings Bioinform. | 3 |
| 2024 | SFINN: inferring gene regulatory network from single-cell and spatial transcriptomic data with shared factor neighborhood and integrated neural networkabstractMOTIVATION: The rise of single-cell RNA sequencing (scRNA-seq) technology presents new opportunities for constructing detailed cell type-specific gene regulatory networks (GRNs) to study cell heterogeneity. However, challenges caused by noises, technical errors, and dropout phenomena in scRNA-seq data pose significant obstacles to GRN inference, making the design of accurate GRN inference algorithms still essential. The recent growth of both single-cell and spatial transcriptomic sequencing data enables the development of supervised deep learning methods to infer GRNs on these diverse single-cell datasets. RESULTS: In this study, we introduce a novel deep learning framework based on shared factor neighborhood and integrated neural network (SFINN) for inferring potential interactions and causalities between transcription factors and target genes from single-cell and spatial transcriptomic data. SFINN utilizes shared factor neighborhood to construct cellular neighborhood network based on gene expression data and additionally integrates cellular network generated from spatial location information. Subsequently, the cell adjacency matrix and gene pair expression are fed into an integrated neural network framework consisting of a graph convolutional neural network and a fully-connected neural network to determine whether the genes interact. Performance evaluation in the tasks of gene interaction and causality prediction against the existing GRN reconstruction algorithms demonstrates the usability and competitiveness of SFINN across different kinds of data. SFINN can be applied to infer GRNs from conventional single-cell sequencing data and spatial transcriptomic data. AVAILABILITY AND IMPLEMENTATION: SFINN can be accessed at GitHub: https://github.com/JGuan-lab/SFINN. Fengfan Zhou, Jinting Guan |
Bioinform. | 3 |
| 2023 | Single-nucleus gene and gene set expression-based similarity network fusion identifies autism molecular subtypesabstractBACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder that is highly phenotypically and genetically heterogeneous. With the accumulation of biological sequencing data, more and more studies shift to molecular subtype-first approach, from identifying molecular subtypes based on genetic and molecular data to linking molecular subtypes with clinical manifestation, which can reduce heterogeneity before phenotypic profiling. RESULTS: In this study, we perform similarity network fusion to integrate gene and gene set expression data of multiple human brain cell types for ASD molecular subtype identification. Then we apply subtype-specific differential gene and gene set expression analyses to study expression patterns specific to molecular subtypes in each cell type. To demonstrate the biological and practical significance, we analyze the molecular subtypes, investigate their correlation with ASD clinical phenotype, and construct ASD molecular subtype prediction models. CONCLUSIONS: The identified molecular subtype-specific gene and gene set expression may be used to differentiate ASD molecular subtypes, facilitating the diagnosis and treatment of ASD. Our method provides an analytical pipeline for the identification of molecular subtypes and even disease subtypes of complex disorders. Guoli Ji, Xilin Gao, Jinting Guan |
BMC Bioinform. | 4 |
| 2022 | MOCSC: a multi-omics data based framework for cancer subtype classificationabstractUsing multi-omics data to achieve the classification of cancer subtypes from the perspective of machine learning can facilitate the understanding of cancer mechanism from the biomolecular aspect, which is of guidance to achieve personalized precision medicine in clinical medicine. To address the problem of cancer subtype classification based on multi-omics data, we propose a framework named MOCSC. Considering the high-dimensional complexity of multi-omics data and the heterogeneity among them, we use the idea of late integration model framework. Specifically, after dividing the training and test sets, for each omics data, the training set is first learned using a stacked sparse denoising autoencoder to extract its feature, and then the extracted feature is supplied to a single-layer neural network to obtain an initial prediction. Then, the initial results of all omics-specific classifiers are integrated and put into a view correlation discovery network for training to obtain the final prediction. Finally, the entire trained model is applied to the test set. The designed model is used to explore the best combination of multi-omics data and then tested on several different datasets. We compare our framework with the existing classification tools, showing its effectiveness in cancer subtype classification problem and also other classification problems based on multi-omics data. Yuanling Ma, Jinting Guan |
BIBM | 2 |
| 2022 | scIAE: an integrative autoencoder-based ensemble classification framework for single-cell RNA-seq dataabstractSingle-cell RNA sequencing (scRNA-seq) allows quantitative analysis of gene expression at the level of single cells, beneficial to study cell heterogeneity. The recognition of cell types facilitates the construction of cell atlas in complex tissues or organisms, which is the basis of almost all downstream scRNA-seq data analyses. Using disease-related scRNA-seq data to perform the prediction of disease status can facilitate the specific diagnosis and personalized treatment of disease. Since single-cell gene expression data are high-dimensional and sparse with dropouts, we propose scIAE, an integrative autoencoder-based ensemble classification framework, to firstly perform multiple random projections and apply integrative and devisable autoencoders (integrating stacked, denoising and sparse autoencoders) to obtain compressed representations. Then base classifiers are built on the lower-dimensional representations and the predictions from all base models are integrated. The comparison of scIAE and common feature extraction methods shows that scIAE is effective and robust, independent of the choice of dimension, which is beneficial to subsequent cell classification. By testing scIAE on different types of data and comparing it with existing general and single-cell-specific classification methods, it is proven that scIAE has a great classification power in cell type annotation intradataset, across batches, across platforms and across species, and also disease status prediction. The architecture of scIAE is flexible and devisable, and it is available at https://github.com/JGuan-lab/scIAE. Qingyang Yin, Jinting Guan, Guoli Ji |
Briefings Bioinform. | 3 |
| 2020 | Gene Screening for Autism Based on Cell-type-specific Predictive ModelsabstractAutism spectrum disorder (ASD), with substantial genetic and phenotypic heterogeneity, is characterized by difficulties in social interaction and communication, and restricted and repetitive behaviors. Recent studies based on bulk RNA-seq data of brains from ASD patients have revealed the affected pathways in ASD, while the cell type heterogeneity of ASD is still needed to be explored. Gene prioritization studies can be conducted for screening gene candidates with high confidence, providing new insights for experimental studies. Based on the single-nucleus RNA-seq data of brains from ASD and healthy individuals, we identify cell-type-specific differential expressed genes by applying three kinds of methods for differential expression analysis and then construct cell-type-specific classification models for ASD adopting the algorithm of stochastic gradient boosting. We find layer 2/3 and 4 excitatory neurons, layer 5/6 cortico-cortical projection neurons, and protoplasmic astrocytes are vulnerable in ASD. Then we calculate gene importance to prioritize cell-type-specific differential expressed genes, and compare the top important genes across different cell types. Our results suggest that causal genes are distinct and dysregulated gene functions are different across brain cells in ASD. The constructed classification models can predict the diagnosis for a nucleus with given cell type, promoting the detection of ASD. The prioritized cell-type-specific genes may be used as potential ASD biomarkers, promoting the development of effective interventions. Yiping Lin, Guoli Ji, Jinting Guan |
BIBM | 4 |
| 2019 | Human brain cell type-specific gene co-expression associated with autism spectrum disorderabstractAutism spectrum disorder (ASD) is a complex neuropsychiatric disorder with substantial phenotypic and genetic heterogeneity. Until now, about a thousand diverse genes have been associated with ASD, while it remains elusive that how disruptions in these different genes can lead to a common clinical phenotype. Therefore, it is essential to understand how ASD candidate genes relate to each other and identify potential shared molecular pathways. Human brain is a highly heterogeneous organ involving multiple cell types. Different functions in different types of cells may be dysregulated in ASD; investigating functional interactions between ASD candidate genes in normal human brain cells may shed new light on the genetic heterogeneity of ASD. To this end, we construct cell type-associated gene co-expression networks based on human brain cell gene expression data. Then we identify seven cell type-specific gene modules and analyze the specific gene functions in each cell type. We also identify six ASD-associated gene modules and study the dysregulated functions in ASD. Lastly, we obtain two ASD-associated cell type-specific gene modules for studying the cell type-specific aberrant functions in ASD. It is found that ASD-associated astrocytes-specific gene modules are relevant to endocytosis, neuron differentiation and cell projection organization, while ASD-associated neurons-specific gene modules are relevant to presynapse, glutamatergic synapse and neuron projection morphogenesis. Our method has been proven to be effective in discovering ASD-associated cell type-specific gene expression pattern. Our findings can promote the study of the heterogeneity of ASD in gene expression between different cell types, providing new insights into the molecular mechanisms underlying the pathogenesis of ASD. Yiping Lin, Shuchao Li, Guoli Ji, Jinting Guan |
BIBM | 4 |
| 2018 | AEGS: identifying aberrantly expressed gene sets for differential variability analysisabstractMotivation: In gene expression studies, differential expression (DE) analysis has been widely used to identify genes with shifted expression mean between groups. Recently, differential variability (DV) analysis has been increasingly applied as analyzing changed expression variability (e.g. the changes in expression variance) between groups may reveal underlying genetic heterogeneity and undetected interactions, which has great implications in many fields of biology. An easy-to-use tool for DV analysis is needed. Results: We develop AEGS for DV analysis, to identify aberrantly expressed gene sets in diseased cases but not in controls. AEGS can rank individual genes in an aberrantly expressed gene set by each gene's relative contribution to the total degree of aberrant expression, prioritizing top genes. AEGS can be used for discovering gene sets with disease-specific expression variability changes. Availability and implementation: AEGS web server is accessible at http://bmi.xmu.edu.cn:8003/AEGS, where a stand-alone AEGS application can also be downloaded. Contact: [email protected]. Jinting Guan, Moliang Chen, Congting Ye, James J. Cai, Guoli Ji |
Bioinform. | 1 |
| 2015 | Genome-wide identification and predictive modeling of polyadenylation sites in eukaryotesabstractPolyadenylation [poly(A)] is a vital step in post-transcriptional processing of pre-mRNA. Alternative polyadenylation is a widespread mechanism of regulating gene expression in eukaryotes. Defining poly(A) sites contributes to the annotation of transcripts' ends and the study of gene regulatory mechanisms. Here, we survey methods for collecting poly(A) sites using high-throughput sequencing technologies and summarize the general processes for genome-wide poly(A) site identifications. We also compare the performances of various poly(A) site prediction models and discuss the relationship between poly(A) site identification from sequencing projects and predictive modeling. Moreover, we attempt to address some potential problems in current researches and propose future directions related to polyadenylation research. Guoli Ji, Jinting Guan, Qingshun Quinn Li |
Briefings Bioinform. | 2 |