VLDB 2026 Research / reviewers in the wild / expert
Bingjun Li
dblp:01/3804
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-modal Spatial Clustering for Spatial Transcriptomics Utilizing High-resolution Histology ImagesabstractUnderstanding the intricate cellular environment within biological tissues is crucial for uncovering insights into complex biological functions. While single-cell RNA sequencing has significantly enhanced our understanding of cellular states, it lacks the spatial context to fully comprehend the cellular environment. Spatial transcriptomics (ST) addresses this limitation by enabling transcriptome-wide profiling while preserving spatial context. One of the principal challenges in ST data analysis is spatial clustering. Modern ST sequencing procedures typically include a high-resolution histology image, which has been shown in previous studies to be closely connected to gene expression profiles. However, current spatial clustering methods often fail to fully utilize the image information, limiting their ability to capture critical spatial and cellular interactions.In this study, we propose the spatial transcriptomics multimodal clustering (stMMC) model, a novel contrastive learningbased deep learning approach that integrates gene expression data with histology image features through a multi-modal parallel graph autoencoder. We tested stMMC against four state-of-the-art baseline models on two public ST datasets. The experiments demonstrated the superior performance of stMMC in terms of ARI and NMI and an ablation study validated the contributions of key components. Bingjun Li, Mostafa Karami, Masum Shah Junayed, Sheida Nabavi |
BIBM | 1 |
| 2024 | Improved allele-specific single-cell copy number estimation in low-coverage DNA-sequencingabstractMOTIVATION: Advances in whole-genome single-cell DNA sequencing (scDNA-seq) have led to the development of numerous methods for detecting copy number aberrations (CNAs), a key driver of genetic heterogeneity in cancer. While most of these methods are limited to the inference of total copy number, some recent approaches now infer allele-specific CNAs using innovative techniques for estimating allele-frequencies in low coverage scDNA-seq data. However, these existing allele-specific methods are limited in their segmentation strategies, a crucial step in the CNA detection pipeline. RESULTS: We present SEACON (Single-cell Estimation of Allele-specific COpy Numbers), an allele-specific copy number profiler for scDNA-seq data. SEACON uses a Gaussian Mixture Model to identify latent copy number states and breakpoints between contiguous segments across cells, filters the segments for high-quality breakpoints using an ensemble technique, and adopts several strategies for tolerating noisy read-depth and allele frequency measurements. Using a wide array of both real and simulated datasets, we show that SEACON derives accurate copy numbers and surpasses existing approaches under numerous experimental conditions, and identify its strengths and weaknesses. AVAILABILITY AND IMPLEMENTATION: SEACON is implemented in Python and is freely available open-source from https://github.com/NabaviLab/SEACON and https://doi.org/10.5281/zenodo.12727008. Samson Weiner, Bingjun Li, Sheida Nabavi |
Bioinform. | 2 |
| 2024 | A multimodal graph neural network framework for cancer molecular subtype classificationabstractBACKGROUND: The recent development of high-throughput sequencing has created a large collection of multi-omics data, which enables researchers to better investigate cancer molecular profiles and cancer taxonomy based on molecular subtypes. Integrating multi-omics data has been proven to be effective for building more precise classification models. Most current multi-omics integrative models use either an early fusion in the form of concatenation or late fusion with a separate feature extractor for each omic, which are mainly based on deep neural networks. Due to the nature of biological systems, graphs are a better structural representation of bio-medical data. Although few graph neural network (GNN) based multi-omics integrative methods have been proposed, they suffer from three common disadvantages. One is most of them use only one type of connection, either inter-omics or intra-omic connection; second, they only consider one kind of GNN layer, either graph convolution network (GCN) or graph attention network (GAT); and third, most of these methods have not been tested on a more complex classification task, such as cancer molecular subtypes. RESULTS: In this study, we propose a novel end-to-end multi-omics GNN framework for accurate and robust cancer subtype classification. The proposed model utilizes multi-omics data in the form of heterogeneous multi-layer graphs, which combine both inter-omics and intra-omic connections from established biological knowledge. The proposed model incorporates learned graph features and global genome features for accurate classification. We tested the proposed model on the Cancer Genome Atlas (TCGA) Pan-cancer dataset and TCGA breast invasive carcinoma (BRCA) dataset for molecular subtype and cancer subtype classification, respectively. The proposed model shows superior performance compared to four current state-of-the-art baseline models in terms of accuracy, F1 score, precision, and recall. The comparative analysis of GAT-based models and GCN-based models reveals that GAT-based models are preferred for smaller graphs with less information and GCN-based models are preferred for larger graphs with extra information. Bingjun Li, Sheida Nabavi |
BMC Bioinform. | 1 |
| 2023 | scGEMOC, A Graph Embedded Contrastive Learning Single-cell Multiomics Clustering ModelabstractRecent advancements in single-cell multiomics sequencing create new research opportunities but also pose challenges, particularly in cell clustering. One major challenge is feature fusion. Early fusion models are robust but ignore the unique distributions of omics and cannot handle various omic dimensions. Most current clustering methods use late fusion, employing independent encoders for each omic. However, the extracted omic features belong to different latent spaces, leading to difficulties in aligning omics. Additionally, current cell clustering methods do not incorporate prior biological knowledge, such as interactions within and across omics, which has been shown plays a key role in defining cell types.To address these shortcomings, we propose a novel, scalable, end-to-end clustering method, called single-cell graph embedding multiomics cluster (scGEMOC). scGEMOC utilizes prior biological knowledge to represent inter- and intra-omics connections as a heterogeneous graph. It applies graph embedding to aggregate omics interaction data as a pseudo omic and employs contrastive learning for effectively aligning omics in the latent space. We evaluated scGEMOC on three public datasets against five state-of-the-art baseline models. scGEMOC achieves superior clustering performance compared to the baseline models on all datasets. An ablation study confirms the significant contribution of each component and identifies the most impactful one. Bingjun Li, Sheida Nabavi |
BIBM | 1 |
| 2021 | Single-cell RNA sequencing data clustering using graph convolutional networksabstractSingle-cell RNA sequencing (scRNAseq) makes it possible to analyze gene expression profiles at the individual cell scale and to discover intrinsic and extrinsic cellular processes in biological research. Cell clustering is one of the most important steps in analyzing scRNAseq data. With rapid developments of single cell sequencing technologies, scRNAseq data grow in size and heterogeneity. However, traditional clustering methods like Kmeans with or without dimension reduction methods, cannot handle high sparse and massive scRNAseq data. Although some deep learning based methods have been proposed to denoise the data and cluster cells simultaneously, learning informative representations of cells for accurate cell clustering is still a challenging problem to be solved. In this work, we propose a deep learning model that combines a deep graph convolutional network (GCN) and a self-supervised mechanism. The GCN considers not only the gene expressions but also the relationship between cells to represent cells. The self-supervised mechanism is employed to provide the clustering assignments of cells. Moreover, we utilize the negative log-likelihood of the negative binomial (NB) function as loss in the data reconstruction due to the assumption that genes expression values can be represented by the NB model. We compared the performance of our proposed method with those of the existing clustering methods for scRNAseq data and conventional clustering methods. Results show that our method achieves better performance in terms of accuracy, adjusted random index (ARI), and normalized mutual information (NMI). Bingjun Li, Sheida Nabavi |
BIBM | 2 |
| 2008 | Hybrid forecasting method of GM(1, 1) disaster model with application to regional ggain productionabstractEach technique has its own drawback and advantage. There is no method that is powerful in any problems. Therefore, the hybridization of two or more different techniques is important to overcome the disadvantages of the individual techniques. In this paper, a data sequence having a linear tendency with upper/positive and lower/negative aberrances is analyzed. Based on the linear regression analysis, the data sequence is classified into three parts: upper/positive aberrant data, lower/negative aberrant data and normal data. Then introducing the grey disaster forecast analysis, we establish three models: a GM(1,1) disaster model based on upper aberrant data, a GM(1,1) disaster model based on lower aberrant data and a linear regression model based on the remaining normal data. Using the established models, we obtain aberrant forecasting values at oncoming aberrant time points by GM(1,1) from upper and lower aberrant data, and normal forecasting value obtained by the linear regression function. Applying it to the prediction of regional grain production, we demonstrate the good performance and effectiveness of the proposed hybrid method. Bingjun Li, Masahiro Inuiguchi |
SMC | 1 |