Shixiong Zhang 0002

dblp:19/906-2 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-0314-9199ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 scMGCL: accurate and efficient integration representation of single-cell multi-omics data
abstract
MOTIVATION: Single-cell multi-omics data integration is essential for understanding cellular states and disease mechanisms, yet integrating heterogeneous data modalities remains a challenge. We present scMGCL, a graph contrastive learning framework for robust integration of single-cell ATAC-seq and RNA-seq data. Our approach leverages self-supervised learning on cell-cell similarity graphs, in which each modality's graph structure serves as an augmentation for the other. This cross-modality contrastive paradigm enables the learning of biologically meaningful, shared representations while preserving modality-specific features. RESULTS: Benchmarking against state-of-the-art methods demonstrates that scMGCL outperforms others in cell-type clustering, label transfer accuracy, and preservation of marker-gene correlations. Additionally, scMGCL significantly improves computational efficiency, reducing runtime and memory usage. The method's effectiveness is further validated through extensive analyses of cell-type similarity and functional consistency, providing a powerful tool for multi-omics data exploration. AVAILABILITY AND IMPLEMENTATION: Code and datasets are released at https://github.com/zlCreator/scMGCL.
Zhenglong Cheng, Risheng Lu, Shixiong Zhang 0002
Bioinform.3
2022 High-throughput single-cell RNA-seq data imputation and characterization with surrogate-assisted automated deep learning
abstract
Single-cell RNA sequencing (scRNA-seq) technologies have been heavily developed to probe gene expression profiles at single-cell resolution. Deep imputation methods have been proposed to address the related computational challenges (e.g. the gene sparsity in single-cell data). In particular, the neural architectures of those deep imputation models have been proven to be critical for performance. However, deep imputation architectures are difficult to design and tune for those without rich knowledge of deep neural networks and scRNA-seq. Therefore, Surrogate-assisted Evolutionary Deep Imputation Model (SEDIM) is proposed to automatically design the architectures of deep neural networks for imputing gene expression levels in scRNA-seq data without any manual tuning. Moreover, the proposed SEDIM constructs an offline surrogate model, which can accelerate the computational efficiency of the architectural search. Comprehensive studies show that SEDIM significantly improves the imputation and clustering performance compared with other benchmark methods. In addition, we also extensively explore the performance of SEDIM in other contexts and platforms including mass cytometry and metabolic profiling in a comprehensive manner. Marker gene detection, gene ontology enrichment and pathological analysis are conducted to provide novel insights into cell-type identification and the underlying mechanisms. The source code is available at https://github.com/li-shaochuan/SEDIM.
Xiangtao Li, Shaochuan Li, Shixiong Zhang 0002, Ka-Chun Wong
Briefings Bioinform.4
2022 DeepMotifSyn: a deep learning approach to synthesize heterodimeric DNA motifs
Jiecong Lin, Xingjian Chen, Shixiong Zhang 0002, Ka-Chun Wong
Briefings Bioinform.4
2022 scWMC: weighted matrix completion-based imputation of scRNA-seq data via prior subspace information
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) can provide insight into gene expression patterns at the resolution of individual cells, which offers new opportunities to study the behavior of different cell types. However, it is often plagued by dropout events, a phenomenon where the expression value of a gene tends to be measured as zero in the expression matrix due to various technical defects. RESULTS: In this article, we argue that borrowing gene and cell information across column and row subspaces directly results in suboptimal solutions due to the noise contamination in imputing dropout values. Thus, to impute more precisely the dropout events in scRNA-seq data, we develop a regularization for leveraging that imperfect prior information to estimate the true underlying prior subspace and then embed it in a typical low-rank matrix completion-based framework, named scWMC. To evaluate the performance of the proposed method, we conduct comprehensive experiments on simulated and real scRNA-seq data. Extensive data analysis, including simulated analysis, cell clustering, differential expression analysis, functional genomic analysis, cell trajectory inference and scalability analysis, demonstrate that our method produces improved imputation results compared to competing methods that benefits subsequent downstream analysis. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/XuYuanchi/scWMC and test data is available at https://doi.org/10.5281/zenodo.6832477. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yanchi Su, Fuzhou Wang, Shixiong Zhang 0002, Ka-Chun Wong, Xiangtao Li
Bioinform.3
2021 Identification of pan-cancer Ras pathway activation with deep learning
abstract
The identification of hidden responders is often an essential challenge in precision oncology. A recent attempt based on machine learning has been proposed for classifying aberrant pathway activity from multiomic cancer data. However, we note several critical limitations there, such as high-dimensionality, data sparsity and model performance. Given the central importance and broad impact of precision oncology, we propose nature-inspired deep Ras activation pan-cancer (NatDRAP), a deep neural network (DNN) model, to address those restrictions for the identification of hidden responders. In this study, we develop the nature-inspired deep learning model that integrates bulk RNA sequencing, copy number and mutation data from PanCanAltas to detect pan-cancer Ras pathway activation. In NatDRAP, we propose to synergize the nature-inspired artificial bee colony algorithm with different gradient-based optimizers in one framework for optimizing DNNs in a collaborative manner. Multiple experiments were conducted on 33 different cancer types across PanCanAtlas. The experimental results demonstrate that the proposed NatDRAP can provide superior performance over other benchmark methods with strong robustness towards diagnosing RAS aberrant pathway activity across different cancer types. In addition, gene ontology enrichment and pathological analysis are conducted to reveal novel insights into the RAS aberrant pathway activity identification and characterization. NatDRAP is written in Python and available at https://github.com/lixt314/NatDRAP1.
Xiangtao Li, Shaochuan Li, Yunhe Wang 0002, Shixiong Zhang 0002, Ka-Chun Wong
Briefings Bioinform.4
2021 Deep embedded clustering with multiple objectives on scRNA-seq data
abstract
In recent years, single-cell RNA sequencing (scRNA-seq) technologies have been widely adopted to interrogate gene expression of individual cells; it brings opportunities to understand the underlying processes in a high-throughput manner. Deep embedded clustering (DEC) was demonstrated successful in high-dimensional sparse scRNA-seq data by joint feature learning and cluster assignment for identifying cell types simultaneously. However, the deep network architecture for embedding clustering is not trivial to optimize. Therefore, we propose an evolutionary multiobjective DEC by synergizing the multiobjective evolutionary optimization to simultaneously evolve the hyperparameters and architectures of DEC in an automatic manner. Firstly, a denoising autoencoder is integrated into the DEC to project the high-dimensional sparse scRNA-seq data into a low-dimensional space. After that, to guide the evolution, three objective functions are formulated to balance the model's generality and clustering performance for robustness. Meanwhile, migration and mutation operators are proposed to optimize the objective functions to select the suitable hyperparameters and architectures of DEC in the multiobjective framework. Multiple comparison analyses are conducted on twenty synthetic data and eight real data from different representative single-cell sequencing platforms to validate the effectiveness. The experimental results reveal that the proposed algorithm outperforms other state-of-the-art clustering methods under different metrics. Meanwhile, marker genes identification, gene ontology enrichment and pathology analysis are conducted to reveal novel insights into the cell type identification and characterization mechanisms.
Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong
Briefings Bioinform.2
2021 Elucidating transcriptomic profiles from single-cell RNA sequencing data using nature-inspired compressed sensing
abstract
Gene-expression profiling can define the cell state and gene-expression pattern of cells at the genetic level in a high-throughput manner. With the development of transcriptome techniques, processing high-dimensional genetic data has become a major challenge in expression profiling. Thanks to the recent widespread use of matrix decomposition methods in bioinformatics, a computational framework based on compressed sensing was adopted to reduce dimensionality. However, compressed sensing requires an optimization strategy to learn the modular dictionaries and activity levels from the low-dimensional random composite measurements to reconstruct the high-dimensional gene-expression data. Considering this, here we introduce and compare four compressed sensing frameworks coming from nature-inspired optimization algorithms (CSCS, ABCCS, BACS and FACS) to improve the quality of the decompression process. Several experiments establish that the three proposed methods outperform benchmark methods on nine different datasets, especially the FACS method. We illustrate therefore, the robustness and convergence of FACS in various aspects; notably, time complexity and parameter analyses highlight properties of our proposed FACS. Furthermore, differential gene-expression analysis, cell-type clustering, gene ontology enrichment and pathology analysis are conducted, which bring novel insights into cell-type identification and characterization mechanisms from different perspectives. All algorithms are written in Python and available at https://github.com/Philyzh8/Nature-inspired-CS.
Zhuohan Yu, Chuang Bian, Genggeng Liu, Shixiong Zhang 0002, Ka-Chun Wong, Xiangtao Li
Briefings Bioinform.4
2021 Evolving Transcriptomic Profiles From Single-Cell RNA-Seq Data Using Nature-Inspired Multiobjective Optimization
abstract
Transcriptomic profiling plays an important role in post-genomic analysis. Especially, the single-cell RNA-seq technology has advanced our understanding of gene expression from cell population level into individual cell level. Many computational methods have been proposed to decipher transcriptomic profiles from those RNA-seq data. However, most of the related algorithms suffer from realistic restrictions such as high dimensionality and premature convergence. In this paper, we propose and formulate an evolutionary multiobjective blind compressed sensing (EMOBCS) to address those problems for evolving transcriptomic profiles from single-cell RNA-seq data. In the proposed framework, to characterize various gene expression profile models, two objective functions including chi-squared kernel score and euclidean distance of different gene expression profiles are formulated. After that, multiobjective blind compressed sensing based on artificial bee colony is designed to optimize the two objective functions on single-cell RNA-seq data by proposing a rank probability model and two new search strategies into the cooperative convolution framework in an unbiased manner. To demonstrate its effectiveness, extensive experiments have been conducted, comparing the proposed algorithm with 14 algorithms including eight state-of-the-art algorithms and six different EMOBCS algorithms under different search strategies on 10 single-cell RNA-seq datasets and one case study. The experimental results reveal that the proposed algorithm is better than or comparable with those compared algorithms. Furthermore, we also conduct the time complexity analysis, convergence analysis, and parameter analysis to demonstrate various properties of EMOBCS.
Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Multiobjective Genome-Wide RNA-Binding Event Identification From CLIP-Seq Data
abstract
RNA-binding proteins (RBPs) are the master regulators of mRNA processing, which are vital players for the post-transcriptional control of gene expression. In recent years, crosslinking immunoprecipitation sequencing (CLIP-seq) technologies have enabled us to sequence massive amounts of genome-wide RNA-binding event data. Its increasing availability provides opportunities to identify protein-RNA interactions on a genome-wide scale. Genome-wide RNA-binding event detection methods have been developed to the understanding of the proteins' functions within cellular processes. Unfortunately, those methods often suffer from realistic restrictions, such as high costs, intensive computation, high dimensionality, numerical instability, and data sparsity. We present a computational method [multiobjective forest algorithm (MFA)] to identify protein-RNA interactions from CLIP-seq data by synergizing multiobjective biogeography-based optimization (BBO) with random forest (RF). Since most of the tree-structured classifiers in RF are unnecessarily bulky with extra time costs and memory consumption, multiobjective BBO is designed to prune the unsuitable tree-structured classifiers dynamically. Moreover, to direct the evolution dynamics of the MFA, two objective functions are formulated to balance model generality and complexity for robust performance. To validate our MFA method, we compare its performance across 31 large-scale CLIP-seq datasets. The experimental results demonstrate that MFA can obtain superior performance over the current state-of-the-art methods. Mechanistic insights are also revealed and discussed to explore the multifaceted aspects of MFA through data source importance analysis, matrix rank estimations, seeding component perturbations, and multiobjective optimization methodology comparisons.
Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong
IEEE Trans. Cybern.2
2021 Nature-Inspired Compressed Sensing for Transcriptomic Profiling From Random Composite Measurements
abstract
Transcriptomic profiling is a high-throughput approach to measure gene expression levels under different experimental conditions at different timings. With the development of the related technologies such as single-cell RNA-Seq, the dimensions of gene expression data are increased to hundreds of thousands or more for high-resolution insights. There is a long-lasting challenge in exploiting the relations between transcriptomic profiles and random composite measurements. To address it, we proposed a mathematical framework based on differential evolution (global search) with the help of compressed sensing (local search) termed as DECS. Exploiting the inherent sparse nature of gene expression data, the proposed DECS method can learn the sparse module dictionaries and levels from the low-dimensional random composite measurements for reconstructing the high-dimensional gene expression data with significant orders of magnitude (e.g. 200 × ). Several experiments were conducted to compare DECS with three benchmark methods, demonstrating that the proposed DECS outperforms the benchmark methods and can recover most of the gene expression patterns. The underlying reasons are discussed and illustrated by revealing the related mechanistic insights through extensive benchmarks on nine GSE datasets and their sensitivity analysis.
Shixiong Zhang 0002, Xiangtao Li, Qiuzhen Lin, Ka-Chun Wong
IEEE Trans. Cybern.1
2020 Nature-Inspired Multiobjective Epistasis Elucidation from Genome-Wide Association Studies
abstract
In recent years, the detection of epistatic interactions of multiple genetic variants on the causes of complex diseases brings a significant challenge in genome-wide association studies (GWAS). However, most of the existing methods still suffer from algorithmic limitations such as single-objective optimization, intensive computational requirement, and premature convergence. In this paper, we propose and formulate an epistatic interaction multi-objective artificial bee colony algorithm based on decomposition (EIMOABC/D) to address those problems for genetic interaction detection in genome-wide association studies. First, to direct the genetic interaction detection, two objective functions are formulated to characterize various epistatic models; rank probability model is proposed to sort each population into different nondomination levels based on the fast nondominated sorting approach. After that, the mutual information based local search algorithm is proposed to guide the population search for disease model evaluations in an unbiased manner. To validate the effectiveness of EIMOABC/D, we compare EIMOABC/D against seven state-of-the-art methods on 77 epistatic models including eight small-scale epistatic models with marginal effects, eight large-scale epistatic models with marginal effects, 60 large-scale epistatic models without any marginal effect, and one case study. The experimental results indicate that our proposed algorithm EIMOABC/D outperforms seven state-of-the-art methods on those epistatic models. Furthermore, time complexity analysis and parameter analysis are conducted to demonstrate various properties of our proposed algorithm.
Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong
IEEE ACM Trans. Comput. Biol. Bioinform.2
2019 Single-cell RNA-seq interpretations using evolutionary multiobjective ensemble pruning
abstract
MOTIVATION: In recent years, single-cell RNA sequencing enables us to discover cell types or even subtypes. Its increasing availability provides opportunities to identify cell populations from single-cell RNA-seq data. Computational methods have been employed to reveal the gene expression variations among multiple cell populations. Unfortunately, the existing ones can suffer from realistic restrictions such as experimental noises, numerical instability, high dimensionality and computational scalability. RESULTS: We propose an evolutionary multiobjective ensemble pruning algorithm (EMEP) that addresses those realistic restrictions. Our EMEP algorithm first applies the unsupervised dimensionality reduction to project data from the original high dimensions to low-dimensional subspaces; basic clustering algorithms are applied in those new subspaces to generate different clustering results to form cluster ensembles. However, most of those cluster ensembles are unnecessarily bulky with the expense of extra time costs and memory consumption. To overcome that problem, EMEP is designed to dynamically select the suitable clustering results from the ensembles. Moreover, to guide the multiobjective ensemble evolution, three cluster validity indices including the overall cluster deviation, the within-cluster compactness and the number of basic partition clusters are formulated as the objective functions to unleash its cell type discovery performance using evolutionary multiobjective optimization. We applied EMEP to 55 simulated datasets and seven real single-cell RNA-seq datasets, including six single-cell RNA-seq dataset and one large-scale dataset with 3005 cells and 4412 genes. Two case studies are also conducted to reveal mechanistic insights into the biological relevance of EMEP. We found that EMEP can achieve superior performance over the other clustering algorithms, demonstrating that EMEP can identify cell populations clearly. AVAILABILITY AND IMPLEMENTATION: EMEP is written in Matlab and available at https://github.com/lixt314/EMEP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiangtao Li, Shixiong Zhang 0002, Ka-Chun Wong
Bioinform.2
2019 Synergizing CRISPR/Cas9 off-target predictions for ensemble insights and practical applications
abstract
MOTIVATION: The RNA-guided CRISPR/Cas9 system has been widely applied to genome editing. CRISPR/Cas9 system can effectively edit the on-target genes. Nonetheless, it has recently been demonstrated that many homologous off-target genomic sequences could be mutated, leading to unexpected gene-editing outcomes. Therefore, a plethora of tools were proposed for the prediction of off-target activities of CRISPR/Cas9. Nonetheless, each computational tool has its own advantages and drawbacks under diverse conditions. It is hardly believed that a single tool is optimal for all conditions. Hence, we would like to explore the ensemble learning potential on synergizing multiple tools with genomic annotations together to enhance its predictive abilities. RESULTS: We proposed an ensemble learning framework which synergizes multiple tools together to predict the off-target activities of CRISPR/Cas9 in different combinations. Interestingly, the ensemble learning using AdaBoost outperformed other individual off-target predictive tools. We also investigated the effect of evolutionary conservation (PhyloP and PhastCons) and chromatin annotations (ChromHMM and Segway) and found that only PhyloP can enhance the predictive capabilities further. Case studies are conducted to reveal ensemble insights into the off-target predictions, demonstrating how the current study can be applied in different genomic contexts. The best prediction predicted by AdaBoost is up to 0.9383 (AUC) and 0.2998 (PRC) that outperforms other classifiers. This is ascribable to the fact that AdaBoost introduces a new weak classifier (i.e. decision stump) in each iteration to learn the DNA sequences that were misclassified as off-targets until a small error rate is reached iteratively. AVAILABILITY AND IMPLEMENTATION: The source codes are freely available on GitHub at https://github.com/Alexzsx/CRISPR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shixiong Zhang 0002, Xiangtao Li, Qiuzhen Lin, Ka-Chun Wong
Bioinform.1
2019 A novel method based on FTS with both GA-FCM and multifactor BPNN for stock forecasting
Wenyu Zhang 0001, Shixiong Zhang 0002, Shuai Zhang 0002, Dejian Yu, NingNing Huang
Soft Comput.2
2017 A multi-factor and high-order stock forecast model based on Type-2 FTS using cuckoo search and self-adaptive harmony search
Wenyu Zhang 0001, Shixiong Zhang 0002, Shuai Zhang 0002, Dejian Yu, NingNing Huang
Neurocomputing2