Shulin Wang

dblp:24/396 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorComputer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 GenBench: A Comprehensive Benchmark Dataset for Sequence-Based Genomic Models
Yanlei Kang, Shulin Wang
ICIC (30)4
2024 GRNNLink: Predicting gene regulatory links from single-cell RNA-seq data using graph recurrent neural network
abstract
Single-cell RNA sequencing (scRNA-seq) technology offers unprecedented opportunities for inferring gene regulatory networks (GRNs) at the genome level. However, scRNA-seq data is highly sparse and has a low signal-to-noise ratio with significant dropout. Many unsupervised or self-supervised models have been proposed to infer GRNs from large RNA-seq datasets, but few are suitable for scRNA-seq data. Recent research confirms that transcription factor (TF)-DNA binding data enables supervised GRN inference. In this paper, we propose a novel framework called GRNNLink, which leverages known GRNs to infer potential regulatory relationships between genes. First, we preprocess the raw scRNA-seq data. Then, we introduce an interactive graph encoder based on a graph recurrent neural network (GRNN) to refine gene features by capturing the correlations between network nodes. Finally, matrix completion is performed using the node correlation features to predict GRNs. To evaluate model performance, we compare GRNNLink with six existing GRN reconstruction methods across seven scRNA-seq datasets. The results demonstrate that our method exhibits high robustness and accuracy.
Shulin Wang
BIBM2
2024 MRF-XGBLC: Large-scale gene regulatory network inference based on multi-model fusion
abstract
Gene regulatory networks (GRNs) are crucial for revealing gene interactions and understanding cellular biological mechanisms. However, the high dimensionality and nonlinearity of gene expression data make accurate inference and reconstruction of large-scale GRNs a core computational challenge in systems biology. This study introduces a novel approach, termed MRF-XGBLC, for the reconstruction of large-scale Gene Regulatory Networks (GRNs) utilizing steady-state and time-series gene expression data through nonlinear ordinary differential equations. Firstly, MRF-XGBLC uses the maximum information coefficient (MIC) for dimensionality reduction and eliminates redundant regulatory relationships by calculating the MIC between factors as a prior step in model processing. Furthermore, recognizing the superior performance of the Lasso-Cox model in survival analysis, the feature fusion algorithm of this paper incorporates a hybrid model of XGBoost (eXtreme Gradient Boosting), RF (Random Forest), and Lasso-Cox (Least Absolute Shrinkage and Selection Operator-Cox proportional hazards regression model) integration to effectively train the nonlinear ordinary differential equations, thus improving the accuracy and stability of the inference algorithm. Extensive experiments on datasets of varying sizes demonstrate significant improvements over state-of-the-art methods. Cross-validation experiments on real gene datasets confirm the robustness and effectiveness of MRF-XGBLC.
Shulin Wang, Shaoliang Peng
BIBM4
2024 On Pipelined GCN with Communication-Efficient Sampling and Inclusion-Aware Caching
abstract
Graph convolutional network (GCN) has achieved enormous success in learning structural information from unstructured data. As graphs become increasingly large, distributed training for GCNs is severely prolonged by frequent cross-worker communications. Existing efforts to improve the training efficiency often come at the expense of GCN performance, while the communication overhead persists. In this paper, we propose PSC-GCN, a holistic pipelined framework for distributed GCN training with communication-efficient sampling and inclusion-aware caching, to address the communication bottleneck while ensuring satisfactory model performance. Specifically, we devise an asynchronous pre-fetching scheme to retrieve stale statistics (features, embedding, gradient) of boundary nodes in advance, such that the embedding aggregation and model update are pipelined with statistics transmission. To alleviate communication volume and staleness effect, we introduce a variance-reduction based sampling policy, which prioritizes inner nodes over boundary ones for reducing the access frequency to remote neighbors, thus mitigating cross-worker statistics exchange. Complementing graph sampling, a feature caching module is co-designed to buffer hot nodes with high inclusion probability, ensuring that frequently sampled nodes will be available in local memory. Extensive evaluations on real-world datasets show the superiority of PSC-GCN over state-of-the-art methods, where we can reduce training time by 72%-80% without sacrificing model accuracy.
Shulin Wang, Xiong Wang 0006, Yuqing Li 0001, Hai Jin 0001
INFOCOM1
2024 SAE-Impute: imputation for single-cell data via subspace regression and auto-encoders
abstract
BACKGROUND: Single-cell RNA sequencing (scRNA-seq) technology has emerged as a crucial tool for studying cellular heterogeneity. However, dropouts are inherent to the sequencing process, known as dropout events, posing challenges in downstream analysis and interpretation. Imputing dropout data becomes a critical concern in scRNA-seq data analysis. Present imputation methods predominantly rely on statistical or machine learning approaches, often overlooking inter-sample correlations. RESULTS: To address this limitation, We introduced SAE-Impute, a new computational method for imputing single-cell data by combining subspace regression and auto-encoders for enhancing the accuracy and reliability of the imputation process. Specifically, SAE-Impute assesses sample correlations via subspace regression, predicts potential dropout values, and then leverages these predictions within an autoencoder framework for interpolation. To validate the performance of SAE-Impute, we systematically conducted experiments on both simulated and real scRNA-seq datasets. These results highlight that SAE-Impute effectively reduces false negative signals in single-cell data and enhances the retrieval of dropout values, gene-gene and cell-cell correlations. Finally, We also conducted several downstream analyses on the imputed single-cell RNA sequencing (scRNA-seq) data, including the identification of differential gene expression, cell clustering and visualization, and cell trajectory construction. CONCLUSIONS: These results once again demonstrate that SAE-Impute is able to effectively reduce the droupouts in single-cell dataset, thereby improving the functional interpretability of the data.
Shulin Wang
BMC Bioinform.3
2024 Fluid-Shuttle: Efficient Cloud Data Transmission Based on Serverless Computing Compression
abstract
Nowadays, there exists a lot of cross-region data transmission demand on the cloud. It is promising to use serverless computing for data compressing to save the total data size. However, it is challenging to estimate the data transmission time and monetary cost with serverless compression. In addition, minimizing the data transmission cost is non-trivial due to the enormous parameter space. This paper focuses on this problem and makes the following contributions: 1) We propose empirical data transmission time and monetary cost models based on serverless compression. It can also predict compression information, e.g., ratio and speed using chunk sampling and machine learning techniques. 2) For single-task cloud data transmission, we propose two efficient parameter search methods based on Sequential Quadratic Programming (SQP) and Eliminate then Divide and Conquer (EDC) with proven error upper bounds. Besides, we propose a parameter fine-tuning strategy to deal with transmission bandwidth variance. 3) Furthermore, for multi-task scenarios, a parameter search method based on dynamic programming and numerical computation is proposed. We have implemented the system called Fluid-Shuttle, which includes straggler optimization, cache optimization, and the autoscaling decompression mechanism. Finally, we evaluate the performance of Fluid-Shuttle with various workloads and applications on the real-world AWS serverless computing platform. Experimental results show that the proposed approach can improve the parameter search efficiency by over$3\times $compared with the state-of-art methods and achieves better parameter quality. In addition, our approach achieves higher time efficiency and lower monetary cost compared with competing cloud data transmission approaches.
Rong Gu 0001, Shulin Wang, Haipeng Dai 0001, Zhaokang Wang, Wenjie Bao, Jiaqi Zheng 0001, Yaofeng Tu, Yihua Huang 0001, Lianyong Qi, Xiaolong Xu 0001, Wan-Chun Dou, Guihai Chen
IEEE/ACM Trans. Netw.2
2023 Ponzi Scheme Identification of Smart Contract Based on Multi Feature Fusion
Xiaoxiao Jiang, Mingdong Xie, Shulin Wang
ICIC (4)3
2023 Time and Cost-Efficient Cloud Data Transmission based on Serverless Computing Compression
Rong Gu 0001, Haipeng Dai 0001, Shulin Wang, Zhaokang Wang, Yaofeng Tu, Yihua Huang 0001, Guihai Chen
INFOCOM4
2022 A heterogeneous graph cross-omics attention model for single-cell representation learning
abstract
Single-cell multi-omics sequencing technologies allow simultaneous measurement of transcriptome and epigenome profiles in the same cell, providing unprecedented opportunities to dissect cell heterogeneity. Despite great efforts, conjoint analysis of single-cell multi-omics data still suffers from sparsity, high dimensionality and binary. In this study, we present a heterogeneous graph cross-omics attention model (scHGA), a computational tool based on a heterogeneous graph neural network combining two attention mechanisms to jointly analyze single-cell multi-omics data based on different protocols data, including SNARE-seq, scMT-seq and sci-CAR. To avoid the cell heterogeneity of single-omics data, scHGA automatically learns a cell association graph to capture neighbor information. The latent representation of aggregated cells generated by hierarchical attention can fuse knowledge across different omics to dissect cellular heterogeneity, providing a better scheme to characterize the features of cells. scHGA is an effective exploration of graph neural networks in single-cell multi-omics analysis, providing new insights into the understanding of single-cell sequencing data.
Yue Liu 0041, Shulin Wang, Wei Zhang 0089, Xiangxiang Zeng, Chee Keong Kwoh 0001
BIBM3
2021 LPI-FKLGCN: Predicting LncRNA-Protein Interactions Through Fast Kernel Learning and Graph Convolutional Network
Wen Li 0009, Shulin Wang, Hu Guo
ISBRA2
2020 GeoAI-based Epidemic Control with Geo-Social Data Sharing on Blockchain
abstract
Epidemics especially those caused by major contagious diseases have entailed huge losses in human history. The fights have thus never stopped to prevent pandemics. Due to its acute outbreak, is generally susceptible to the population regardless of ages, so strict quarantine of the infections becomes the most effective means for the epidemic control, which has been proved in the prevention of other contagious diseases such as SARS and H1N1. The key strategy widely used to find infected and suspected patients is still the epidemiological tracking of confirmed cases. However, this may fail to identify infections especially when patients do not show any symptoms. Therefore, the approach to rapid, effective, and simple infection identification is essential to prevent the spread of a contagious disease. This paper proposes to leverage a social apps and Geospatial artificial intelligence (GeoAI) with Blockchain to effectively identify infections with privacy concern. Since people widely use social apps, a large scale of social data with geospatial information could be easily collected and kept on Blockchain with privacy preservation, which thus provides a framework of decentralized, tamper-proof, and privacy-preserved information sharing. With the support of GeoAI, which analyzes the spatial distribution of diseases from the shared data, we could study the influence factors based on spatial propagation of contagious diseases for infection identification. Since WeChat is widely used in China, we take COVID-19 as an example to use the experiments on real-life datasets demonstrate the effectiveness of our method, and provide insight into epidemic control in terms of geo-social data sharing.
Shaoliang Peng, Li Xiong 0015, Qiang Qu 0001, Shulin Wang
HealthCom6
2020 A deep metric learning algorithm for similarity measure of the gene expression profile
abstract
Clustering gene expression profiles is a fundamental task in the genome and biomedical research. With the development of RNA-seq and gene chip technology, mass gene expression profile data has been generated, which puts forward two requirements for related research of gene expression profile: i) accurate analysis of drug R&D requires high accuracy of similarity analysis, ii) large-scale analysis of data requires as little running time as possible. We propose a faster, more accurate method called DeepCDNet, which is based on the framework of the Siamese network. DeepCDNet uses the DenseNet structure and optimized loss function to achieve rapid convergence, and the similarity between expression spectra is calculated by a cosine function. The experiment results show that: i) our method breaks through the limitation of high dimensions of gene expression profile and can quickly and accurately learn the required gene characteristics, ii) The accuracy of our method in similarity analysis is greatly improved, iii) as the dimension of data increases, the advantage of our method on time cost gradually becomes more prominent, and time consumption is less.
Shaoliang Peng, Yaning Yang, Fei Li 0040, Hao Hong, Kenli Li 0001, Shulin Wang
HealthCom8
2020 Identification of Human LncRNA-Disease Association by Fast Kernel Learning-Based Kronecker Regularized Least Squares
Wen Li 0009, Shulin Wang, Junlin Xu, Jialiang Yang
ICIC (2)2
2019 A Selection Method for Denoising Auto Encoder Features Using Cross Entropy
Shulin Wang, Jiawei Luo 0001
ICIC (3)4
2019 The Detection of Gene Modules with Overlapping Characteristic via Integrating Multi-omics Data in Six Cancers
Xinguo Lu, Qiumai Miao, Zhenghao Zhu, Shulin Wang
ICIC (2)7
2018 Feature selection in machine learning: A new perspective
Jiawei Luo 0001, Shulin Wang
Neurocomputing3
2017 Feature Selection Based on Density Peak Clustering Using Information Distance Measure
Shilong Chao, Shulin Wang, Jiawei Luo 0001
ICIC (2)4
2016 A Clustering Based Feature Selection Method Using Feature Information Distance for Text Data
Shilong Chao, Shulin Wang
ICIC (1)4
2016 Identifying miRNA-mRNA Regulatory Modules Based on Overlapping Neighborhood Expansion from Multiple Types of Genomic Data
Jiawei Luo 0001, Buwen Cao, Shulin Wang
ICIC (1)4
2016 Detecting overlapping protein complexes in weighted protein-protein interaction networks using pseudo-clique extension based on fuzzy relation
abstract
Detecting overlapping protein complexes in protein-protein interaction (PPI) networks can provide insight into cellular functional organization and thus elucidate underlying cellular mechanisms. Recently, various algorithms for protein complex detection have been developed for PPI networks. However, the majority of algorithms primarily depend on network topological features and/or gene expression profile, failing to consider the inherent biological meanings between protein pairs. In this paper, we propose a method of pseudo-clique extension based on fuzzy relation (PCE-FR) that detects protein complexes from PPI networks weighted with the biological significance hidden in protein pairs. The proposed algorithm operates in three stages: it first forms the non-overlapping protein sub-structure based on fuzzy relation and then expands each sub-structure by adding neighbor proteins to maximize the cohesive score. Finally, highly overlapped candidate protein complexes are merged to form the final protein complex set. We apply PCE-FR to two yeast PPI networks and a human PPI network and validate our results by using CYC2008 and CHPC2012, respectively. Experimental results show that our method outperforms classical algorithms such as CFinder, ClusterONE, CMC, RRW, HC-PIN and ProRank+, and that it achieves ideal overall performance in terms of Precision, Accuracy, and Separation.
Buwen Cao, Jiawei Luo 0001, Cheng Liang 0001, Shulin Wang
IJCNN4
2015 Spectral Clustering of High-Dimensional Data via k-Nearest Neighbor Based Sparse Representation Coefficients
Shulin Wang, Jianwen Fang
ICIC (3)2
2015 Clustering High-Dimensional Data via Spectral Clustering Using Collaborative Representation Coefficients
Shulin Wang, Jinchao Gu
ICIC (2)1
2015 Spectral clustering of high-dimensional data via Nonnegative Matrix Factorization
abstract
Spectral clustering has become a popular subspace clustering algorithm in machine learning and data mining, which aims at finding a low-dimensional representation by utilizing the spectrum of a Laplacian matrix. It is a key to construct a discriminative and reliable affinity matrix for spectral clustering to achieve impressive clustering quality. As the real word data increase with higher dimension of features and larger number of data samples, it is a challenge to construct a good affinity matrix. Recently, sparse representation based spectral clustering (SRSC) has proven its efficiency for clustering and lead to promising clustering results in high-dimensional data. SRSC constructs affinity matrix by using sparse representation coefficient vectors. However, it is very time consuming. Additionally, the dimension of the sparse coefficient vector is equal to the number of samples, which may make the affinity matrix not discriminative enough. Therefore, it is inefficient to apply SRSC in clustering large scale datasets. To remedy these issues, we propose a new spectral clustering algorithm which constructs affinity matrix via Nonnegative Matrix Factorization (NMF) coefficient vectors. We call our algorithm as NMF based spectral clustering (NMFSC). The dimension of NMF coefficient vector is independent on the number of the samples and significantly smaller than that of sparse coefficient vector. Therefore, the affinity matrix can be constructed via NMF coefficient vector with much lower computational cost. The experimental results on several public gene expression profiling (GEP) datasets demonstrate the advantage of NMF coefficient over sparse representation coefficient and suggest that NMFSC is promising in clustering high-dimensional data.
Shulin Wang, Jianwen Fang
IJCNN1
2012 Protein-Protein Interaction Affinity Prediction Based on Interface Descriptors and Machine Learning
Xueling Li, Xiao-Lai Li, Hong-Qiang Wang, Shulin Wang
ICIC (2)5
2012 Protein-Protein Binding Affinity Prediction Based on an SVR Ensemble
Xueling Li, Xiao-Lai Li, Hong-Qiang Wang, Shulin Wang
ICIC (1)5
2012 Research on Virus Detection Technology Based on Ensemble Neural Network and SVM
Boyun Zhang, Jianping Yin, Shulin Wang
ICIC (3)3
2012 Multiple ant colony algorithm method for selecting tag SNPs
Xiong Li 0002, Wen Zhu, Renfa Li, Shulin Wang
J. Biomed. Informatics5
2011 Gray Scale Potential Theory of Sparse Image
Wensheng Tang, Shaohua Jiang, Shulin Wang
ICIC (1)3
2011 Network Security Situation Assessment Based on Stochastic Game Model
Boyun Zhang, Wensheng Tang, Shulin Wang
ICIC (1)6
2011 Network Security Situation Assessment Based on HMM
Boyun Zhang, Shulin Wang, Dingxing Zhang
ICIC (2)3
2011 Network Security Situation Assessment Based on Hidden Semi-Markov Model
Boyun Zhang, Shulin Wang
ICIC (1)4
2009 Analysis of Enterprise Workflow Solutions
Cui-e Chen, Shulin Wang
ICIC (2)2
2009 A Novel Method to Robust Tumor Classification Based on MACE Filter
Shulin Wang, Yihai Zhu
ICIC (2)1
2008 Neighborhood Rough Set Model Based Gene Selection for Multi-subtype Tumor Classification
Shulin Wang, Xueling Li, Shanwen Zhang
ICIC (1)1
2008 A Novel Hybrid Method of Gene Selection and Its Application on Tumor Classification
Zhu-Hong You, Shulin Wang, Jie Gui, Shanwen Zhang
ICIC (2)2
2008 Palmprint Linear Feature Extraction and Identification Based on Ridgelet Transforms and Rough Sets
Shanwen Zhang, Shulin Wang, Xueling Li
ICIC (2)2
2007 Malicious Codes Detection Based on Ensemble Learning
Boyun Zhang, Jianping Yin, Jingbo Hao, Dingxing Zhang, Shulin Wang
ATC5
2007 Feature Extraction and Classification of Tumor Based on Wavelet Package and Support Vector Machines
Shulin Wang, Ji Wang 0001, Huowang Chen, Shutao Li 0001
PAKDD1
2006 SVM-Based Tumor Classification with Gene Expression Data
Shulin Wang, Ji Wang 0001, Huowang Chen, Boyun Zhang
ADMA1
1995 Research and design of a fuzzy neural expert system
Shulin Wang
J. Comput. Sci. Technol.2
1990 An Integrated Framework of Inducing Rules from Examples
Yihua Wu, Shulin Wang
ML2