EDBT 2026 Demo / reviewers in the wild / expert
Shunfang Wang
dblp:157/2950
· DBLP profile ↗
28ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-1927-8753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 20 since 2021Security and privacy · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Few-Sample GNSS Signal LOS/NLOS Identification via Physical-Semantic Distillation From Large Language ModelsabstractGlobal Navigation Satellite Systems (GNSS) are widely used for positioning and navigation, yet NLOS signals severely degrade performance in urban environments. Existing deep learning-based LOS/NLOS identification methods require large labeled datasets and lack physical interpretability. To address this issue, we propose a few-sample GNSS signal LOS/NLOS identification framework based on physical-semantic distillation from large language models (LLMs), where physics-informed reasoning generates soft NLOS labels to guide a lightweight student model. Experiments show a 4.5% accuracy improvement on a public dataset with 1% training data, and 2.4% gain on a self-collected dataset under the same setting. Shunfang Wang, Ming Huang 0003 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Style-Aware Blending and Prototype-Based Cross-Contrast Consistency for Semi-Supervised Medical Image SegmentationabstractWeak-strong consistency learning strategies are widely employed in semi-supervised medical image segmentation to train models by leveraging limited labeled data and enforcing weak-to-strong consistency. However, most existing methods primarily focus on designing and combining various perturbation schemes, overlooking the intrinsic potential and limitations of the framework itself. In this paper, we identify two critical deficiencies: (1) separated training data streams, which lead to confirmation bias dominated by the labeled stream; and (2) incomplete utilization of supervisory signals, which limits exploration of strong-to-weak consistency. To address these challenges, we propose a style-aware blending and prototype-based crosscontrast consistency learning framework. Specifically, inspired by the empirical observation that the distribution mismatch between labeled and unlabeled data can be characterized by their statistical moments, we design a style-guided distribution blending module to bridge the independent training data streams. Meanwhile, considering the potential noise in strong pseudolabels, we introduce a prototype-based cross-contrast strategy to enable the model to learn informative supervisory signals from both weak-to-strong and strong-to-weak predictions, while mitigating the adverse effects of noise. Extensive experiments demonstrate the effectiveness and superiority of our framework across multiple medical image segmentation benchmarks under various semi-supervised settings. The code is available at https://gndlwch2w.github.io/spc-demo. Chaowei Chen, Xiang Zhang 0037, Honglie Guo, Shunfang Wang |
BIBM | 4 |
| 2025 | IGCLAPS: an interpretable graph contrastive learning method with adaptive positive sampling for scRNA-seq data analysisabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) technology enables biological research at single-cell resolution. Cell clustering is a crucial task in scRNA-seq data analysis since it provides insights into cell heterogeneity. Although existing methods have made significant progress in this task, it remains challenging to fully utilize the relationship among cells. RESULTS: We propose Interpretable Graph Contrastive Learning method with Adaptive Positive Sampling (IGCLAPS), a novel end-to-end graph contrastive clustering method for scRNA-seq data analysis. Specifically, IGCLAPS learns low-dimensional embeddings with a graph transformer, based on which a dual-head graph contrastive learning module is used to perform dimension reduction and cell clustering simultaneously. Besides, an accurate definition of positive sample pairs is crucial in contrastive learning, we devise an adaptive positive sampling module, which dynamically identifies true positive sample pairs based on both expression similarity and soft cluster labels generated by the contrastive learning module. Extensive experiments on a series of real datasets including cell clustering, visualization, and differential expression analysis demonstrate that IGCLAPS can effectively enhance clustering performance and generate interpretable gene expression patterns of scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: The source codes of IGCLAPS are available at https://github.com/ZhengWeihuaYNU/IGCLAPS. Wenwen Min, Shunfang Wang |
Bioinform. | 3 |
| 2025 | Multi-Omics Correlation Reconstruction of Complete Graph Forms Based on the Self-Expressive Learning Network for Cancer Subtype PredictionabstractMulti-omics cancer subtype prediction can identify cancer subtypes effectively and has the advantage of correlating genotype and phenotype. However, insufficient exploration of the correlation of information among different omics levels may lead to poor prediction of cancer subtypes. To address this issue, we propose a novel framework, termed Multi-Omics correlation reconstruction (MOCR), which performs reconstruction of complete graph forms based on a self-expressive learning network. Specifically, MOCR first employs autoencoders to unify the dimensions across different omics types. It then leverages parallel query and key networks (QKNets) to learn representations for each omics. These representations are passed into a correlation reconstruction module (CRModule), which computes self-expressive coefficients that jointly capture omics-self characteristics and inter-omics relationships. QKNets and CRModule form a correlative self-expressive learning network, enabling better utilization of the advantages of multi-omics. Importantly, the CRModule's complete graph reconstruction of omics correlations models each omics pair exactly once, thereby avoiding redundancy. Finally, spectral clustering is applied to derive cancer subtypes. We have evaluated our method on nine TCGA cancer datasets and three simulation datasets. The results showed that the MOCR had significant advantages in cancer subtype identification. Junran Zhao, Yueyi Cai, Shunfang Wang |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | scDVAE:Single-Cell Data Clustering Based on Variational Autoencoder With Disentangled Latent RepresentationsabstractSingle-cell RNA sequencing (scRNA-seq) technology enables the analysis of gene expression in individual cells, allowing for a deeper exploration of heterogeneity in organisms and complex diseases. Cell clustering is a crucial step in single-cell analysis, enabling the identification of cellular heterogeneity. However, the high dimensionality, sparsity, and dropout events in single-cell data have brought enormous challenges to clustering analysis. Building on the proven success of deep generative models in learning meaningful representations from low-dimensional latent spaces, we introduce scDVAE, a novel deep generative approach that leverages a variational autoencoder with disentangled latent representations for single-cell clustering. First, each latent representation generated by the encoder is disentangled into clustering features and generative features. In this way, the clustering features can enhance the performance of the clustering task without interference from the generative task. Second, we employ a Student's t-mixture model as the prior distribution for the clustering features to enhance the robustness of our method against dropout events. In addition, we introduce a hybrid data augmentation strategy to generate augmented scRNA-seq data, which enhances dataset diversity while also helping to reduce noise. Our experimental studies on 10 real-world datasets demonstrate that scDVAE significantly improves clustering performance compared to state-of-the-art methods. Xiaohan Zou, Shunfang Wang |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Interval-Shared Information Integration and False-Negative Association Reduction in Multi-Source MiRNA-Disease Association PredictionabstractNumerous studies have demonstrated that microRNAs (miRNAs) play crucial roles in the development and progression of various diseases, making the identification of miRNA-disease association (MDA) essential for understanding human disease etiology. While several computational models have been developed to predict MDAs, challenges persist-particularly the limited consideration of information interactions among multi-source similarities and the presence of "false-negative" associations in the original topology. To address these issues, we propose ISFNMDA, a model designed to infer potential MDAs by leveraging multi-view collaborative learning for feature extraction and optimizing association topology through graph structure momentum contrastive learning. Specifically, multi-source similarities of miRNAs and diseases are mapped into a unified feature space via encoders. The Pearson correlation coefficient is employed to derive pairwise constraints between nodes, facilitating information interactions and constructing interval-shared information constraints. Subsequently, an inference graph learner models the representations to generate an inferred graph topology. By maximizing mutual information between the inferred topology and the original "false-negative" associations through momentum contrastive learning, the model effectively reduces spurious correlations. The final comprehensive representations and optimized graph structure are then used to predict potential MDAs. Experimental results demonstrate that ISFNMDA outperforms existing methods, and case studies further validate its predictive capability. Qinghang Cui, Honglie Guo, Yueyi Cai, Shunfang Wang |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | MSVM-UNet: Multi-Scale Vision Mamba UNet for Medical Image SegmentationabstractState Space Models (SSMs), particularly Mamba, have demonstrated significant potential in medical image segmentation due to their capability to model long-range dependencies with linear computational complexity. However, achieving accurate medical image segmentation necessitates the effective learning of both multi-scale detailed feature representations and global contextual dependencies. Although existing research has attempted to address this challenge by integrating CNNs and SSMs to leverage their respective strengths, they have not developed specialized modules to effectively capture multi-scale feature representations, nor have they sufficiently addressed the directional sensitivity issue when applying Mamba to 2D image data. To overcome these limitations, we propose a Multi-Scale Vision Mamba UNet model for medical image segmentation, termed MSVM-UNet. Specifically, by introducing multi-scale convolutions in the VSS blocks, we can more effectively capture and aggregate multi-scale feature representations from the hierarchical features of the VMamba encoder and better handle 2D visual data. Additionally, the large kernel patch expanding (LKPE) layers achieve more efficient upsampling of feature maps by simultaneously integrating spatial and channel information. Extensive experiments on the Synapse and ACDC datasets demonstrate that our approach is more effective than some state-of-the-art methods in capturing and aggregating multi-scale feature representations and modeling long-range dependencies between pixels. Our implementation is available at https://github.com/gndlwch2w/msvm-unet. Chaowei Chen, Shiquan Min, Shunfang Wang |
BIBM | 4 |
| 2024 | Masked adversarial neural network for cell type deconvolution in spatial transcriptomicsabstractAccurately determining cell type composition in disease-relevant tissues is crucial for identifying disease targets. Most existing spatial transcriptomics (ST) technologies cannot achieve single-cell resolution, making it challenging to accurately determine cell types. To address this issue, various deconvolution methods have been developed. Most of these methods use single-cell RNA sequencing (scRNA-seq) data from the same tissue as a reference to infer cell types in ST data spots. However, they often overlook the differences between scRNA-seq and ST data. To overcome this limitation, we propose a Masked Adversarial Neural Network (MACD). MACD employs adversarial learning to align real ST data with simulated ST data generated from scRNA-seq data. By mapping them into a unified latent space, it can minimize the differences between the two types of data. Additionally, MACD uses masking techniques to effectively learn the features of real ST data and mitigate noise. We evaluated MACD on 32 simulated datasets, demonstrating its accuracy in performing cell type deconvolution. All code and public datasets used in this paper are available at https://github.com/wenwenmin/MACD. Shunfang Wang, Wenwen Min |
BIBM | 3 |
| 2024 | Masked Conditional Diffusion Model with GNN for Spatial Transcriptomics Data ImputationabstractSpatially resolved transcriptomics represents a significant advancement in single-cell analysis by offering both gene expression data and their corresponding physical locations. However, this high degree of spatial resolution entails a drawback, as the resulting spatial transcriptomic data at the cellular level is notably plagued by a high incidence of missing values. Furthermore, most existing imputation methods either overlook the spatial information between spots or compromise the overall gene expression data distribution. To address these challenges, our primary focus is on effectively utilizing the spatial location information within spatial transcriptomic data to impute missing values, while preserving the overall data distribution. We introduce stMCDI, a masked conditional diffusion model for spatial transcriptomics data imputation, which employs a denoising network trained using randomly masked data portions as guidance, with the unmasked data serving as conditions. Additionally, it utilizes a GNN encoder to integrate the spatial position information, thereby enhancing model performance. Compared with baseline methods, our model achieves state-of-the-art performance in all evaluation metrics on six real-world datasets. The results obtained from spatial transcriptomics datasets elucidate the performance of our methods relative to existing approaches. Our code can be accessed at https://github.com/wenwenmin/stMCDI. Wenwen Min, Shunfang Wang, Changmiao Wang, Taosheng Xu |
BIBM | 3 |
| 2024 | scASDC: Attention Enhanced Structural Deep Clustering for Single-cell RNA-seq DataabstractSingle-cell RNA sequencing (scRNA-seq) data analysis is pivotal for understanding cellular heterogeneity. However, the high sparsity and complex noise patterns inherent in scRNA-seq data present significant challenges for traditional clustering methods. To address these issues, we propose a deep clustering method, Attention-Enhanced Structural Deep Embedding Graph Clustering (scASDC), which integrates multiple advanced modules to improve clustering accuracy and robustness. Our approach employs a multi-layer graph convolutional network (GCN) to capture high-order structural relationships between cells, termed as the graph autoencoder module. We introduce a ZINB-based autoencoder module that extracts content information from the data and learns latent representations of gene expression. These modules are further integrated through an attention fusion mechanism, ensuring effective combination of gene expression and structural information at each layer of the GCN. Additionally, a self-supervised learning module is incorporated to enhance the robustness of the learned embeddings. Extensive experiments demonstrate that scASDC outperforms existing state-of-the-art methods, providing a robust and effective solution for single-cell clustering tasks. All code and public datasets used in this paper are available at https://github.com/wenwenmin/scASDC. Wenwen Min, Taosheng Xu, Guangsheng Wu, Shunfang Wang |
BIBM | 6 |
| 2024 | Polymorphic multi-head attention aggregation network for skin lesion segmentationabstractSkin cancer is one of the most common malignant tumors, and accurately segmenting the lesion area from dermoscopic images is of great clinical significance. With the rapid development of computer technology, Transformer-based models have dominated the field of automatic skin lesion segmentation. However, Transformer-based models typically focus on capturing global dependencies, lacking explicitly encoded convolutional layers as in RNNs or CNNs. This paper proposes a polymorphic multi-head attention aggregation network for skin lesion segmentation (PMAA-Net). It establishes a polymorphic multi-head attention mechanism (PMA), where each self-attention head is designed with convolutional layers that can encode gradients and textures to obtain various local features. Meanwhile, two attention heads are designed to capture spatial and channel-wise features respectively. This allows the model to capture local features and spatial-dimensional features. We further introduce a multi-head attention aggregation module (AGG) to aggregate multiple attention heads. This transforms the attention maps into a high-rank mixture distribution, significantly enhancing the feature representation capability of the attention heads. Our experimental results on three publicly available skin lesion segmentation datasets show that PMAA-Net outperforms other mainstream methods, especially those focusing on extracting local structural features. The codes are available at https://github.com/yuli0501/PMAA-Net. Chaowei Chen, Wenwen Min, Che Zhao, Shunfang Wang |
BIBM | 5 |
| 2024 | Learning an Adaptive Self-expressive Fusion Model for Multi-omics Cancer Subtype Prediction
Yueyi Cai, Junran Zhao, Shunfang Wang |
ISBRA (1) | 4 |
| 2024 | Deeply integrating latent consistent representations in high-noise multi-omics data for cancer subtypingabstractCancer is a complex and high-mortality disease regulated by multiple factors. Accurate cancer subtyping is crucial for formulating personalized treatment plans and improving patient survival rates. The underlying mechanisms that drive cancer progression can be comprehensively understood by analyzing multi-omics data. However, the high noise levels in omics data often pose challenges in capturing consistent representations and adequately integrating their information. This paper proposed a novel variational autoencoder-based deep learning model, named Deeply Integrating Latent Consistent Representations (DILCR). Firstly, multiple independent variational autoencoders and contrastive loss functions were designed to separate noise from omics data and capture latent consistent representations. Subsequently, an Attention Deep Integration Network was proposed to integrate consistent representations across different omics levels effectively. Additionally, we introduced the Improved Deep Embedded Clustering algorithm to make integrated variable clustering friendly. The effectiveness of DILCR was evaluated using 10 typical cancer datasets from The Cancer Genome Atlas and compared with 14 state-of-the-art integration methods. The results demonstrated that DILCR effectively captures the consistent representations in omics data and outperforms other integration methods in cancer subtyping. In the Kidney Renal Clear Cell Carcinoma case study, cancer subtypes were identified by DILCR with significant biological significance and interpretability. Yueyi Cai, Shunfang Wang |
Briefings Bioinform. | 2 |
| 2024 | Hilbert signal envelope-based multi-features methods for GNSS spoofing detection
Zixiao Peng, Shunfang Wang, Ming Huang 0003 |
Comput. Secur. | 4 |
| 2024 | One-step multi-view clustering guided by weakened view-specific distribution
Yueyi Cai, Shunfang Wang |
Expert Syst. Appl. | 2 |
| 2024 | Boundary-Aware Gradient Operator Network for Medical Image SegmentationabstractMedical image segmentation is a crucial task in computer-aided diagnosis. Although convolutional neural networks (CNNs) have made significant progress in the field of medical image segmentation, the convolution kernels of CNNs are optimized from random initialization without explicitly encoding gradient information, leading to a lack of specificity for certain features, such as blurred boundary features. Furthermore, the frequently applied down-sampling operation also loses the fine structural features in shallow layers. Therefore, we propose a boundary-aware gradient operator network (BG-Net) for medical image segmentation, in which the gradient convolution (GConv) and the boundary-aware mechanism (BAM) modules are developed to simulate image boundary features and the remote dependencies between channels. The GConv module transforms the gradient operator into a convolutional operation that can extract gradient features; it attempts to extract more features such as images boundaries and textures, thereby fully utilizing limited input to capture more features representing boundaries. In addition, the BAM can increase the amount of global contextual information while suppressing invalid information by focusing on feature dependencies and the weight ratios between channels. Thus, the boundary perception ability of BG-Net is improved. Finally, we use a multi-modal fusion mechanism to effectively fuse lightweight gradient convolution and U-shaped branch features into a multilevel feature, enabling global dependencies and low-level spatial details to be effectively captured in a shallower manner. We conduct extensive experiments on eight datasets that broadly cover medical images to evaluate the effectiveness of the proposed BG-Net. The experimental results demonstrate that BG-Net outperforms the state-of-the-art methods, particularly those focused on boundary segmentation. Wenwen Min, Shunfang Wang |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | TransVCOX: Bridging Transformer Encoder and Pre-trained VAE for Robust Cancer Multi-Omics Survival AnalysisabstractTraditional survival analysis models, such as the COX proportional hazards model, face challenges in processing multimodal data, identifying nonlinear relationships, and recognizing complex data patterns. The rise of deep learning, particularly Transformers and variational autoencoders (VAEs), has showcased its potential in analyzing cancer multi-omics data comprehensively. However, many individual cancer datasets suffer from limited sample sizes, preventing some deep learning models from extracting in-depth data representations and resulting in subpar performance. To address this issue, we advocate the adoption of pre-training and fine-tuning techniques, which effectively mitigate performance deficits due to sparse cancer data samples. We introduce TransVCOX, a deep survival analysis model integrating a Transformer encoder with VAE. This model leverages pre-training and fine-tuning approaches to predict patients’ survival risk using cancer multi-omics data. Rigorous tests on eight unique cancer datasets from TCGA revealed: (1) TransVCOX outperforms other deep learning and conventional COX models. (2) Pre-training significantly reduces model overfitting and enhances performance. (3) VAE encoding, compared to positional encoding, offers a richer decision-making foundation. (4) The performance boost doesn’t linearly correlate with the addition of Transformer blocks. These findings underline TransVCOX’s promising capability for predicting cancer patients’ survival risks using multi-omics data. The implementation can be accessed at https://github.com/wenwenmin/TransVCOX. Wenwen Min, Shunfang Wang |
BIBM | 5 |
| 2023 | BFP-Net: Boundary Feature Pyramid for Medical Image SegmentationabstractIn this paper, we propose a novel method, namely boundary feature pyramid network (BFP-Net), which can effectively segment targets with blurred boundaries. Specifically, we first propose a global feature fusion module (GFFM) at the top of BFP-Net to fuse feature maps within different scales. It can learn more global region localization of the target and help the learning of boundaries more effectively. Then, we propose a series of boundary enhancement modules (BEMs) at the decoder to effectively extract and integrate boundary information during the upsampling process, thereby enhancing the ability to capture fine details (such as the boundaries). Furthermore, we introduce a boundary-enhanced composite loss function to effectively segment both the regions and their boundaries within different scales. Finally, extensive experiments on two widely-used datasets demonstrate that BFP-Net is more effective in fusing contextual information and guiding feature map boundaries compared to previous competitive methods. Shiquan Min, Xiang Zhang 0037, Shunfang Wang |
BIBM | 3 |
| 2023 | TsImpute: an accurate two-step imputation method for single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) technology has enabled discovering gene expression patterns at single cell resolution. However, due to technical limitations, there are usually excessive zeros, called "dropouts," in scRNA-seq data, which may mislead the downstream analysis. Therefore, it is crucial to impute these dropouts to recover the biological information. RESULTS: We propose a two-step imputation method called tsImpute to impute scRNA-seq data. At the first step, tsImpute adopts zero-inflated negative binomial distribution to discriminate dropouts from true zeros and performs initial imputation by calculating the expected expression level. At the second step, it conducts clustering with this modified expression matrix, based on which the final distance weighted imputation is performed. Numerical results based on both simulated and real data show that tsImpute achieves favorable performance in terms of gene expression recovery, cell clustering, and differential expression analysis. AVAILABILITY AND IMPLEMENTATION: The R package of tsImpute is available at https://github.com/ZhengWeihuaYNU/tsImpute. Wenwen Min, Shunfang Wang |
Bioinform. | 3 |
| 2021 | IMAL: An Improved Meta-learning Approach for Few-shot Classification of Plant DiseasesabstractThe timely identification of plant diseases is crucial for the production of crops. For this problem, many excellent and state-of-the-art algorithms based on deep learning have emerged currently. However, these algorithms still have problems such as poor generalization, difficulty in learning and adapting to new tasks, and extreme reliance on large-scale data. This study introduces an improved meta-learning approach(IMAL) for the few-shot classification of plant diseases, which can produce good generalization performance on new tasks with only a small amount of data and several steps of gradient update. In IMAL, the model-agnostic meta-learning approach with strong generalization capability is used as the overall algorithm framework, a fresh loss function called soft-center loss is adopted to conquer the problem of the poor distinguishing ability of the softmax classifier for features, and the Parametric Rectified Linear Unit (PReLU) activation function is utilized to enhance the model fitting ability with negligible additional computational cost and overfitting risk. The experiment results of plant diseases identification confirmed that the proposed IMAL approach is superior to many current few-shot learning approaches. Yingtao Wang, Shunfang Wang |
BIBE | 2 |
| 2021 | Parameter Transfer Learning Measured by Image Similarity to Detect CT of COVID-19
Shunfang Wang |
ISBRA | 2 |
| 2021 | Predicting antifreeze proteins with weighted generalized dipeptide composition and multi-regression feature selection ensembleabstractBACKGROUND: Antifreeze proteins (AFPs) are a group of proteins that inhibit body fluids from growing to ice crystals and thus improve biological antifreeze ability. It is vital to the survival of living organisms in extremely cold environments. However, little research is performed on sequences feature extraction and selection for antifreeze proteins classification in the structure and function prediction, which is of great significance. RESULTS: In this paper, to predict the antifreeze proteins, a feature representation of weighted generalized dipeptide composition (W-GDipC) and an ensemble feature selection based on two-stage and multi-regression method (LRMR-Ri) are proposed. Specifically, four feature selection algorithms: Lasso regression, Ridge regression, Maximal information coefficient and Relief are used to select the feature sets, respectively, which is the first stage of LRMR-Ri method. If there exists a common feature subset among the above four sets, it is the optimal subset; otherwise we use Ridge regression to select the optimal subset from the public set pooled by the four sets, which is the second stage of LRMR-Ri. The LRMR-Ri method combined with W-GDipC was performed both on the antifreeze proteins dataset (binary classification), and on the membrane protein dataset (multiple classification). Experimental results show that this method has good performance in support vector machine (SVM), decision tree (DT) and stochastic gradient descent (SGD). The values of ACC, RE and MCC of LRMR-Ri and W-GDipC with antifreeze proteins dataset and SVM classifier have reached as high as 95.56%, 97.06% and 0.9105, respectively, much higher than those of each single method: Lasso, Ridge, Mic and Relief, nearly 13% higher than single Lasso for ACC. CONCLUSION: The experimental results show that the proposed LRMR-Ri and W-GDipC method can significantly improve the accuracy of antifreeze proteins prediction compared with other similar single feature methods. In addition, our method has also achieved good results in the classification and prediction of membrane proteins, which verifies its widely reliability to a certain extent. Shunfang Wang, Xinnan Xia, Zicheng Cao |
BMC Bioinform. | 1 |
| 2021 | Gene prediction of aging-related diseases based on DNN and MashupabstractBACKGROUND: At present, the bioinformatics research on the relationship between aging-related diseases and genes is mainly through the establishment of a machine learning multi-label model to classify each gene. Most of the existing methods for predicting pathogenic genes mainly rely on specific types of gene features, or directly encode multiple features with different dimensions, use the same encoder to concatenate and predict the final results, which will be subject to many limitations in the applicability of the algorithm. Possible shortcomings of the above include: incomplete coverage of gene features by a single type of biomics data, overfitting of small dimensional datasets by a single encoder, or underfitting of larger dimensional datasets. METHODS: We use the known gene disease association data and gene descriptors, such as gene ontology terms (GO), protein interaction data (PPI), PathDIP, Kyoto Encyclopedia of genes and genomes Genes (KEGG), etc, as input for deep learning to predict the association between genes and diseases. Our innovation is to use Mashup algorithm to reduce the dimensionality of PPI, GO and other large biological networks, and add new pathway data in KEGG database, and then combine a variety of biological information sources through modular Deep Neural Network (DNN) to predict the genes related to aging diseases. RESULT AND CONCLUSION: The results show that our algorithm is more effective than the standard neural network algorithm (the Area Under the ROC curve from 0.8795 to 0.9153), gradient enhanced tree classifier and logistic regression classifier. In this paper, we firstly use DNN to learn the similar genes associated with the known diseases from the complex multi-dimensional feature space, and then provide the evidence that the assumed genes are associated with a certain disease. Junhua Ye, Shunfang Wang, Xianjun Tang |
BMC Bioinform. | 2 |
| 2020 | Multi-feature fusion and dimensional reduction based on the two-step deep ontology and the conjoint triad for the identification of cancerlectinsabstractCancerlectins play an important role in the differentiation and the transfer of tumor cell. The accurate identification of cancerlectins is helpful to clarify the development direction of cancer treatment. This paper proposes a feature expression algorithm which combines multi-feature information to distinguish cancerlectins from non-cancerlectins. The main work is as follows. First, develops a two-step feature extraction algorithm based on multi-information fusion, in which the first step is to use DeepGo algorithm to extract the feature of protein and get the GO annotation of each sample, and the second step is to extract the GO term vectors for the given protein sequence. The experimental results show that this method can provide more protein information and identify lectins better. Second, due to the high dimensionality of the feature vector after fusion, this paper firstly used linear discriminant analysis algorithm to reduce its dimension and improved the prediction accuracy to 87.57%. Third, combining with protein physicochemical information, and fusing the conjoint triad with the extracted features from the above two-step method, we obtained a new type of feature vector with multiple information, CTD+DGG, which furtherly improved the prediction accuracy by 1.95% based on the previous methods. Shunfang Wang |
BIBM | 1 |
| 2020 | G-DipC: An Improved Feature Representation Method for Short Sequences to Predict the Type of Cargo in Cell-Penetrating PeptidesabstractCell-penetrating peptides (CPPs) are functional short peptides with high carrying capacity. CPP sequences with targeting functions for the highly efficient delivery of drugs to target cells. In this paper, which is focused on the prediction of the cargo category of CPPs, a biocomputational model is constructed to efficiently distinguish the category of cargo carried by CPPs as macromolecular carriers among the seven known deliverable cargo categories. Based on dipeptide composition (DipC), an improved feature representation method, general dipeptide composition (G-DipC) is proposed for short peptide sequences and can effectively increase the abundance of features represented. Then linear discriminant analysis (LDA) is applied to mine some important low-dimensional features of G-DipC and a predictive model is built with the XGBoost algorithm. Experimental results with five-fold cross validation show that G-DipC improves accuracy by 25 and 5 percent compared with amino acid composition (AAC) and DipC, respectively. G-DipC is even found to be better than tripeptide composition (TipC). Thus, the proposed model provides a novel resource for the study of cell-penetrating peptides, and the improved dipeptide composition G-DipC can be widely adapted to determine the feature representation of other biological sequences. Shunfang Wang, Zicheng Cao, Yaoting Yue |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Accurate classification of membrane protein types based on sequence and evolutionary information using deep learningabstractBACKGROUND: Membrane proteins play an important role in the life activities of organisms. Knowing membrane protein types provides clues for understanding the structure and function of proteins. Though various computational methods for predicting membrane protein types have been developed, the results still do not meet the expectations of researchers. RESULTS: We propose two deep learning models to process sequence information and evolutionary information, respectively. Both models obtained better results than traditional machine learning models. Furthermore, to improve the performance of the sequence information model, we also provide a new vector representation method to replace the one-hot encoding, whose overall success rate improved by 3.81% and 6.55% on two datasets. Finally, a more effective model is obtained by fusing the above two models, whose overall success rate reached 95.68% and 92.98% on two datasets. CONCLUSION: The final experimental results show that our method is more effective than existing methods for predicting membrane protein types, which can help laboratory researchers to identify the type of novel membrane proteins. Shunfang Wang, Zicheng Cao |
BMC Bioinform. | 2 |
| 2019 | Prediction of protein structural classes by different feature expressions based on 2-D wavelet denoising and fusionabstractBACKGROUND: Protein structural class predicting is a heavily researched subject in bioinformatics that plays a vital role in protein functional analysis, protein folding recognition, rational drug design and other related fields. However, when traditional feature expression methods are adopted, the features usually contain considerable redundant information, which leads to a very low recognition rate of protein structural classes. RESULTS: We constructed a prediction model based on wavelet denoising using different feature expression methods. A new fusion idea, first fuse and then denoise, is proposed in this article. Two types of pseudo amino acid compositions are utilized to distill feature vectors. Then, a two-dimensional (2-D) wavelet denoising algorithm is used to remove the redundant information from two extracted feature vectors. The two feature vectors based on parallel 2-D wavelet denoising are fused, which is known as PWD-FU-PseAAC. The related source codes are available at https://github.com/Xiaoheng-Wang12/Wang-xiaoheng/tree/master. CONCLUSIONS: Experimental verification of three low-similarity datasets suggests that the proposed model achieves notably good results as regarding the prediction of protein structural classes. Shunfang Wang, Xiaoheng Wang |
BMC Bioinform. | 1 |
| 2014 | Improved 2DLDA Algorithm and Its Application in Face RecognitionabstractThe face recognition method based on linear discriminant analysis (LDA) is always faced with the high-dimensional small sample size problem in image processing. The two-dimensional linear discriminant analysis (2DLDA) which was proposed recently can extract the feature from the original image matrix directly and decrease the dimensionality of the original matrix in a great extent. But there are still some problems such as the overlap of the neighbor samples and the deviation of the class-center. This paper improves the 2DLDA algorithm and proposes the MF2DLDA algorithm. The MF2DLDA algorithm improves the recognition rate by redefining the between-class scatter matrix which can retains the most discriminative features and using the median matrix instead of the mean matrix which can weaken the negative effect that the anomalous data has on the computing of the median matrix. The results of the experiments show that the new algorithm is feasible. Shunfang Wang |
TrustCom | 2 |