VLDB 2026 Research / reviewers in the wild / expert
Yanbin Yin
dblp:30/7005
· DBLP profile ↗
26ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0001-7667-881XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 15 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language ModelsabstractYanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang, Yifei Shao, Shibo Hao, Yi Gu, Jieyuan Liu, Somanshu Singla, Tianyang Liu, Eric P. Xing, Zhengzhong Liu, Haojian Jin, Zhiting Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yanbin Yin, Kun Zhou 0002, Zhen Wang 0041, Yifei Shao, Shibo Hao, Yi Gu 0002, Jieyuan Liu, Somanshu Singla, Tianyang Liu 0003, Eric P. Xing, Zhengzhong Liu 0001, Haojian Jin, Zhiting Hu |
ACL (1) | 1 |
| 2025 | Neuron based Personality Trait Induction in Large Language ModelsabstractLarge language models (LLMs) have become increasingly proficient at simulating various personality traits, an important capability for supporting related applications (e.g., role-playing). To further improve this capacity, in this paper, we present a neuron based approach for personality trait induction in LLMs, with three major technical contributions. First, we construct PERSONALITYBENCH, a large-scale dataset for identifying and evaluating personality traits in LLMs. This dataset is grounded in the Big Five personality traits from psychology and designed to assess the generative capabilities of LLMs towards specific personality traits. Second, by leveraging PERSONALITYBENCH, we propose an efficient method for identifying personality-related neurons within LLMs by examining the opposite aspects of a given trait. Third, we develop a simple yet effective induction method that manipulates the values of these identified personality-related neurons, which enables fine-grained control over the traits exhibited by LLMs without training and modifying model parameters. Extensive experiments validates the efficacy of our neuron identification and trait induction methods. Notably, our approach achieves comparable performance as fine-tuned models, offering a more efficient and flexible solution for personality trait induction in LLMs. Yanbin Yin, Wayne Xin Zhao, Ji-Rong Wen |
ICLR | 3 |
| 2024 | TMGE: a multi-view graph embedding model for prediction task of genes associated with neurodegenerative diseasesabstractPredicting genes associated with diseases is of great significance for the early diagnosis and efficient treatment of neurodegenerative diseases. However, with the increasing complexity of omics data, it is difficult to integrate many aspects of relationships and effectively mine the genes associated with neurodegenerative diseases. To address this problem, we propose a novel multi-view graph embedding model, named TMGE, which can effectively integrate multiple networks to learn gene representations for prediction task of genes associated with neurodegenerative diseases. First, we design view-specific and view-shared encoders to extract view-specific content and capture consistent information across views; this enables the separation and preservation of diverse information within multi-view data for more accurate representations. Secondly, we introduce cross-linked encoders to enhance the integration of multi-view information, enabling the imputation of missing data across views by capturing cross-view dependencies. Furthermore, we employ adversarial learning to reduce the distribution discrepancy between diverse views at the global level, and then optimize the mean square error of local pairs with the help of the cross-linked encoder structure to achieve local alignment. We conduct experiments to verify the effectiveness of TMGE for predicting genes associated with neurodegenerative diseases using two disease datasets of Alzheimer and Huntington. Yuanxiang Jiang, Mingjing Han, Yanbin Yin, Han Zhang 0017 |
BIBM | 3 |
| 2024 | BERMAD: batch effect removal for single-cell RNA-seq data using a multi-layer adaptation autoencoder with dual-channel frameworkabstractMOTIVATION: Removal of batch effect between multiple datasets from different experimental platforms has become an urgent problem, since single-cell RNA sequencing (scRNA-seq) techniques developed rapidly. Although there have been some methods for this problem, most of them still face the challenge of under-correction or over-correction. Specifically, handling batch effect in highly nonlinear scRNA-seq data requires a more powerful model to address under-correction. In the meantime, some previous methods focus too much on removing difference between batches, which may disturb the biological signal heterogeneity of datasets generated from different experiments, thereby leading to over-correction. RESULTS: In this article, we propose a novel multi-layer adaptation autoencoder with dual-channel framework to address the under-correction and over-correction problems in batch effect removal, which is called BERMAD and can achieve better results of scRNA-seq data integration and joint analysis. First, we design a multi-layer adaptation architecture to model distribution difference between batches from different feature granularities. The distribution matching on various layers of autoencoder with different feature dimensions can result in more accurate batch correction outcome. Second, we propose a dual-channel framework, where the deep autoencoder processing each single dataset is independently trained. Hence, the heterogeneous information that is not shared between different batches can be retained more completely, which can alleviate over-correction. Comprehensive experiments on multiple scRNA-seq datasets demonstrate the effectiveness and superiority of our method over the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The code implemented in Python and the data used for experiments have been released on GitHub (https://github.com/zhanglabNKU/BERMAD) and Zenodo (https://zenodo.org/records/10695073) with detailed instructions. Xiangxin Zhan, Yanbin Yin, Han Zhang 0017 |
Bioinform. | 2 |
| 2024 | A Dual-Modality Complex-Valued Fusion Method for Predicting Side Effects of Drug-Drug Interactions Based on Graph Neural NetworkabstractPredicting potential side effects of drug-drug interactions (DDIs), which is a major concern in clinical treatment, can increase therapeutic efficacy. In recent studies, how to use the multi-modal drug features is important for DDI prediction. Thus, it remains a challenge to explore an efficient computational method to achieve the feature fusion cross- and intra-modality. In this paper, we propose a dual-modality complex-valued fusion method (DMCF-DDI) for predicting the side effects of DDIs, using the form and properties of complex-vector to enhance the representations of DDIs. Firstly, DMCF-DDI applies two Graph Convolutional Network (GCN) encoders to learn molecular structure and topological features from fingerprint and knowledge graphs, respectively. Secondly, an asymmetric skip connection (ASC) uses distinct semantic-level features to construct the complex-valued drug pair representations (DPRs). Then, the complex-vector multiplication is used as a fusion operator to obtain the fine-grained DPRs. Finally, we calculate the prediction probability of DDIs by Hermitian inner product in the complex space. Compared with other methods, DMCF-DDI achieves superior performance in all situations using a fusion operator with the lowest parameter numbers. For the case study, we select six diseases and common side effects in clinical treatment to verify identification ability of our model. We also prove the advantage of ASC and complex-valued fusion can achieve to align the cross-modal fused positive DPRs through a comprehensive analysis on the phase-modulus distribution histogram of DPRs. In the end, we explain the reason for alignment based on the similarity of features and node neighbors. Chuanze Kang, Han Zhang 0017, Yanbin Yin |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Hierarchical Semantic Augmentation Graph Neural Network for Drug-Disease Association PredictionabstractAs an essential step in drug intervention discovery, predicting the drug-disease associations (DDAs) explores the potential therapeutic associations in given dugs and diseases. Since the various links in drugs and diseases contain high-order relations and complex therapeutic semantics, Graph Neural Networks (GNNs) have been introduced to DDA predictions and achieved great success. However, most previous approaches require the nodes of given drugs and diseases to have smooth attributes, which is difficult to meet in practical applications. Besides, GNN-based models suffer from the problem of semantic confusion for DDA prediction in heterogeneous graph. These challenges limit the model validity to discover therapeutic semantics in drug-disease networks. To address these challenges in DDA, we propose a novel graph neural network model called HSAGNN to augment node semantics hierarchically with three key steps by applying semantic-guided idea of SGNN method, including topological embedding learning, attribute completion, and semantic-guided aggregation. HSAGNN first learns the topological embedding and adopts the learned topological relationships to complete missing attributes with attention mechanism, which allows the node to contain richer information for neighbor aggregation. Then, the model aggregates the neighbor information with semantic-guided aggregation in both node and semantic levels. Here, HSAGNN injects the learned common knowledge as jumping knowledge to alleviate the semantic confusion. We evaluate the model in DDA tasks with various baselines and explore the model validity with extensive studies. The experimental results show that HSAGNN can discover the potential therapeutic associations by augmented semantics. Mingjing Han, Yanbin Yin, Han Zhang 0017 |
BIBM | 2 |
| 2023 | Unsupervised Feature Selection by Fusing Spectral Clustering and Locality Preserving ProjectionabstractDue to feature redundancy in high dimensional data, the unsupervised feature selection methods for dimension reduction have attracted considerable attention. The current feature selection frameworks consider the global information of data, but ignore mutual screening of global and local information, and there have been no important breakthroughs on this approach research recently. We propose a novel feature selection method based on iterative optimization between the pseudo label matrix from spectral clustering and the local projection information (SNUFS), and then prove the convergence of the method. The pseudo label matrix and the local projection are designed in objective function for mutual screening and guiding the regularization feature selection by iterative approach. Our method selects features most relevant to the pseudo label and preserves the local structure of original data from feature selection matrix, where alternate iteration of different optimization items including pseudo label matrix achieve mutual screening. For this method, we give the objective function, iterative optimization fusion approach and convergence analysis in detail. Furthermore, we use K-Nearest Neighbor (KNN) and K-means to implement locality preserving projection and obtain two specific algorithms. Experiments on four real-world datasets in different fields demonstrate that our algorithms can effectively improve the accuracy of feature selection. In particular, our algorithm of KNN implementation is more effective and outperforms other major algorithms. Xiongwen Quan, Mingjing Han, Xia Guo, Han Zhang 0017, Yanbin Yin |
BIBM | 6 |
| 2023 | scMHNN: a novel hypergraph neural network for integrative analysis of single-cell epigenomic, transcriptomic and proteomic dataabstractTechnological advances have now made it possible to simultaneously profile the changes of epigenomic, transcriptomic and proteomic at the single cell level, allowing a more unified view of cellular phenotypes and heterogeneities. However, current computational tools for single-cell multi-omics data integration are mainly tailored for bi-modality data, so new tools are urgently needed to integrate tri-modality data with complex associations. To this end, we develop scMHNN to integrate single-cell multi-omics data based on hypergraph neural network. After modeling the complex data associations among various modalities, scMHNN performs message passing process on the multi-omics hypergraph, which can capture the high-order data relationships and integrate the multiple heterogeneous features. Followingly, scMHNN learns discriminative cell representation via a dual-contrastive loss in self-supervised manner. Based on the pretrained hypergraph encoder, we further introduce the pre-training and fine-tuning paradigm, which allows more accurate cell-type annotation with only a small number of labeled cells as reference. Benchmarking results on real and simulated single-cell tri-modality datasets indicate that scMHNN outperforms other competing methods on both cell clustering and cell-type annotation tasks. In addition, we also demonstrate scMHNN facilitates various downstream tasks, such as cell marker detection and enrichment analysis. Wei Li 0184, Bin Xiang, Fan Yang 0081, Yu Rong 0001, Yanbin Yin, Jianhua Yao 0001, Han Zhang 0017 |
Briefings Bioinform. | 5 |
| 2023 | Genome mining for anti-CRISPR operons using machine learningabstractMOTIVATION: Encoded by (pro-)viruses, anti-CRISPR (Acr) proteins inhibit the CRISPR-Cas immune system of their prokaryotic hosts. As a result, Acr proteins can be employed to develop more controllable CRISPR-Cas genome editing tools. Recent studies revealed that known acr genes often coexist with other acr genes and with phage structural genes within the same operon. For example, we found that 47 of 98 known acr genes (or their homologs) co-exist in the same operons. None of the current Acr prediction tools have considered this important genomic context feature. We have developed a new software tool AOminer to facilitate the improved discovery of new Acrs by fully exploiting the genomic context of known acr genes and their homologs. RESULTS: AOminer is the first machine learning based tool focused on the discovery of Acr operons (AOs). A two-state HMM (hidden Markov model) was trained to learn the conserved genomic context of operons that contain known acr genes or their homologs, and the learnt features could distinguish AOs and non-AOs. AOminer allows automated mining for potential AOs from query genomes or operons. AOminer outperformed all existing Acr prediction tools with an accuracy = 0.85. AOminer will facilitate the discovery of novel anti-CRISPR operons. AVAILABILITY AND IMPLEMENTATION: The webserver is available at: http://aca.unl.edu/AOminer/AOminer_APP/. The python program is at: https://github.com/boweny920/AOminer. Minal Khatri, Jinfang Zheng, Jitender S. Deogun, Yanbin Yin |
Bioinform. | 5 |
| 2023 | Semantic-guided graph neural network for heterogeneous graph embedding
Mingjing Han, Han Zhang 0017, Wei Li 0184, Yanbin Yin |
Expert Syst. Appl. | 4 |
| 2022 | Mutual Information Estimation-Based Disentangled Representation Network for Medical Image FusionabstractDeep learning-based method for medical image fusion has become a hot topic in recent years. However, they ignore the expression of the most important features in image fusion and only extract the general features for medical image fusion, which will restrict the expression of unique information on the fusion image. To address this restriction, we propose a novel disentangled representation network for medical image fusion with mutual information estimation, which extract the disentangled features of medical image fusion, i.e., the shared and exclusive features between multi-model medical images. In our method, we use the cross mutual information method to obtain the shared features of each modality pair, which enforce the fusion network to achieve the maximum of mutual information estimation for multi-modal medical images. The exclusive features are extracted by the adversarial objective method and it constrains the fusion network with the optimization to the minimum of mutual information estimation between shared and exclusive features. These disentangled features effectively take the interpretative advantages and make the fusion image retaining more details from source images as well as improving the visual quality of fusion image. Our method has achieved better results than several state-of-the-art methods. Both qualitative and quantitative experiments have proved the superiority of our method. Wanwan Huang, Han Zhang 0017, Yanbin Yin |
BIBM | 4 |
| 2022 | LR-GNN: a graph neural network based on link representation for predicting molecular associationsabstractIn biomedical networks, molecular associations are important to understand biological processes and functions. Many computational methods, such as link prediction methods based on graph neural networks (GNNs), have been successfully applied in discovering molecular relationships with biological significance. However, it remains a challenge to explore a method that relies on representation learning of links for accurately predicting molecular associations. In this paper, we present a novel GNN based on link representation (LR-GNN) to identify potential molecular associations. LR-GNN applies a graph convolutional network (GCN)-encoder to obtain node embedding. To represent associations between molecules, we design a propagation rule that captures the node embedding of each GCN-encoder layer to construct the LR. Furthermore, the LRs of all layers are fused in output by a designed layer-wise fusing rule, which enables LR-GNN to output more accurate results. Experiments on four biomedical network data, including lncRNA-disease association, miRNA-disease association, protein-protein interaction and drug-drug interaction, show that LR-GNN outperforms state-of-the-art methods and achieves robust performance. Case studies are also presented on two datasets to verify the ability to predict unknown associations. Finally, we validate the effectiveness of the LR by visualization. Chuanze Kang, Han Zhang 0017, Shenwei Huang, Yanbin Yin |
Briefings Bioinform. | 5 |
| 2022 | Critical assessment of pan-genomic analysis of metagenome-assembled genomesabstractPan-genome analyses of metagenome-assembled genomes (MAGs) may suffer from the known issues with MAGs: fragmentation, incompleteness and contamination. Here, we conducted a critical assessment of pan-genomics of MAGs, by comparing pan-genome analysis results of complete bacterial genomes and simulated MAGs. We found that incompleteness led to significant core gene (CG) loss. The CG loss remained when using different pan-genome analysis tools (Roary, BPGA, Anvi'o) and when using a mixture of MAGs and complete genomes. Contamination had little effect on core genome size (except for Roary due to in its gene clustering issue) but had major influence on accessory genomes. Importantly, the CG loss was partially alleviated by lowering the CG threshold and using gene prediction algorithms that consider fragmented genes, but to a less degree when incompleteness was higher than 5%. The CG loss also led to incorrect pan-genome functional predictions and inaccurate phylogenetic trees. Our main findings were supported by a study of real MAG-isolate genome data. We conclude that lowering CG threshold and predicting genes in metagenome mode (as Anvi'o does with Prodigal) are necessary in pan-genome analysis of MAGs. Development of new pan-genome analysis tools specifically for MAGs are needed in future studies. Tang Li 0002, Yanbin Yin |
Briefings Bioinform. | 2 |
| 2022 | MGEGFP: a multi-view graph embedding method for gene function prediction based on adaptive estimation with GCNabstractIn recent years, a number of computational approaches have been proposed to effectively integrate multiple heterogeneous biological networks, and have shown impressive performance for inferring gene function. However, the previous methods do not fully represent the critical neighborhood relationship between genes during the feature learning process. Furthermore, it is difficult to accurately estimate the contributions of different views for multi-view integration. In this paper, we propose MGEGFP, a multi-view graph embedding method based on adaptive estimation with Graph Convolutional Network (GCN), to learn high-quality gene representations among multiple interaction networks for function prediction. First, we design a dual-channel GCN encoder to disentangle the view-specific information and the consensus pattern across diverse networks. By the aid of disentangled representations, we develop a multi-gate module to adaptively estimate the contributions of different views during each reconstruction process and make full use of the multiplexity advantages, where a diversity preservation constraint is designed to prevent the over-fitting problem. To validate the effectiveness of our model, we conduct experiments on networks from the STRING database for both yeast and human datasets, and compare the performance with seven state-of-the-art methods in five evaluation metrics. Moreover, the ablation study manifests the important contribution of the designed dual-channel encoder, multi-gate module and the diversity preservation constraint in MGEGFP. The experimental results confirm the superiority of our proposed method and suggest that MGEGFP can be a useful tool for gene function prediction. Wei Li 0184, Han Zhang 0017, Minghe Li, Mingjing Han, Yanbin Yin |
Briefings Bioinform. | 5 |
| 2022 | HDMC: a novel deep learning-based framework for removing batch effects in single-cell RNA-seq dataabstractMOTIVATION: With the development of single-cell RNA sequencing (scRNA-seq) techniques, increasingly more large-scale gene expression datasets become available. However, to analyze datasets produced by different experiments, batch effects among different datasets must be considered. Although several methods have been recently published to remove batch effects in scRNA-seq data, two problems remain to be challenging and not completely solved: (i) how to reduce the distribution differences of different batches more accurately; and (ii) how to align samples from different batches to recover the cell type clusters. RESULTS: We proposed a novel deep-learning approach, which is a hierarchical distribution-matching framework assisted with contrastive learning to address these two problems. Firstly, we design a hierarchical framework for distribution matching based on a deep autoencoder. This framework employs an adversarial training strategy to match the global distribution of different batches. This provides an improved foundation to further match the local distributions with a maximum mean discrepancy-based loss. For local matching, we divide cells in each batch into clusters and develop a contrastive learning mechanism to simultaneously align similar cluster pairs and keep noisy pairs apart from each other. This allows to obtain clusters with all cells of the same type (true positives), and avoid clusters with cells of different type (false positives). We demonstrate the effectiveness of our method on both simulated and real datasets. Results show that our new method significantly outperforms the state-of-the-art methods and has the ability to prevent overcorrection. AVAILABILITY AND IMPLEMENTATION: The python code to generate results and figures in this article is available at https://github.com/zhanglabNKU/HDMC, the data underlying this article is also available at this github repository. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao Wang 0099, Jia Wang 0051, Han Zhang 0017, Shenwei Huang, Yanbin Yin |
Bioinform. | 5 |
| 2021 | Predicting lncRNA-protein interactions based on graph autoencoders and collaborative trainingabstractLong non-coding RNAs(lncRNAs) play an important role in various biological processes. lncRNAs usually perform their molecular functions by interacting with proteins. Therefore, it is essential to predict potential lncRNA-protein associations for disease prevention and disease treatment. Label-propagation-based methods are widely used for predicting associations among biological entities. However, in these approaches, similarity computation and label propagation are separate procedures, which decrease the effectiveness of label propagation. Moreover, the prediction accuracy of existing models is also not ideal. In this paper, we proposed an end-to-end deep learning lncRNA-protein Interaction predictor through Graph Autoencoders and Collaborative training (LPIGAC). Different from previous studies, our model implemented two graph autoencoders on lncRNA graph and protein graph respectively, and trained these two graph autoencoders collaboratively. Graph autoencoders on lncRNA graph and protein graph are competent to reconstruct score matrix through initial association matrix, which is equivalent to propagate labels on graphs. This end-to-end framework can strengthen the robustness and precision of the label propagation procedure. Cross validations indicate that LPIGAC outperforms current lncRNA-protein associations prediction methods. Case studies demonstrate that LPIGAC is competent to detect potential lncRNA-protein associations. Source code of our paper is available at https://github.com/zhanglabNKULPIGAC. Zhuangwei Shi, Han Zhang 0017, Yanbin Yin |
BIBM | 4 |
| 2021 | Adversarial Dual-Channel Variational Graph Autoencoder for Synthetic Lethality Prediction in Human CancersabstractSynthetic Lethality (SL) is a type of vital gene interaction that can lead to various human diseases including cancers. Therefore, SL gene pair prediction can aid in the prevention and treatment of cancer. A number of computational approaches, especially Graph Neural Network (GNN) based methods, have been proposed for this link prediction problem on the graph. However, these GNN-based methods only consider embedding as deterministic vectors and do not take data distribution into account. Here we propose an Adversarial Dual-Channel Variational Graph Autoencoder based on semi-implicit variational inference for SL prediction in human cancers. We consider node embedding as a random variable that has an explicit Gaussian distribution. Then we design a dual-channel GCN encoder to inject stochasticity into the distribution parameters and allow latent embedding to exceed the Gaussian distribution. This hierarchical scheme leads to a more flexible posterior of latent embedding and enhances the model representation capacity. To further obtain a robust and stable representation, an adversarial module is devised for variance regularization. Experimental results compared with other state-of-the-art methods confirm the effectiveness of our proposed method. Moreover, we conduct a case study to demonstrate that our model can be very useful to predict novel SL pairs. Wei Li 0184, Han Zhang 0017, Jian Liu 0040, Yanbin Yin |
BIBM | 5 |
| 2021 | A representation learning model based on variational inference and graph autoencoder for predicting lncRNA-disease associationsabstractBACKGROUND: Numerous studies have demonstrated that long non-coding RNAs are related to plenty of human diseases. Therefore, it is crucial to predict potential lncRNA-disease associations for disease prognosis, diagnosis and therapy. Dozens of machine learning and deep learning algorithms have been adopted to this problem, yet it is still challenging to learn efficient low-dimensional representations from high-dimensional features of lncRNAs and diseases to predict unknown lncRNA-disease associations accurately. RESULTS: We proposed an end-to-end model, VGAELDA, which integrates variational inference and graph autoencoders for lncRNA-disease associations prediction. VGAELDA contains two kinds of graph autoencoders. Variational graph autoencoders (VGAE) infer representations from features of lncRNAs and diseases respectively, while graph autoencoders propagate labels via known lncRNA-disease associations. These two kinds of autoencoders are trained alternately by adopting variational expectation maximization algorithm. The integration of both the VGAE for graph representation learning, and the alternate training via variational inference, strengthens the capability of VGAELDA to capture efficient low-dimensional representations from high-dimensional features, and hence promotes the robustness and preciseness for predicting unknown lncRNA-disease associations. Further analysis illuminates that the designed co-training framework of lncRNA and disease for VGAELDA solves a geometric matrix completion problem for capturing efficient low-dimensional representations via a deep learning approach. CONCLUSION: Cross validations and numerical experiments illustrate that VGAELDA outperforms the current state-of-the-art methods in lncRNA-disease association prediction. Case studies indicate that VGAELDA is capable of detecting potential lncRNA-disease associations. The source code and data are available at https://github.com/zhanglabNKU/VGAELDA . Zhuangwei Shi, Han Zhang 0017, Xiongwen Quan, Yanbin Yin |
BMC Bioinform. | 5 |
| 2020 | Bayesian Multi-scale Convolutional Neural Network for Motif Occupancy IdentificationabstractConvolutional neural network (CNN) has been successfully used for the identification of motif occupancy. However, the CNN architecture requires varying length instead of fixed-length filters due to different motif lengths. Moreover, plain neural networks with single point estimation for weights suffer from over-fitting, which is more likely to occur as increasing parameters for multi-scale modeling.Hence, we have designed a Bayesian Multi-scale CNN. The model employs convolutional filters of different scales to extract latent features of DNA sequence, and incorporates Bayesian architecture which regards multi-scale weights as random variables. We further stack two sequential convolutional operations for mean and variance respectively, and apply Bayes by Back prop for posterior estimation of weights. Results have shown that our method not only improved the prediction performance for motif occupancy identification, but also prevented over-fitting due to the capability of Bayesian neural network. The model has also developed a measure of uncertainty estimation for model assessment. Wei Li 0184, Han Zhang 0017, Xiongwen Quan, Jing Xu 0008, Yanbin Yin |
BIBM | 6 |
| 2020 | GDASC: a GPU parallel-based web server for detecting hidden batch factorsabstractSUMMARY: We developed GDASC, a web version of our former DASC algorithm implemented with GPU. It provides a user-friendly web interface for detecting batch factors. Based on the good performance of DASC algorithm, it is able to give the most accurate results. For two steps of DASC, data-adaptive shrinkage and semi-non-negative matrix factorization, we designed parallelization strategies facing convex clustering solution and decomposition process. It runs more than 50 times faster than the original version on the representative RNA sequencing quality control dataset. With its accuracy and high speed, this server will be a useful tool for batch effects analysis. AVAILABILITY AND IMPLEMENTATION: http://bioinfo.nankai.edu.cn/gdasc.php. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao Wang 0099, Haidong Yi, Jia Wang 0051, Zhandong Liu, Yanbin Yin, Han Zhang 0017 |
Bioinform. | 5 |
| 2020 | eCAMI: simultaneous classification and motif identification for enzyme annotationabstractMOTIVATION: Carbohydrate-active enzymes (CAZymes) are extremely important to bioenergy, human gut microbiome, and plant pathogen researches and industries. Here we developed a new amino acid k-mer-based CAZyme classification, motif identification and genome annotation tool using a bipartite network algorithm. Using this tool, we classified 390 CAZyme families into thousands of subfamilies each with distinguishing k-mer peptides. These k-mers represented the characteristic motifs (in the form of a collection of conserved short peptides) of each subfamily, and thus were further used to annotate new genomes for CAZymes. This idea was also generalized to extract characteristic k-mer peptides for all the Swiss-Prot enzymes classified by the EC (enzyme commission) numbers and applied to enzyme EC prediction. RESULTS: This new tool was implemented as a Python package named eCAMI. Benchmark analysis of eCAMI against the state-of-the-art tools on CAZyme and enzyme EC datasets found that: (i) eCAMI has the best performance in terms of accuracy and memory use for CAZyme and enzyme EC classification and annotation; (ii) the k-mer-based tools (including PPR-Hotpep, CUPP and eCAMI) perform better than homology-based tools and deep-learning tools in enzyme EC prediction. Lastly, we confirmed that the k-mer-based tools have the unique ability to identify the characteristic k-mer peptides in the predicted enzymes. AVAILABILITY AND IMPLEMENTATION: https://github.com/yinlabniu/eCAMI and https://github.com/zhanglabNKU/eCAMI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jing Xu 0008, Han Zhang 0017, Jinfang Zheng, Philippe Dovoedo, Yanbin Yin |
Bioinform. | 5 |
| 2019 | Antimicrobial peptide identification using multi-scale convolutional networkabstractBACKGROUND: Antibiotic resistance has become an increasingly serious problem in the past decades. As an alternative choice, antimicrobial peptides (AMPs) have attracted lots of attention. To identify new AMPs, machine learning methods have been commonly used. More recently, some deep learning methods have also been applied to this problem. RESULTS: In this paper, we designed a deep learning model to identify AMP sequences. We employed the embedding layer and the multi-scale convolutional network in our model. The multi-scale convolutional network, which contains multiple convolutional layers of varying filter lengths, could utilize all latent features captured by the multiple convolutional layers. To further improve the performance, we also incorporated additional information into the designed model and proposed a fusion model. Results showed that our model outperforms the state-of-the-art models on two AMP datasets and the Antimicrobial Peptide Database (APD)3 benchmark dataset. The fusion model also outperforms the state-of-the-art model on an anti-inflammatory peptides (AIPs) dataset at the accuracy. CONCLUSIONS: Multi-scale convolutional network is a novel addition to existing deep neural network (DNN) models. The proposed DNN model and the modified fusion model outperform the state-of-the-art models for new AMP discovery. The source code and data are available at https://github.com/zhanglabNKU/APIN. Jing Xu 0008, Yanbin Yin, Xiongwen Quan, Han Zhang 0017 |
BMC Bioinform. | 3 |
| 2016 | ORFanFinder: automated identification of taxonomically restricted orphan genesabstractMOTIVATION: Orphan genes, also known as ORFans, are newly evolved genes in a genome that enable the organism to adapt to specific living environment. The gene content of every sequenced genome can be classified into different age groups, based on how widely/narrowly a gene's homologs are distributed in the context of species taxonomy. Those having homologs restricted to organisms of particular taxonomic ranks are classified as taxonomically restricted ORFans. RESULTS: Implementing this idea, we have developed an open source program named ORFanFinder and a free web server to allow automated classification of a genome's gene content and identification of ORFans at different taxonomic ranks. ORFanFinder and its web server will contribute to the comparative genomics field by facilitating the study of the origin of new genes and the emergence of lineage-specific traits in both prokaryotes and eukaryotes. AVAILABILITY AND IMPLEMENTATION: http://cys.bios.niu.edu/orfanfinder CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Ekstrom, Yanbin Yin |
Bioinform. | 2 |
| 2013 | An integrated toolkit for accurate prediction and analysis of cis-regulatory motifs at a genome scaleabstractMOTIVATION: We present an integrated toolkit, BoBro2.0, for prediction and analysis of cis-regulatory motifs. This toolkit can (i) reliably identify statistically significant cis-regulatory motifs at a genome scale; (ii) accurately scan for all motif instances of a query motif in specified genomic regions using a novel method for P-value estimation; (iii) provide highly reliable comparisons and clustering of identified motifs, which takes into consideration the weak signals from the flanking regions of the motifs; and (iv) analyze co-occurring motifs in the regulatory regions. RESULTS: We have carried out systematic comparisons between motif predictions using BoBro2.0 and the MEME package. The comparison results on Escherichia coli K12 genome and the human genome show that BoBro2.0 can identify the statistically significant motifs at a genome scale more efficiently, identify motif instances more accurately and get more reliable motif clusters than MEME. In addition, BoBro2.0 provides correlational analyses among the identified motifs to facilitate the inference of joint regulation relationships of transcription factors. AVAILABILITY: The source code of the program is freely available for noncommercial uses at http://code.google.com/p/bobro/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qin Ma 0003, Bingqiang Liu, Chuan Zhou 0009, Yanbin Yin, Ying Xu 0001 |
Bioinform. | 4 |
| 2011 | Measurement and prediction of DTMB reception quality in single frequency networksabstractThis paper presents a method to measure and predict the reception quality of the digital terrestrial/television multimedia broadcasting (DTMB) system in a single frequency network (SFN) environment. Measurement of the signal-to-noise ratio (SNR) threshold at several market commercial receivers was carried out in the laboratory. After analyzing the channel capacity loss produced by the multiple transmitter reception, a prediction model of SNR threshold is presented. Simulation results are also performed to verify the accuracy of the prediction method. This model and method could be adopted for SFN planning to ensure a good service quality in the SFN environment. Keqian Yan, Wenbo Ding 0001, Yanbin Yin, Fang Yang 0001, Changyong Pan |
IWCMC | 4 |
| 2010 | GolgiP: prediction of Golgi-resident proteins in plantsabstractUNLABELLED: We present a novel Golgi-prediction server, GolgiP, for computational prediction of both membrane- and non-membrane-associated Golgi-resident proteins in plants. We have employed a support vector machine-based classification method for the prediction of such Golgi proteins, based on three types of information, dipeptide composition, transmembrane domain(s) (TMDs) and functional domain(s) of a protein, where the functional domain information is generated through searching against the Conserved Domains Database, and the TMD information includes the number of TMDs, the length of TMD and the number of TMDs at the N-terminus of a protein. Using GolgiP, we have made genome-scale predictions of Golgi-resident proteins in 18 plant genomes, and have made the preliminary analysis of the predicted data. AVAILABILITY: The GolgiP web service is publically available at http://csbl1.bmb.uga.edu/GolgiP/. Wen-Chi Chou, Yanbin Yin |
Bioinform. | 2 |