VLDB 2026 Research / reviewers in the wild / expert
Yang Yang 0030
dblp:48/450-30
· DBLP profile ↗
55ranked-venue papers
13as first author
28since 2021 · last 2026
0000-0001-5720-773XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 44 · 8 first-author · 24 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Align then clip: Refining graph for face clustering
Yanlun Tu, Guoliang Cao, Jialiang Shen, Min Wang 0024, Wentao Liu 0002, Chen Qian 0006, Yang Yang 0030 |
Neural Networks | 8 |
| 2026 | Harmonizing Generalization and Specialization: Uncertainty-Informed Collaborative Learning for Semi-Supervised Medical Image SegmentationabstractVision foundation models have demonstrated strong generalization in medical image segmentation by leveraging large-scale, heterogeneous pretraining. However, they often struggle to generalize to specialized clinical tasks under limited annotations or rare pathological variations, due to a mismatch between general priors and task-specific requirements. To address this, we propose Uncertainty-informed Collaborative Learning (UnCoL), a dual-teacher framework that harmonizes generalization and specialization in semi-supervised medical image segmentation. Specifically, UnCoL distills both visual and semantic representations from a frozen foundation model to transfer general knowledge, while concurrently maintaining a progressively adapting teacher to capture fine-grained and task-specific representations. To balance guidance from both teachers, pseudo-label learning in UnCoL is adaptively regulated by predictive uncertainty, which selectively suppresses unreliable supervision and stabilizes learning in ambiguous regions. Experiments on diverse 2D and 3D benchmarks show that UnCoL consistently outperforms existing methods across most datasets and metrics, while achieving comparable performance in only a few cases. Moreover, our model delivers near fully supervised performance with markedly reduced annotation requirements. Code is available at: https://github.com/VivienLu/UnCoL. Wenjing Lu, Yang Yang 0030 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Design RNA with Specified Secondary Structures Using CDMsabstractAs RNA function is strongly tied to its secondary structure, designing RNA molecules with specified structures stands as a key challenge in computational biology. Existing approaches often either yield limited results or incur high computational costs. Building on the success of Denoising Diffusion Probabilistic Models (DDPMs) and their variants, including Conditional Diffusion Models (CDMs) in diverse fields such as image generation, we apply CDMs to generate RNA sequences matching specified secondary structures, using One-Hot or static vocabulary encoding, classifier-free guidance, and a Transformer denoiser to capture long-range dependencies. Experiments on the Rfam dataset show that our best model achieves a task-solving rate of 0.846, outperforming the current next-best method (≤ 0.648). Our models exhibit better performance with a simpler architecture, demonstrating the effectiveness of CDMs in sequence design problems. Future work may focus on optimizing the conditional encoding method and the architecture of the denoiser, as well as developing biological relevance metrics for generated sequences. Zhenran Xiao, Letian Chen, Yichong Li, Yang Yang 0030 |
SMC | 4 |
| 2025 | DRAG: design RNAs as hierarchical graphs with reinforcement learningabstractThe rapid development of RNA vaccines and therapeutics puts forward intensive requirements on the sequence design of RNAs. RNA sequence design, or RNA inverse folding, aims to generate RNA sequences that can fold into specific target structures. To date, efficient and high-accuracy prediction models for secondary structures of RNAs have been developed. They provide a basis for computational RNA sequence design methods. Especially, reinforcement learning (RL) has emerged as a promising approach for RNA design due to its ability to learn from trial and error in generation tasks and work without ground truth data. However, existing RL methods are limited in considering complex hierarchical structures in RNA design environments. To address the above limitation, we propose DRAG, an RL method that builds design environments for target secondary structures with hierarchical division based on graph neural networks. Through extensive experiments on benchmark datasets, DRAG exhibits remarkable performance compared with current machine-learning approaches for RNA sequence design. This advantage is particularly evident in long and intricate tasks involving structures with significant depth. Yichong Li, Xiaoyong Pan, Hong-Bin Shen, Yang Yang 0030 |
Briefings Bioinform. | 4 |
| 2024 | Enhancing RBP Binding Site Prediction on Long RNA Sequences through Large Language ModelsabstractRNA binding proteins (RBPs) are crucial in various biological processes and gene regulation. Accurately predicting RBP binding sites is essential for advancing research in biology and medicine. Despite the success of deep learning models in predicting these sites, many RNA sequences exceed the manageable length for existing models. These models typically segment long RNA sequences into shorter sections to predict protein binding, which disrupts RNA structure and loses sequence information. In this study, we develop a language model that processes long-sequence RNA without segmentation, maintaining sequence integrity. Inspired by the Enformer model, we use 128 bases as a token, enabling the processing of RNA sequences up to 20,480 base pairs. This model effectively learns and analyzes the interactions within long sequences, significantly enhancing RBP binding site prediction. We perform three prediction tasks: whole sequence prediction, segmented multi-label prediction, and segmented single-label prediction, achieving promising results. Additionally, we address the issue of imbalanced positive and negative sample distribution in our dataset, exploring methods to mitigate its impact on model classification performance. Junkun Guo, Yang Yang 0030 |
BIBM | 2 |
| 2024 | RBP-Former: Joint Prediction of RNA-protein Binding Sites on Full-length RNA Transcripts for Multiple RBPsabstractRNA-binding proteins (RBPs) are essential for gene expression, and the complex RNA-protein interaction mechanisms require analysis of global RNA information. Therefore, accurate prediction of RBP binding sites on full-length RNA transcripts is crucial for understanding these mechanisms and their roles in diseases. While machine learning methods can predict RBP binding to RNA fragments, extending this to full-length transcripts presents challenges due to sequence length and data imbalance. In this paper, we introduce RBP-Former, a binding site joint prediction model designed specifically for full-length RNA transcripts that can be used for multiple RBPs. This model processes information at both coarse and fine-grained levels to fully exploit sequence data and its interactions with multiple RBPs. We develop multi-level imbalance learning strategies, achieving favorable results on imbalanced data. Our method outperforms existing methods in predicting binding sites on full-length RNA transcripts for multiple RBPs, demonstrating its effectiveness in handling imbalanced label and sample distributions. Yichong Li, Xiaoyong Pan, Yang Yang 0030 |
BIBM | 5 |
| 2024 | UP-SAM: Uncertainty-Informed Adaptation of Segment Anything Model for Semi-Supervised Medical Image SegmentationabstractSemi-supervised segmentation is extensively employed in medical image analysis due to its ability to leverage a small amount of labeled data alongside abundant unlabeled data. However, its performance is hindered by the inadequate knowledge of the data domain learned from limited labeled data and the absence of effective strategies for exploiting unlabeled regions, especially when annotations are extremely scarce. To address these challenges, the Segment Anything Model (SAM) has emerged as a promising solution. As a foundation model enriched by extensive and diverse domain knowledge, SAM has been leveraged to mitigate the epistemic uncertainty (EU) of semi-supervised segmentation models, while aleatoric uncertainty (AU) is often ignored. In this paper, we propose a novel semi-supervised medical image segmentation framework called UP-SAM, which adapts SAM for dual uncertainty assessments. The framework achieves effective collaboration between large foundation models and domain-specific models, leading to a simultaneous reduction in the impact of EU and AU. The experiments on the left atrium and pancreas datasets demonstrate the superior efficacy of UP-SAM against baseline methods. Particularly, UP-SAM exhibits substantial advantages over other semi-supervised learning models when dealing with exceedingly scarce labeled data. Code is available at https://github.com/VivienLu/UP-SAM. Wenjing Lu, Yang Yang 0030 |
BIBM | 3 |
| 2024 | Enhancing Chest X-ray Diagnostics with Neighbor-assisted Multimodal IntegrationabstractIn the field of medical imaging analysis, particularly in interpreting chest X-rays, deep learning models have shown remarkable progress. Nonetheless, these models often face challenges such as limited annotation and inadequate utilization of public data resources. This is particularly apparent with databases containing multimodal data, such as images and medical reports, where the effective integration of this multimodal information remains difficult. To address these limitations, we propose the Neighbor-Assisted Multimodal Attention Network (NAMAN), a novel approach designed to leverage retrieval augmentation techniques to enhance disease classification performance. NAMAN combines nearest neighbor search with multimodal fusion, utilizing both visual features from similar X-ray images and textual information from corresponding medical records. The experimental results demonstrate the efficacy of incorporating retrieved neighbor information and multimodal integration mechanisms in NAMAN. Our ablation studies offer insights into the optimal configuration of the model, including the effects of various attention mechanisms and the number of retrieved neighbors. This work contributes to the expanding field of retrieval-augmented approaches in medical imaging, presenting a promising avenue for leveraging large-scale, multimodal medical databases to enhance diagnostic accuracy and reliability. Chenjie Xu, Yang Zhang 0040, Yang Yang 0030 |
BIBM | 6 |
| 2024 | An Efficient Prototype-Based Clustering Approach for Edge Pruning in Graph Neural Networks to Battle Over-Smoothing
Wenjing Lu, Yang Yang 0030 |
IJCAI | 3 |
| 2024 | Improving diagnosis and outcome prediction of gastric cancer via multimodal learning using whole slide pathological images and gene expression
Yuzhang Xie, Qingqing Sang, Qian Da, Guoshuai Niu, Yunqin Chen, Bing-Ya Liu, Yang Yang 0030, Wentao Dai |
Artif. Intell. Medicine | 10 |
| 2023 | Isoform Function Prediction Based on Heterogeneous Graph Attention NetworksabstractIsoforms refer to different mRNA molecules transcribed from the same gene, which can be translated into proteins with varying structures and functions. Predicting the functions of isoforms is an essential topic in bioinformatics as it can provide valuable insights into the intricate mechanisms of gene regulation and biological processes. Conventionally, gene function labels are standardized in Gene Ontology (GO) terms. However, traditional methods for predicting isoform function are largely limited by the absence of isoform-specific labels, sparse annotations, and the vast number of GO terms. To address these issues, we propose HANIso, a deep learning-based method for isoform function prediction. HANIso leverages a pretrained protein language model to extract features from protein sequences. It also integrates heterogeneous information, such as isoform sequence features, GO annotations, and isoform interaction data, using a Heterogeneous Graph Attention Network (HAN). This allows the model to learn the importance of different sources of information and their semantic relationships through the attention mechanism. Our method can predict function labels at both the gene level and isoform level. We conduct experiments on two species datasets, and the results demonstrate that our method outperforms existing methods on both AUROC and AUPRC. HANIso has the potential to overcome the limitations of traditional methods and provide a more accurate and comprehensive understanding of isoform function. Kuo Guo, Hong-Bin Shen, Yang Yang 0030 |
BIBM | 5 |
| 2023 | Enhancing Cancer Gene Prediction through Aligned Fusion of Multiple PPI Networks Using Graph Transformer ModelsabstractDespite the considerable progress that has been made in cancer research, identifying cancer genes remains a significant challenge due to the intricate nature of the disease. Given the importance of incorporating gene interaction relationships in identifying potential cancer genes, Graph Neural Networks (GNNs) have garnered increasing attention for their ability to model gene associations effectively. Particularly, recent studies have demonstrated the superiority of GNNs in deciphering the knowledge embedded in Protein-Protein Interaction (PPI) networks for cancer gene prediction. However, these studies primarily focused on a single PPI network, overlooking the valuable insights encapsulated within other PPI networks. Additionally, previous endeavors often centered around pan-cancer datasets, neglecting the importance of predicting specific cancer genes. To address these limitations, we present a novel method called MPIT, which employs Graph Transformer Networks (GTNs) to identify specific cancer driver genes. MPIT effectively integrates data from diverse PPI and multi-omics data via the alignment and fusion of gene representations learned from different PPI networks. We collect three distinct cancer cell line datasets to assess the model performance. Our experimental findings demonstrate the superiority of MPIT over the existing methods, achieving the state-of-the-art performance across all three datasets. Zebei Han, Gufeng Yu, Yang Yang 0030 |
BIBM | 3 |
| 2023 | UPCoL: Uncertainty-Informed Prototype Consistency Learning for Semi-supervised Medical Image Segmentation
Wenjing Lu, Jiahao Lei, Peng Qiu, Rui Sheng, Jinhua Zhou, Xinwu Lu, Yang Yang 0030 |
MICCAI (4) | 7 |
| 2023 | MIGGRI: A multi-instance graph neural network model for inferring gene regulatory networks for Drosophila from spatial expression imagesabstractRecent breakthrough in spatial transcriptomics has brought great opportunities for exploring gene regulatory networks (GRNs) from a brand-new perspective. Especially, the local expression patterns and spatio-temporal regulation mechanisms captured by spatial expression images allow more delicate delineation of the interplay between transcript factors and their target genes. However, the complexity and size of spatial image collections pose significant challenges to GRN inference using image-based methods. Extracting regulatory information from expression images is difficult due to the lack of supervision and the multi-instance nature of the problem, where a gene often corresponds to multiple images captured from different views. While graph models, particularly graph neural networks, have emerged as a promising method for leveraging underlying structure information from known GRNs, incorporating expression images into graphs is not straightforward. To address these challenges, we propose a two-stage approach, MIGGRI, for capturing comprehensive regulatory patterns from image collections for each gene and known interactions. Our approach involves a multi-instance graph neural network (GNN) model for GRN inference, which first extracts gene regulatory features from spatial expression images via contrastive learning, and then feeds them to a multi-instance GNN for semi-supervised learning. We apply our approach to a large set of Drosophila embryonic spatial gene expression images. MIGGRI achieves outstanding performance in the inference of GRNs for early eye development and mesoderm development of Drosophila, and shows robustness in the scenarios of missing image information. Additionally, we perform interpretable analysis on image reconstruction and functional subgraphs that may reveal potential pathways or coordinate regulations. By leveraging the power of graph neural networks and the information contained in spatial expression images, our approach has the potential to advance our understanding of gene regulation in complex biological systems. Gufeng Yu, Yang Yang 0030 |
PLoS Comput. Biol. | 3 |
| 2023 | HAMIL: Hierarchical aggregation-based multi-instance learning for microscopy image classification
Yang Yang 0030, Yanlun Tu, Houchao Lei |
Pattern Recognit. | 1 |
| 2022 | Learning Time-Series Images of Niacin Skin-Flushing Test for the Diagnosis of Schizophrenia and Affective DisorderabstractNiacin skin-flushing test is a promising method for fast and objective diagnosis of schizophrenia and affective disorder. Compared to healthy controls, the patients have attenuated flushing extent and reduced flushing rate. The traditional analysis of niacin test relies on manually evaluated flushing extent, while flushing patterns sh own in the sk in test images have not yet been considered. To exploit the potential of raw images for diagnosis and avoid subjective and time-consuming estimation of the flushing levels, we propose a CNN-LSTM hybrid model to learn the time-series of images captured in the niacin skin-flushing test. Moreover, an attention layer is introduced to enhance the model performance and interpretability. We collect a data set of 796 participants, and the experimental results show that the new model outperforms the traditional 4-point scaled measurement and conventional deep learning methods by large margins. Besides, the attention scores yielded by the model reveals key time points during the flushing reaction. Shiwen Dong, Chunling Wan, Yang Yang 0030 |
BIBM | 5 |
| 2022 | Survival Prediction for Gastric Cancer via Multimodal Learning of Whole Slide Images and Gene ExpressionabstractGastric cancer (GC) is one of the most common malignancies worldwide. As histopathology tissue analysis is considered as the gold standard in cancer studies, whole slide images (WSIs) have been widely used for GC diagnosis and prognosis, while multimodal studies for GC patients have been very few. Especially, WSIs and gene expression are complementary modalities of data, thus fusion of these two modalities has great potential in the prediction of survival outcomes and other computer-aided tasks, like the mechanism study and clinical treatment for GC patients. However, multimodal learning requires good data fusion strategies and also suffers from the missing data issue. To address these issues, we propose GC-SPLeM, to predict risk scores for patients, which consists of three parts, WSI feature extraction, modal-fusing network, and GNN-based predictor. We conduct experiments on a GC dataset built by ourselves and a public dataset for survival prediction. For both datasets, GC-SPLeM outperforms the state-of-the-art single-modality learning method and multimodal learning method by large margins (over 5% on C-index). We find that the GNN plays an important role in performance enhancement. Through learning the graph of patients, topological structure and neighborhood clinical information are encoded into feature representations of patients. GC-SPLeM not only improves the survival prediction results but also has advantages in dealing with incomplete data over other methods. The Source code, sample data, and gene list of this study are available at https://github.com/constantjxyz/GC-SPLeM. Yuzhang Xie, Guoshuai Niu, Qian Da, Wentao Dai, Yang Yang 0030 |
BIBM | 5 |
| 2022 | FUSSNet: Fusing Two Sources of Uncertainty for Semi-supervised Medical Image Segmentation
Jinyi Xiang, Peng Qiu, Yang Yang 0030 |
MICCAI (8) | 3 |
| 2022 | SIFLoc: a self-supervised pre-training method for enhancing the recognition of protein subcellular localization in immunofluorescence microscopic imagesabstractWith the rapid growth of high-resolution microscopy imaging data, revealing the subcellular map of human proteins has become a central task in the spatial proteome. The cell atlas of the Human Protein Atlas (HPA) provides precious resources for recognizing subcellular localization patterns at the cell level, and the large-scale annotated data enable learning via advanced deep neural networks. However, the existing predictors still suffer from the imbalanced class distribution and the lack of labeled data for minor classes. Thus, it is necessary to develop new methods for coping with these issues. We leverage the self-supervised learning protocol to address these problems. Especially, we propose a pre-training scheme to enhance the conventional supervised learning framework called SIFLoc. The pre-training is featured by a hybrid data augmentation method and a modified contrastive loss function, aiming to learn good feature representations from microscopic images. The experiments are performed on a large-scale immunofluorescence microscopic image dataset collected from the HPA database. Using the same deep neural networks as the classifier, the model pre-trained via SIFLoc not only outperforms the model without pre-training by a large margin but also shows advantages over the state-of-the-art self-supervised learning methods. Especially, SIFLoc improves the prediction accuracy for minor organelles significantly. Yanlun Tu, Houchao Lei, Hong-Bin Shen, Yang Yang 0030 |
Briefings Bioinform. | 4 |
| 2022 | GraphLoc: a graph neural network model for predicting protein subcellular localization from immunohistochemistry imagesabstractMOTIVATION: Recognition of protein subcellular distribution patterns and identification of location biomarker proteins in cancer tissues are important for understanding protein functions and related diseases. Immunohistochemical (IHC) images enable visualizing the distribution of proteins at the tissue level, providing an important resource for the protein localization studies. In the past decades, several image-based protein subcellular location prediction methods have been developed, but the prediction accuracies still have much space to improve due to the complexity of protein patterns resulting from multi-label proteins and the variation of location patterns across cell types or states. RESULTS: Here, we propose a multi-label multi-instance model based on deep graph convolutional neural networks, GraphLoc, to recognize protein subcellular location patterns. GraphLoc builds a graph of multiple IHC images for one protein, learns protein-level representations by graph convolutions and predicts multi-label information by a dynamic threshold method. Our results show that GraphLoc is a promising model for image-based protein subcellular location prediction with model interpretability. Furthermore, we apply GraphLoc to the identification of candidate location biomarkers and potential members for protein networks. A large portion of the predicted results have supporting evidence from the existing literatures and the new candidates also provide guidance for further experimental screening. AVAILABILITY AND IMPLEMENTATION: The dataset and code are available at: www.csbio.sjtu.edu.cn/bioinf/GraphLoc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jin-Xian Hu, Yang Yang 0030, Ying-Ying Xu, Hong-Bin Shen |
Bioinform. | 2 |
| 2022 | Accurate inference of gene regulatory interactions from spatial gene expression with deep contrastive learningabstractMOTIVATION: Reverse engineering of gene regulatory networks (GRNs) has long been an attractive research topic in system biology. Computational prediction of gene regulatory interactions has remained a challenging problem due to the complexity of gene expression and scarce information resources. The high-throughput spatial gene expression data, like in situ hybridization images that exhibit temporal and spatial expression patterns, has provided abundant and reliable information for the inference of GRNs. However, computational tools for analyzing the spatial gene expression data are highly underdeveloped. RESULTS: In this study, we develop a new method for identifying gene regulatory interactions from gene expression images, called ConGRI. The method is featured by a contrastive learning scheme and deep Siamese convolutional neural network architecture, which automatically learns high-level feature embeddings for the expression images and then feeds the embeddings to an artificial neural network to determine whether or not the interaction exists. We apply the method to a Drosophila embryogenesis dataset and identify GRNs of eye development and mesoderm development. Experimental results show that ConGRI outperforms previous traditional and deep learning methods by a large margin, which achieves accuracies of 76.7% and 68.7% for the GRNs of early eye development and mesoderm development, respectively. It also reveals some master regulators for Drosophila eye development. AVAILABILITYAND IMPLEMENTATION: https://github.com/lugimzheng/ConGRI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lujing Zheng, Zhenhuan Liu, Yang Yang 0030, Hong-Bin Shen |
Bioinform. | 3 |
| 2022 | ProtPlat: an efficient pre-training platform for protein classification based on FastTextabstractBACKGROUND: For the past decades, benefitting from the rapid growth of protein sequence data in public databases, a lot of machine learning methods have been developed to predict physicochemical properties or functions of proteins using amino acid sequence features. However, the prediction performance often suffers from the lack of labeled data. In recent years, pre-training methods have been widely studied to address the small-sample issue in computer vision and natural language processing fields, while specific pre-training techniques for protein sequences are few. RESULTS: In this paper, we propose a pre-training platform for representing protein sequences, called ProtPlat, which uses the Pfam database to train a three-layer neural network, and then uses specific training data from downstream tasks to fine-tune the model. ProtPlat can learn good representations for amino acids, and at the same time achieve efficient classification. We conduct experiments on three protein classification tasks, including the identification of type III secreted effectors, the prediction of subcellular localization, and the recognition of signal peptides. The experimental results show that the pre-training can enhance model performance effectively and ProtPlat is competitive to the state-of-the-art predictors, especially for small datasets. We implement the ProtPlat platform as a web service ( https://compbio.sjtu.edu.cn/protplat ) that is accessible to the public. CONCLUSIONS: To enhance the feature representation of protein amino acid sequences and improve the performance of sequence-based classification tasks, we develop ProtPlat, a general platform for the pre-training of protein sequences, which is featured by a large-scale supervised training based on Pfam database and an efficient learning model, FastText. The experimental results of three downstream classification tasks demonstrate the efficacy of ProtPlat. Yang Yang 0030 |
BMC Bioinform. | 2 |
| 2021 | Predicting RNA-RBP Interactions by Using a Pseudo-Siamese NetworkabstractLarge scale RNA-protein binding data made computational identification of RNA-protein interactions possible, which provides great convenience for revealing the interplay between non-coding RNAs and RNA-binding proteins (RBPs). Various machine learning methods have been applied to the prediction of RNA-protein interactions (RPIs). However, most of the RPI predictors were designed for linear RNAs, while methods for circular RNAs (circRNAs) are few. Besides, most of the predictors use only RNA sequence or only protein sequences as input. Automatic feature learning from RNA and RBP sequences using deep neural networks has great potential to improve RPI predictors. In this paper, we propose a new method, PSi-bind, to predict the binding relationship between RNAs and RBPs, which is featured by a pseudo-Siamese neural network. We feed both RNA and RBP sequence embedding features to the network and perform an end-to-end learning to discover the latent association between RNAs and RBPs. Especially, we train our model on a large-scale circRNA dataset. The experimental results show that the model achieves very high accuracy on independent test set, and outperforms the state-of-the-art RPI predictors by a large margin. Liangliang Yuan, Yang Yang 0030 |
BIBM | 3 |
| 2021 | Prediction of Protein Subcellular Localization from Microscopic Images via Few-Shot Learning
Francesco Arcamone, Yanlun Tu, Yang Yang 0030 |
ISBRA | 3 |
| 2021 | Recognizing binding sites of poorly characterized RNA-binding proteins on circular RNAs using attention Siamese networkabstractCircular RNAs (circRNAs) interact with RNA-binding proteins (RBPs) to play crucial roles in gene regulation and disease development. Computational approaches have attracted much attention to quickly predict highly potential RBP binding sites on circRNAs using the sequence or structure statistical binding knowledge. Deep learning is one of the popular learning models in this area but usually requires a lot of labeled training data. It would perform unsatisfactorily for the less characterized RBPs with a limited number of known target circRNAs. How to improve the prediction performance for such small-size labeled characterized RBPs is a challenging task for deep learning-based models. In this study, we propose an RBP-specific method iDeepC for predicting RBP binding sites on circRNAs from sequences. It adopts a Siamese neural network consisting of a lightweight attention module and a metric module. We have found that Siamese neural network effectively enhances the network capability of capturing mutual information between circRNAs with pairwise metric learning. To further deal with the small-sample size problem, we have performed the pretraining using available labeled data from other RBPs and also demonstrate the efficacy of this transfer-learning pipeline. We comprehensively evaluated iDeepC on the benchmark datasets of RBP-binding circRNAs, and the results suggest iDeepC achieving promising results on the poorly characterized RBPs. The source code is available at https://github.com/hehew321/iDeepC. Hehe Wu, Xiaoyong Pan, Yang Yang 0030, Hong-Bin Shen |
Briefings Bioinform. | 3 |
| 2021 | CrepHAN: cross-species prediction of enhancers by using hierarchical attention networksabstractMOTIVATION: Enhancers are important functional elements in genome sequences. The identification of enhancers is a very challenging task due to the great diversity of enhancer sequences and the flexible localization on genomes. Till now, the interactions between enhancers and genes have not been fully understood yet. To speed up the studies of the regulatory roles of enhancers, computational tools for the prediction of enhancers have emerged in recent years. Especially, thanks to the ENCODE project and the advances of high-throughput experimental techniques, a large amount of experimentally verified enhancers have been annotated on the human genome, which allows large-scale predictions of unknown enhancers using data-driven methods. However, except for human and some model organisms, the validated enhancer annotations are scarce for most species, leading to more difficulties in the computational identification of enhancers for their genomes. RESULTS: In this study, we propose a deep learning-based predictor for enhancers, named CrepHAN, which is featured by a hierarchical attention neural network and word embedding-based representations for DNA sequences. We use the experimentally supported data of the human genome to train the model, and perform experiments on human and other mammals, including mouse, cow and dog. The experimental results show that CrepHAN has more advantages on cross-species predictions, and outperforms the existing models by a large margin. Especially, for human-mouse cross-predictions, the area under the receiver operating characteristic (ROC) curve (AUC) score of ROC curve is increased by 0.033∼0.145 on the combined tissue dataset and 0.032∼0.109 on tissue-specific datasets. AVAILABILITY AND IMPLEMENTATION: bcmi.sjtu.edu.cn/∼yangyang/CrepHAN.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianwei Hong, Ruitian Gao, Yang Yang 0030 |
Bioinform. | 3 |
| 2021 | FlyIT: Drosophila Embryogenesis Image Annotation based on Image Tiling and Convolutional Neural NetworksabstractWith the rise of image-based transcriptomics, spatial gene expression data has become increasingly important for understanding gene regulations from the tissue level down to the cell level. Especially, the gene expression images of Drosophila embryos provide a new data source in the study of Drosophila embryogenesis. It is imperative to develop automatic annotation tools since manual annotation is labor-intensive and requires professional knowledge. Although a lot of image annotation methods have been proposed in the computer vision field, they may not work well for gene expression images, due to the great difference between these two annotation tasks. Besides the apparent difference on images, the annotation is performed at the gene level rather than the image level, where the expression patterns of a gene are recorded in multiple images. Moreover, the annotation terms often correspond to local expression patterns of images, yet they are assigned collectively to groups of images and the relations between the terms and single images are unknown. In order to learn the spatial expression patterns comprehensively for genes, we propose a new method, called FlyIT (image annotation based on Image Tiling and convolutional neural networks for fruit Fly). We implement two versions of FlyIT, learning at image-level and gene-level, respectively. The gene-level version employs an image tiling strategy to get a combined image feature representation for each gene. FlyIT uses a pre-trained ResNet model to obtain feature representation and a new loss function to deal with the class imbalance problem. As the annotation of Drosophila images is a multi-label classification problem, the new loss function considers the difficulty levels for recognizing different labels of the same sample and adjusts the sample weights accordingly. The experimental results on the FlyExpress database show that both the image tiling strategy and the deep architecture lead to the great enhancement of the annotation performance. FlyIT outperforms the existing annotators by a large margin (over 9 percent on AUC and 12 percent on macro F1 for predicting the top 10 terms). It also shows advantages over other deep learning models, including both single-instance and multi-instance learning frameworks. Tiange Li, Yang Yang 0030, Hong-Bin Shen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | KenDTI: An Ensemble Model for Predicting Drug-Target Interaction by Integrating Multi-Source InformationabstractThe identification of drug-target interactions (DTIs) is an essential step in the process of drug discovery. As experimental validation suffers from high cost and low success rate, various computational models have been exploited to infer potential DTIs. The performance of DTI prediction depends heavily on the features extracted from drugs and target proteins. The existing predictors vary in input information and each has its own advantages. Therefore, combining the advantages of individual models and generating high-quality representations for drug-target pairs are effective ways to improve the performance of DTI prediction. In this study, we exploit both biochemical characteristics of drugs via network integration and molecular sequences via word embeddings, then we develop an ensemble model, KenDTI, based on two types of methods, i.e., network-based and classification-based. We assess the performance of KenDTI on two large-scale datasets, The experimental results show that KenDTI outperforms the state-of-the-art DTI predictors by a large margin. Moreover, KenDTI is robust against missing data in input networks and lack of prior knowledge. It is able to predict for drug-candidate chemical compounds with scarce information. Zhimiao Yu, Jiarui Lu, Yang Yang 0030 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | NiuEM: A Nested-iterative Unsupervised Learning Model for Single-particle Cryo-EM Image ProcessingabstractCryo-electron microscopy (cryo-EM) has become a mainstream technology for solving spatial structures of biomacromolecules, while the processing of cryo-EM images is a very challenging task. One of the great challenges is the high noise in the images. A common method is to cluster the images with close projecting angles to get mean images, which are used for 3D reconstruction. However, due to the extremely low signal-to-noise-ratio, common clustering methods often fail to obtain high-quality mean images, leading to poorly reconstructed structures. In this study, we present a new unsupervised learning framework, called NiuEM, to discriminate images captured from different angles and yield cluster-mean images. NiuEM first generates pseudo-labels and then exploits both contrastive loss and cross-entropy loss for training convolutional layers to learn feature representations. Moreover, the pseudo-labels are updated iteratively to enhance the reliability of labels. We assess the performance of NiuEM on four data sets via both visualized and quantitative experiments. Especially, two kinds of metrics are adopted to measure the performance, regarding the clustering quality and the resolution of reconstructed 3D models, respectively. The experimental results show that NiuEM achieves very competitive clustering accuracy in the comparison with the state-of-the-art image clustering methods. Moreover, the cluster mean images yielded by NiuEM lead to better initial 3D models compared with the mainstream reconstruction tools. Jia-Ming Cai, Wangjie Zheng, Yang Yang 0030, Hong-Bin Shen |
BIBM | 4 |
| 2020 | Tripfly: Predicting Gene-gene Interaction of Drosophila Eye Development Using Triplet LossabstractThe reconstruction of gene regulatory network (GRN) is of significance in system biology. In recent years, benefiting from the advances of deep learning technologies, image-based gene expression data, which contains spatial expression patterns, has become a new resource in network inference. Most of the existing image-based GRN inference models are based on unsupervised models, due to the lack of labeled data. And a few methods employ supervised learning models, whose performance is limited by the scale of training data.In this study, in order to predict the gene regulatory network of the eye development of Drosophila embryos, we develop a weakly supervised learning method. We generate image triplets of genes according to their orientation and developing stage. Then we build a deep convolutional neural network, using triplet loss to train a siamese network and extract the relationship between genes. The new method achieves promising results in the prediction of gene regulatory relationship in the eye development of Drosophila with a total accuracy of over 72%. Zhenhuan Liu, Jiafeng Chen, Yang Yang 0030 |
BIBM | 3 |
| 2020 | An Unsupervised Iterative Model for Single-Particle Cryo-EM Image Denoising Based on Siamese Neural NetworkabstractCryo-electron microscopy (cryo-EM) has become an important technology in the field of structural biology. Through the continuous development and improvement of hardware and software, more and more molecular biological structures close to atomic resolution have been resolved. In order to obtain accurate and reliable three-dimensional structure, clustering analysis of cryo-EM images is a very important and critical step.Different from traditional images, cryo-EM images have very high noise, in which the particles have random horizontal position and rotation directions. Most of the existing cryo-EM image processing tools implement traditional clustering methods, which do not perform well in such a highly-noisy scenario. In this paper, combined with the traditional K-means clustering algorithm, a new iterative clustering algorithm based on the unsupervised generative model is proposed. The iterative algorithm is mainly based on the idea of siamese network. First, we use K-means algorithm and Resnet to extract pre-labels for unlabeled data, and then in each subsequent iteration, we use siamese network to continuously extract and update the feature matrix of each image. After each epoch, K-means is adopted to cluster the image data based on the new feature representations. The contrastive loss is used as the loss function. The experimental results show that our method significantly improves the signal-to-noise ratio of images, and has better clustering performance compared with traditional methods. Wangjie Zheng, Yang Yang 0030 |
BIBM | 2 |
| 2020 | ImPLoc: a multi-instance deep learning model for the prediction of protein subcellular localization based on immunohistochemistry imagesabstractMOTIVATION: The tissue atlas of the human protein atlas (HPA) houses immunohistochemistry (IHC) images visualizing the protein distribution from the tissue level down to the cell level, which provide an important resource to study human spatial proteome. Especially, the protein subcellular localization patterns revealed by these images are helpful for understanding protein functions, and the differential localization analysis across normal and cancer tissues lead to new cancer biomarkers. However, computational tools for processing images in this database are highly underdeveloped. The recognition of the localization patterns suffers from the variation in image quality and the difficulty in detecting microscopic targets. RESULTS: We propose a deep multi-instance multi-label model, ImPLoc, to predict the subcellular locations from IHC images. In this model, we employ a deep convolutional neural network-based feature extractor to represent image features, and design a multi-head self-attention encoder to aggregate multiple feature vectors for subsequent prediction. We construct a benchmark dataset of 1186 proteins including 7855 images from HPA and 6 subcellular locations. The experimental results show that ImPLoc achieves significant enhancement on the prediction accuracy compared with the current computational methods. We further apply ImPLoc to a test set of 889 proteins with images from both normal and cancer tissues, and obtain 8 differentially localized proteins with a significance level of 0.05. AVAILABILITY AND IMPLEMENTATION: https://github.com/yl2019lw/ImPloc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Yang 0030, Hong-Bin Shen |
Bioinform. | 2 |
| 2020 | Artificial intelligence-based multi-objective optimization protocol for protein structure refinementabstractMOTIVATION: Protein structure refinement is an important step of protein structure prediction. Existing approaches have generally used a single scoring function combined with Monte Carlo method or Molecular Dynamics algorithm. The one-dimension optimization of a single energy function may take the structure too far away without a constraint. The basic motivation of our study is to reduce the bias problem caused by minimizing only a single energy function due to the very diversity of different protein structures. RESULTS: We report a new Artificial Intelligence-based protein structure Refinement method called AIR. Its fundamental idea is to use multiple energy functions as multi-objectives in an effort to correct the potential inaccuracy from a single function. A multi-objective particle swarm optimization algorithm-based structure refinement is designed, where each structure is considered as a particle in the protocol. With the refinement iterations, the particles move around. The quality of particles in each iteration is evaluated by three energy functions, and the non-dominated particles are put into a set called Pareto set. After enough iteration times, particles from the Pareto set are screened and part of the top solutions are outputted as the final refined structures. The multi-objective energy function optimization strategy designed in the AIR protocol provides a different constraint view of the structure, by extending the one-dimension optimization to a new three-dimension space optimization driven by the multi-objective particle swarm optimization engine. Experimental results on CASP11, CASP12 refinement targets and blind tests in CASP 13 turn to be promising. AVAILABILITY AND IMPLEMENTATION: The AIR is available online at: www.csbio.sjtu.edu.cn/bioinf/AIR/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ling Geng, Yu-Jun Zhao, Yang Yang 0030, Yang Zhang 0040, Hong-Bin Shen |
Bioinform. | 4 |
| 2019 | AnnoFly: annotating Drosophila embryonic images based on an attention-enhanced RNN modelabstractMOTIVATION: In the post-genomic era, image-based transcriptomics have received huge attention, because the visualization of gene expression distribution is able to reveal spatial and temporal expression pattern, which is significantly important for understanding biological mechanisms. The Berkeley Drosophila Genome Project has collected a large-scale spatial gene expression database for studying Drosophila embryogenesis. Given the expression images, how to annotate them for the study of Drosophila embryonic development is the next urgent task. In order to speed up the labor-intensive labeling work, automatic tools are highly desired. However, conventional image annotation tools are not applicable here, because the labeling is at the gene-level rather than the image-level, where each gene is represented by a bag of multiple related images, showing a multi-instance phenomenon, and the image quality varies by image orientations and experiment batches. Moreover, different local regions of an image correspond to different CV annotation terms, i.e. an image has multiple labels. Designing an accurate annotation tool in such a multi-instance multi-label scenario is a very challenging task. RESULTS: To address these challenges, we develop a new annotator for the fruit fly embryonic images, called AnnoFly. Driven by an attention-enhanced RNN model, it can weight images of different qualities, so as to focus on the most informative image patterns. We assess the new model on three standard datasets. The experimental results reveal that the attention-based model provides a transparent approach for identifying the important images for labeling, and it substantially enhances the accuracy compared with the existing annotation methods, including both single-instance and multi-instance learning methods. AVAILABILITY AND IMPLEMENTATION: http://www.csbio.sjtu.edu.cn/bioinf/annofly/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Yang 0030, Qingwei Fang, Hong-Bin Shen |
Bioinform. | 1 |
| 2019 | Predicting gene regulatory interactions based on spatial gene expression data and deep learningabstractReverse engineering of gene regulatory networks (GRNs) is a central task in systems biology. Most of the existing methods for GRN inference rely on gene co-expression analysis or TF-target binding information, where the determination of co-expression is often unreliable merely based on gene expression levels, and the TF-target binding data from high-throughput experiments may be noisy, leading to a high ratio of false links and missed links, especially for large-scale networks. In recent years, the microscopy images recording spatial gene expression have become a new resource in GRN reconstruction, as the spatial and temporal expression patterns contain much abundant gene interaction information. Till now, the spatial expression resources have been largely underexploited, and only a few traditional image processing methods have been employed in the image-based GRN reconstruction. Moreover, co-expression analysis using conventional measurements based on image similarity may be inaccurate, because it is the local-pattern consistency rather than global-image-similarity that determines gene-gene interactions. Here we present GripDL (Gene regulatory interaction prediction via Deep Learning), which incorporates high-confidence TF-gene regulation knowledge from previous studies, and constructs GRNs for Drosophila eye development based on Drosophila embryonic gene expression images. Benefitting from the powerful representation ability of deep neural networks and the supervision information of known interactions, the new method outperforms traditional methods with a large margin and reveals new intriguing knowledge about Drosophila eye development. Yang Yang 0030, Qingwei Fang, Hong-Bin Shen |
PLoS Comput. Biol. | 1 |
| 2018 | IterVM: An Iterative Model for Single-Particle Cryo-EM Image Clustering Based on Variational Autoencoder and Multi-Reference Alignment
Guowei Ji, Yang Yang 0030, Hong-Bin Shen |
BIBM | 2 |
| 2018 | HMIML: Hierarchical Multi-Instance Multi-Label Learning of Drosophila Embryogenesis Images Using Convolutional Neural Networks
Tiange Li, Yang Yang 0030, Hong-Bin Shen |
BIBM | 2 |
| 2018 | Prediction of MicroRNA Subcellular Localization by Using a Sequence-to-Sequence ModelabstractThe subcellular localization of microRNAs (miR-NAs) is closely related with their biological functions. Some recent studies have discovered that microRNAs can target to various cellular compartments, and have abundant localization patterns in cells. However, to the best of our knowledge, there has been no computational tool for predicting miRNA subcellular locations to date. The major reason is that the lack of useful information source largely limits the prediction performance using traditional statistical learning approaches. In this study, we regard this prediction task as a Sequence-to-Sequence learning process and propose an attention-based encoder-decoder model, miRLocator, to identify subcellular locations of human miRNAs. The designed miRLocator uses a bidirectional long short-term memory (BiLSTM) module to encode the input sequences, and an LSTM module to decode these context vectors as location sets. Especially, a new encoding method for RNAs, RNA2Vec, and an entropy-based method are incorporated in the model to determine the input and output representations, respectively. The experimental results show that miRLocator achieves promising prediction accuracy with the limited input information, and outperforms the models using hand-designed features and conventional RNN models. Yiqun Xiao, Jiaxun Cai, Yang Yang 0030, Hai Zhao 0001, Hong-Bin Shen |
ICDM | 3 |
| 2018 | Prediction of Type III Secreted Effectors Based on Word Embeddings for Protein Sequences
Xiaofeng Fu, Yiqun Xiao, Yang Yang 0030 |
ISBRA | 3 |
| 2018 | The lncLocator: a subcellular localization predictor for long non-coding RNAs based on a stacked ensemble classifierabstractMotivation: The long non-coding RNA (lncRNA) studies have been hot topics in the field of RNA biology. Recent studies have shown that their subcellular localizations carry important information for understanding their complex biological functions. Considering the costly and time-consuming experiments for identifying subcellular localization of lncRNAs, computational methods are urgently desired. However, to the best of our knowledge, there are no computational tools for predicting the lncRNA subcellular locations to date. Results: In this study, we report an ensemble classifier-based predictor, lncLocator, for predicting the lncRNA subcellular localizations. To fully exploit lncRNA sequence information, we adopt both k-mer features and high-level abstraction features generated by unsupervised deep models, and construct four classifiers by feeding these two types of features to support vector machine (SVM) and random forest (RF), respectively. Then we use a stacked ensemble strategy to combine the four classifiers and get the final prediction results. The current lncLocator can predict five subcellular localizations of lncRNAs, including cytoplasm, nucleus, cytosol, ribosome and exosome, and yield an overall accuracy of 0.59 on the constructed benchmark dataset. Availability and implementation: The lncLocator is available at www.csbio.sjtu.edu.cn/bioinf/lncLocator. Supplementary information: Supplementary data are available at Bioinformatics online. Xiaoyong Pan, Yang Yang 0030, Hong-Bin Shen |
Bioinform. | 3 |
| 2018 | MiRGOFS: a GO-based functional similarity measurement for miRNAs, with applications to the prediction of miRNA subcellular localization and miRNA-disease associationabstractMotivation: Benefiting from high-throughput experimental technologies, whole-genome analysis of microRNAs (miRNAs) has been more and more common to uncover important regulatory roles of miRNAs and identify miRNA biomarkers for disease diagnosis. As a complementary information to the high-throughput experimental data, domain knowledge like the Gene Ontology and KEGG pathway is usually used to guide gene function analysis. However, functional annotation for miRNAs is scarce in the public databases. Till now, only a few methods have been proposed for measuring the functional similarity between miRNAs based on public annotation data, and these methods cover a very limited number of miRNAs, which are not applicable to large-scale miRNA analysis. Results: In this paper, we propose a new method to measure the functional similarity for miRNAs, called miRGOFS, which has two notable features: (i) it adopts a new GO semantic similarity metric which considers both common ancestors and descendants of GO terms; (i) it computes similarity between GO sets in an asymmetric manner, and weights each GO term by its statistical significance. The miRGOFS-based predictor achieves an F1 of 61.2% on a benchmark dataset of miRNA localization, and AUC values of 87.7 and 81.1% on two benchmark sets of miRNA-disease association, respectively. Compared with the existing functional similarity measurements of miRNAs, miRGOFS has the advantages of higher accuracy and larger coverage of human miRNAs (over 1000 miRNAs). Availability and implementation: http://www.csbio.sjtu.edu.cn/bioinf/MiRGOFS/. Supplementary information: Supplementary data are available at Bioinformatics online. Yang Yang 0030, Xiaofeng Fu, Wenhao Qu, Yiqun Xiao, Hong-Bin Shen |
Bioinform. | 1 |
| 2017 | Hum-mPLoc 3.0: prediction enhancement of human protein subcellular localization through modeling the hidden correlations of gene ontology and functional domain featuresabstractMotivation: Protein subcellular localization prediction has been an important research topic in computational biology over the last decade. Various automatic methods have been proposed to predict locations for large scale protein datasets, where statistical machine learning algorithms are widely used for model construction. A key step in these predictors is encoding the amino acid sequences into feature vectors. Many studies have shown that features extracted from biological domains, such as gene ontology and functional domains, can be very useful for improving the prediction accuracy. However, domain knowledge usually results in redundant features and high-dimensional feature spaces, which may degenerate the performance of machine learning models. Results: In this paper, we propose a new amino acid sequence-based human protein subcellular location prediction approach Hum-mPLoc 3.0, which covers 12 human subcellular localizations. The sequences are represented by multi-view complementary features, i.e. context vocabulary annotation-based gene ontology (GO) terms, peptide-based functional domains, and residue-based statistical features. To systematically reflect the structural hierarchy of the domain knowledge bases, we propose a novel feature representation protocol denoted as HCM (Hidden Correlation Modeling), which will create more compact and discriminative feature vectors by modeling the hidden correlations between annotation terms. Experimental results on four benchmark datasets show that HCM improves prediction accuracy by 5-11% and F 1 by 8-19% compared with conventional GO-based methods. A large-scale application of Hum-mPLoc 3.0 on the whole human proteome reveals proteins co-localization preferences in the cell. Availability and Implementation: www.csbio.sjtu.edu.cn/bioinf/Hum-mPLoc3/. Contacts: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Yang Yang 0030, Hong-Bin Shen |
Bioinform. | 2 |
| 2016 | Feature selection based on functional group structure for microRNA expression data analysisabstractFeature selection methods have been widely used in gene expression analysis to identify differentially expressed genes and explore potential biomarkers for complex diseases. While a lot of studies have shown that incorporating feature structure information can greatly enhance the performance of feature selection algorithms, and genes naturally fall into groups with regard to common function and co-regulation, only a few of gene expression studies utilized the structured properties. And, as far as we know, there has been no such study on microRNA (miRNA) expression analysis due to the lack of available functional annotation for miRNAs. In this study, we focus on miRNA expression analysis because of its importance in the diagnosis, prognosis prediction and new therapeutic target detection for complex diseases. MiRNAs tend to work in groups to play their regulation roles, thus the miRNA expression data also has group structure. We utilize the GO-based semantic similarity to infer miRNA functional groups, and propose a new feature selection method taking group structure into consideration, called MiRFFS (MiRNA Functional group-based Feature Selection). We also apply the group information to the sparse group Lasso method, and compare MiRFFS with the sparse group Lasso as well as some existing feature selection methods. The results on three miRNA microarray profiles of breast cancer show that MiRFFS can achieve a compact feature subset with high classification accuracy. Yang Yang 0030 |
BIBM | 1 |
| 2016 | Missing value imputation for microRNA expression data by using a GO-based similarity measureabstractBACKGROUND: Missing values are commonly present in microarray data profiles. Instead of discarding genes or samples with incomplete expression level, missing values need to be properly imputed for accurate data analysis. The imputation methods can be roughly categorized as expression level-based and domain knowledge-based. The first type of methods only rely on expression data without the help of external data sources, while the second type incorporates available domain knowledge into expression data to improve imputation accuracy. In recent years, microRNA (miRNA) microarray has been largely developed and used for identifying miRNA biomarkers in complex human disease studies. Similar to mRNA profiles, miRNA expression profiles with missing values can be treated with the existing imputation methods. However, the domain knowledge-based methods are hard to be applied due to the lack of direct functional annotation for miRNAs. With the rapid accumulation of miRNA microarray data, it is increasingly needed to develop domain knowledge-based imputation algorithms specific to miRNA expression profiles to improve the quality of miRNA data analysis. RESULTS: We connect miRNAs with domain knowledge of Gene Ontology (GO) via their target genes, and define miRNA functional similarity based on the semantic similarity of GO terms in GO graphs. A new measure combining miRNA functional similarity and expression similarity is used in the imputation of missing values. The new measure is tested on two miRNA microarray datasets from breast cancer research and achieves improved performance compared with the expression-based method on both datasets. CONCLUSIONS: The experimental results demonstrate that the biological domain knowledge can benefit the estimation of missing values in miRNA profiles as well as mRNA profiles. Especially, functional similarity defined by GO terms annotated for the target genes of miRNAs can be useful complementary information for the expression-based method to improve the imputation accuracy of miRNA array data. Our method and data are available to the public upon request. Yang Yang 0030, Zhuangdi Xu |
BMC Bioinform. | 1 |
| 2015 | Deceptive Opinion Spam Detection Using Deep Level Linguistic FeaturesabstractThis paper focuses on improving a specific opinion spam detection task, deceptive spam. In addition to traditional word form and other shallow syntactic features, we introduce two types of deep level linguistic features. The first type of features are derived from a shallow discourse parser trained on Penn Discourse Treebank (PDTB), which can capture inter-sentence information. The second type is based on the relationship between sentiment analysis and spam detection. The experimental results over the benchmark dataset demonstrate that both of the proposed deep features achieve improved performance over the baseline. Changge Chen, Hai Zhao 0001, Yang Yang 0030 |
NLPCC | 3 |
| 2011 | Feature Reduction Using a Topic Model for the Prediction of Type III Secreted Effectors
Sihui Qi, Yang Yang 0030, Anjun Song |
ICONIP (1) | 2 |
| 2010 | Computational prediction of type III secreted proteins from gram-negative bacteriaabstractBACKGROUND: Type III secretion system (T3SS) is a specialized protein delivery system in gram-negative bacteria that injects proteins (called effectors) directly into the eukaryotic host cytosol and facilitates bacterial infection. For many plant and animal pathogens, T3SS is indispensable for disease development. Recently, T3SS has also been found in rhizobia and plays a crucial role in the nodulation process. Although a great deal of efforts have been done to understand type III secretion, the precise mechanism underlying the secretion and translocation process has not been fully understood. In particular, defined secretion and translocation signals enabling the secretion have not been identified from the type III secreted effectors (T3SEs), which makes the identification of these important virulence factors notoriously challenging. The availability of a large number of sequenced genomes for plant and animal-associated bacteria demands the development of efficient and effective prediction methods for the identification of T3SEs using bioinformatics approaches. RESULTS: We have developed a machine learning method based on the N-terminal amino acid sequences to predict novel type III effectors in the plant pathogen Pseudomonas syringae and the microsymbiont rhizobia. The extracted features used in the learning model (or classifier) include amino acid composition, secondary structure and solvent accessibility information. The method achieved a precision of over 90% on P. syringae in a cross validation study. In combination with a promoter screen for the type III specific promoters, this classifier trained on the P. syringae data was applied to predict novel T3SEs from the genomic sequences of four rhizobial strains. This application resulted in 57 candidate type III secreted proteins, 17 of which are confirmed effectors. CONCLUSION: Our experimental results demonstrate that the machine learning method based on N-terminal amino acid sequences combined with a promoter screen could prove to be a very effective computational approach for predicting novel type III effectors in gram-negative bacteria. Our method and data are available to the public upon request. Yang Yang 0030, Jiayuan Zhao, Robyn L. Morgan, Tao Jiang 0001 |
BMC Bioinform. | 1 |
| 2010 | Protein Subcellular Multi-Localization Prediction Using a Min-Max Modular Support Vector MachineabstractPrediction of protein subcellular localization is an important issue in computational biology because it provides important clues for the characterization of protein functions. Currently, much research has been dedicated to developing automatic prediction tools. Most, however, focus on mono-locational proteins, i.e., they assume that proteins exist in only one location. It should be noted that many proteins bear multi-locational characteristics and carry out crucial functions in biological processes. This work aims to develop a general pattern classifier for predicting multiple subcellular locations of proteins. We use an ensemble classifier, called the min-max modular support vector machine (M(3)-SVM), to solve protein subcellular multi-localization problems; and, propose a module decomposition method based on gene ontology (GO) semantic information for M(3)-SVM. The amino acid composition with secondary structure and solvent accessibility information is adopted to represent features of protein sequences. We apply our method to two multi-locational protein data sets. The M(3)-SVMs show higher accuracy and efficiency than traditional SVMs using the same feature vectors. And the GO decomposition also helps to improve prediction accuracy. Moreover, our method has a much higher rate of accuracy than existing subcellular localization predictors in predicting protein multi-localization. Yang Yang 0030, Bao-Liang Lu |
Int. J. Neural Syst. | 1 |
| 2009 | Computational prediction of novel non-coding RNAs in Arabidopsis thalianaabstractBACKGROUND: Non-coding RNA (ncRNA) genes do not encode proteins but produce functional RNA molecules that play crucial roles in many key biological processes. Recent genome-wide transcriptional profiling studies using tiling arrays in organisms such as human and Arabidopsis have revealed a great number of transcripts, a large portion of which have little or no capability to encode proteins. This unexpected finding suggests that the currently known repertoire of ncRNAs may only represent a small fraction of ncRNAs of the organisms. Thus, efficient and effective prediction of ncRNAs has become an important task in bioinformatics in recent years. Among the available computational methods, the comparative genomic approach seems to be the most powerful to detect ncRNAs. The recent completion of the sequencing of several major plant genomes has made the approach possible for plants. RESULTS: We have developed a pipeline to predict novel ncRNAs in the Arabidopsis (Arabidopsis thaliana) genome. It starts by comparing the expressed intergenic regions of Arabidopsis as provided in two whole-genome high-density oligo-probe arrays from the literature with the intergenic nucleotide sequences of all completely sequenced plant genomes including rice (Oryza sativa), poplar (Populus trichocarpa), grape (Vitis vinifera), and papaya (Carica papaya). By using multiple sequence alignment, a popular ncRNA prediction program (RNAz), wet-bench experimental validation, protein-coding potential analysis, and stringent screening against various ncRNA databases, the pipeline resulted in 16 families of novel ncRNAs (with a total of 21 ncRNAs). CONCLUSION: In this paper, we undertake a genome-wide search for novel ncRNAs in the genome of Arabidopsis by a comparative genomics approach. The identified novel ncRNAs are evolutionarily conserved between Arabidopsis and other recently sequenced plants, and may conduct interesting novel biological functions. Yang Yang 0030, Binglian Zheng, Zhidong Deng, Bao-Liang Lu, Tao Jiang 0001 |
BMC Bioinform. | 2 |
| 2008 | Classification of Protein Sequences Based on Word Segmentation Methods
Yang Yang 0030, Bao-Liang Lu, Wen-Yun Yang |
APBC | 1 |
| 2007 | Incorporating Domain Knowledge into a Min-Max Modular Support Vector Machine for Protein Subcellular Localization
Yang Yang 0030, Bao-Liang Lu |
ICONIP (2) | 1 |
| 2006 | A Comparative Study on Feature Extraction from Protein Sequences for Subcellular Localization PredictionabstractOne of the central problems in computational biology is to identify the protein function in an automated and high-throughput fashion. A key step in this process is to predict subcellular compartment the protein belongs to, since the protein localization closely correlates with its function. A wide variety of methods for protein subcellular localization has been proposed over recent years. They fall into two categories, sequence-based and database-based. The first one is to extract useful features from amino acid sequences and strives to discover the principles behind protein localization process. The second one is more apt to conduct data mining from existing public annotation databases. This paper focuses on the sequence-based approach and exploits the discriminative ability contained in amino acid sequences for protein subcellular localization. By using support vector machines (SVMs) as predictors, we conducted comparisons among amino acid composition approach, amino acid tuple approach, voting scheme, and a new characteristic representation of proteins proposed in this paper. Our experiments are carried out on 7579 eukaryotic protein sequences from 12 subcellular locations. The highest accuracy, 82.8% across 5-fold cross validation, is obtained by voting scheme using five predictors. This is the best performance achieved on this dataset using sequence-based approach. Our experiments demonstrate that there are considerable potentials on improving prediction accuracy by exploiting protein sequences, which have not been fully utilized so far, and more explorations are still needed in this direction Wen-Yun Yang, Bao-Liang Lu, Yang Yang 0030 |
CIBCB | 3 |
| 2006 | Prediction of Protein Subcellular Multi-locations with a Min-Max Modular Support Vector Machine
Yang Yang 0030, Bao-Liang Lu |
ISNN (2) | 1 |
| 2005 | Extracting Features from Protein Sequences Using Chinese Segmentation Techniques for Subcellular Localization
Yang Yang 0030, Bao-Liang Lu |
CIBCB | 1 |
| 2005 | Structure Pruning Strategies for Min-Max Modular Network
Yang Yang 0030, Bao-Liang Lu |
ISNN (1) | 1 |