EDBT 2026 Demo / reviewers in the wild / expert
Shanfeng Zhu
dblp:20/6446
· DBLP profile ↗
66ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0002-6067-5312ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 42 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 12 · 5 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Theory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Investigation of synonym expansion and self-alignment pretraining for enhancing Human Phenotype Ontology concept recognition
Weiqi Zhai, Rongze Jiang, Xiaodi Huang 0001, Junyi Bian, Shanfeng Zhu |
Artif. Intell. Medicine | 5 |
| 2026 | MoE-TCR: Mixture-of-experts framework for pan-specific TCR-epitope binding prediction
Shanfeng Zhu |
Pattern Recognit. | 5 |
| 2025 | DualGOFiller: A Dual-Channel Graph Neural Network with Contrastive Learning for Enhancing Function Prediction in Partially Annotated Proteins
Hancheng Liu, Weiqi Zhai, Shanfeng Zhu |
RECOMB | 4 |
| 2025 | ContraDTI: Improved drug-target interaction prediction via multi-view contrastive learning
Zhirui Liao, Shanfeng Zhu |
Artif. Intell. Medicine | 3 |
| 2025 | GOAnnotator: accurate protein function annotation using automatically retrieved literatureabstractSUMMARY: Automated protein function prediction/annotation (AFP) is vital for understanding biological processes and advancing biomedical research. Existing text-based AFP methods including the state-of-the-art method, GORetriever, rely on expert-curated relevant literature, which is costly and time-consuming, and cover only a small portion of the proteins in UniProt. To overcome this limitation, we propose GOAnnotator, a novel framework for automated protein function annotation. It consists of two key modules: PubRetriever, a hybrid system for retrieving and re-ranking relevant literature, and GORetriever+, an enhanced module for identifying Gene Ontology (GO) terms from the retrieved texts. Extensive experiments over three benchmark datasets demonstrate that GOAnnotator delivers high-quality functional annotations, surpassing GORetriever in realistic situations by uncovering unique literature and predicting additional functions. These results highlight its great potential to streamline and enhance annotation of protein functions without relying on manual curation. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/ZhuLab-Fudan/GOAnnotator. Huiying Yan, Hancheng Liu, Shanfeng Zhu |
Bioinform. | 4 |
| 2025 | Constrained multi-scale dense connections for biomedical image segmentation
Yanchun Zhang, Hailong Qiu, Xiaomeng Li 0001, Shanfeng Zhu, Meiping Huang, Jian Zhuang, Yiyu Shi 0001, Xiaowei Xu 0004 |
Pattern Recognit. | 6 |
| 2025 | MultiFusion2HPO: A Multimodal Deep Learning Approach for Enhancing Human Protein-Phenotype Association PredictionabstractAccurately identifying associations between human genes (proteins) and clinical phenotypes is critical for advancing drug development and precision medicine. While the human phenotype ontology (HPO) standardizes clinical phenotypes, current computational approaches for predicting human protein-phenotype associations suffer from two limitations: (1) underutilization of multimodal protein-related information and (2) lack of state-of-the-art deep learning representations tailored to diverse data modalities, such as text and sequence. To overcome these limitations, we introduce MultiFusion2HPO, a novel multimodal model that integrates diverse features and advanced learning methods from multiple data sources to enhance the prediction of human protein-HPO associations. MultiFusion2HPO leverages five critical modalities: textual information (TFIDF-D2V and BioLinkBERT embeddings), protein sequence data (InterPro and ESM2), protein-protein interaction (PPI) networks, gene ontology (GO) annotation, and gene expression. Comprehensive experiments on benchmark datasets demonstrate the superiority of MultiFusion2HPO over the state-of-the-art methods, DeepPheno and HPOLabeler. These results underscore the effectiveness of integrating multimodal protein data to improve the accuracy of human protein-HPO association predictions. Weiqi Zhai, Yongjun Deng, Xiaodi Huang 0001, Shanfeng Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | PubLabeler: Enhancing Automatic Classification of Publications in UniProtKB Using Protein Textual Description and PubMedBERTabstractIn UniProtKB, each protein is linked to numerous publications covering topics such as sequence, function, and structure, which are annotated manually or through automated methods. Given the vast number of proteins and literature, manual annotation is time-consuming and labour-intensive. Although UniProtKB offers automated annotations, their quality often falls short. Therefore, developing an accurate automated classifier to identify the topics of publications associated with each protein is imperative for advancing biomedical knowledge discovery. Classifying publications in UniProtKB involves protein-publication pairs characterized by multi-label, label co-occurrence, and class imbalance, which increases complexity. This paper proposes a novel method called PubLabeler, which simultaneously considers protein description and scientific literature texts as input. PubLabeler employs the PubMedBERT model to encode input texts and integrates label co-occurrence information into the model parameters. Additionally, it uses focal loss to update parameters, allowing the model to focus more on classes with a few instances. Using newly annotated literature from Swiss-Prot in 2023 as a test set, PubLabeler achieved superior results in both micro and macro metrics, showing a 28.5% improvement in macro-F1 compared to UniProtKB's automated annotation method, UPCLASS. Furthermore, we validated PubLabeler's effectiveness in TrEMBL annotation, showcasing its comprehensive prediction results compared to TrEMBL's automated annotations. These findings highlight PubLabeler's reliability and potential to advance protein-related information extraction and knowledge discovery. Junyi Bian, Xiaodi Huang 0001, Shanfeng Zhu |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | DMNER: Biomedical Named Entity Recognition by Detection and MatchingabstractBiomedical Named Entity Recognition (NER) is a crucial task in extracting information from biomedical texts. However, the diversity of professional terminology, semantic complexity, and the widespread presence of synonyms pose significant challenges. Traditional methods that rely on sequence labeling training datasets often struggle to handle these complexities. To address this, we introduce a novel framework for BioNER, termed DMNER, which leverages external knowledge and operates in two steps: entity boundary detection and entity category identification through matching. The core of DMNER is its second step, which determines the entity category by retrieving similar entities and their categories from a knowledge dictionary using semantic similarity matching. Our experiments on 10 biomedical datasets demonstrate that DMNER outperforms baselines across these tasks, proving its effectiveness and adaptability. DMNER is versatile and can be applied to various NER tasks, including supervised NER, distantly supervised NER, and NER on multiple datasets with disjoint label sets. The DMNER code is publicly available1. Junyi Bian, Rongze Jiang, Weiqi Zhai, Tianyang Huang, Xiaodi Huang 0001, Shanfeng Zhu |
BIBM | 7 |
| 2024 | VANER: Leveraging Large Language Model for Versatile and Adaptive Biomedical Named Entity RecognitionabstractThe prevalent solution for BioNER involves using representation learning techniques combined with sequence labeling. However, such methods are inherently task-specific, demonstrate poor generalizability, and often require a dedicated model for each dataset. To leverage the versatile capabilities of recent large language models (LLMs), several approaches have explored generative techniques for entity extraction. Yet, these approaches often fall short compared to previous sequence labeling approaches. In this paper, we utilize the open-sourced LLM LLaMA2 as the backbone model, and design specific instructions to distinguish between different types of entities and datasets. By combining the LLM’s understanding of instructions with sequence labeling techniques, we train a model using a mix of datasets capable of extracting various types of entities. Given that the backbone LLMs lacks specialized medical knowledge, we also integrate external entity knowledge bases and employ instruction tuning to enable the model to densely recognize curated entities. Our parameter-efficient training model, VANER, significantly outperforms previous LLMs-based models. For the first time, as an LLM-based model, VANER surpasses the majority of conventional state-of-the-art BioNER systems, achieving the highest F1 scores across three datasets. Junyi Bian, Weiqi Zhai, Xiaodi Huang 0001, Jiaxuan Zheng, Shanfeng Zhu |
ECAI | 5 |
| 2024 | GORetriever: reranking protein-description-based GO candidates by literature-driven deep information retrieval for protein function annotationabstractSUMMARY: The vast majority of proteins still lack experimentally validated functional annotations, which highlights the importance of developing high-performance automated protein function prediction/annotation (AFP) methods. While existing approaches focus on protein sequences, networks, and structural data, textual information related to proteins has been overlooked. However, roughly 82% of SwissProt proteins already possess literature information that experts have annotated. To efficiently and effectively use literature information, we present GORetriever, a two-stage deep information retrieval-based method for AFP. Given a target protein, in the first stage, candidate Gene Ontology (GO) terms are retrieved by using annotated proteins with similar descriptions. In the second stage, the GO terms are reranked based on semantic matching between the GO definitions and textual information (literature and protein description) of the target protein. Extensive experiments over benchmark datasets demonstrate the remarkable effectiveness of GORetriever in enhancing the AFP performance. Note that GORetriever is the key component of GOCurator, which has achieved first place in the latest critical assessment of protein function annotation (CAFA5: over 1600 teams participated), held in 2023-2024. AVAILABILITY AND IMPLEMENTATION: GORetriever is publicly available at https://github.com/ZhuLab-Fudan/GORetriever. Huiying Yan, Hancheng Liu, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 5 |
| 2024 | GoSum: extractive summarization of long documents by reinforcement learning and graph-organized discourse state
Junyi Bian, Xiaodi Huang 0001, Tianyang Huang, Shanfeng Zhu |
Knowl. Inf. Syst. | 5 |
| 2023 | Phen2Disease: a phenotype-driven model for disease and gene prioritization by bidirectional maximum matching semantic similaritiesabstractHuman Phenotype Ontology (HPO)-based approaches have gained popularity in recent times as a tool for genomic diagnostics of rare diseases. However, these approaches do not make full use of the available information on disease and patient phenotypes. We present a new method called Phen2Disease, which utilizes the bidirectional maximum matching semantic similarity between two phenotype sets of patients and diseases to prioritize diseases and genes. Our comprehensive experiments have been conducted on six real data cohorts with 2051 cases (Cohort 1, n = 384; Cohort 2, n = 281; Cohort 3, n = 185; Cohort 4, n = 784; Cohort 5, n = 208; and Cohort 6, n = 209) and two simulated data cohorts with 1000 cases. The results of the experiments showed that Phen2Disease outperforms the three state-of-the-art methods when only phenotype information and HPO knowledge base are used, particularly in cohorts with fewer average numbers of HPO terms. We also observed that patients with higher information content scores have more specific information, leading to more accurate predictions. Moreover, Phen2Disease provides high interpretability with ranked diseases and patient HPO terms presented. Our method provides a novel approach to utilizing phenotype data for genomic diagnostics of rare diseases, with potential for clinical impact. Phen2Disease is freely available on GitHub at https://github.com/ZhuLab-Fudan/Phen2Disease. Weiqi Zhai, Xiaodi Huang 0001, Nan Shen, Shanfeng Zhu |
Briefings Bioinform. | 4 |
| 2023 | Sc2Mol: a scaffold-based two-step molecule generator with variational autoencoder and transformerabstractMOTIVATION: Finding molecules with desired pharmaceutical properties is crucial in drug discovery. Generative models can be an efficient tool to find desired molecules through the distribution learned by the model to approximate given training data. Existing generative models (i) do not consider backbone structures (scaffolds), resulting in inefficiency or (ii) need prior patterns for scaffolds, causing bias. Scaffolds are reasonable to use, and it is imperative to design a generative model without any prior scaffold patterns. RESULTS: We propose a generative model-based molecule generator, Sc2Mol, without any prior scaffold patterns. Sc2Mol uses SMILES strings for molecules. It consists of two steps: scaffold generation and scaffold decoration, which are carried out by a variational autoencoder and a transformer, respectively. The two steps are powerful for implementing random molecule generation and scaffold optimization. Our empirical evaluation using drug-like molecule datasets confirmed the success of our model in distribution learning and molecule optimization. Also, our model could automatically learn the rules to transform coarse scaffolds into sophisticated drug candidates. These rules were consistent with those for current lead optimization. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/zhiruiliao/Sc2Mol. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhirui Liao, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 4 |
| 2023 | DeepMHCI: an anchor position-aware deep interaction model for accurate MHC-I peptide binding affinity predictionabstractMOTIVATION: Computationally predicting major histocompatibility complex class I (MHC-I) peptide binding affinity is an important problem in immunological bioinformatics, which is also crucial for the identification of neoantigens for personalized therapeutic cancer vaccines. Recent cutting-edge deep learning-based methods for this problem cannot achieve satisfactory performance, especially for non-9-mer peptides. This is because such methods generate the input by simply concatenating the two given sequences: a peptide and (the pseudo sequence of) an MHC class I molecule, which cannot precisely capture the anchor positions of the MHC binding motif for the peptides with variable lengths. We thus developed an anchor position-aware and high-performance deep model, DeepMHCI, with a position-wise gated layer and a residual binding interaction convolution layer. This allows the model to control the information flow in peptides to be aware of anchor positions and model the interactions between peptides and the MHC pseudo (binding) sequence directly with multiple convolutional kernels. RESULTS: The performance of DeepMHCI has been thoroughly validated by extensive experiments on four benchmark datasets under various settings, such as 5-fold cross-validation, validation with the independent testing set, external HPV vaccine identification, and external CD8+ epitope identification. Experimental results with visualization of binding motifs demonstrate that DeepMHCI outperformed all competing methods, especially on non-9-mer peptides binding prediction. AVAILABILITY AND IMPLEMENTATION: DeepMHCI is publicly available at https://github.com/ZhuLab-Fudan/DeepMHCI. Ronghui You, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 4 |
| 2022 | META-DDIE: predicting drug-drug interaction events with few-shot learningabstractDrug-drug interactions (DDIs) are one of the major concerns in pharmaceutical research, and a number of computational methods have been developed to predict whether two drugs interact or not. Recently, more attention has been paid to events caused by the DDIs, which is more useful for investigating the mechanism hidden behind the combined drug usage or adverse reactions. However, some rare events may only have few examples, hindering them from being precisely predicted. To address the above issues, we present a few-shot computational method named META-DDIE, which consists of a representation module and a comparing module, to predict DDI events. We collect drug chemical structures and DDIs from DrugBank, and categorize DDI events into hundreds of types using a standard pipeline. META-DDIE uses the structures of drugs as input and learns the interpretable representations of DDIs through the representation module. Then, the model uses the comparing module to predict whether two representations are similar, and finally predicts DDI events with few labeled examples. In the computational experiments, META-DDIE outperforms several baseline methods and especially enhances the predictive capability for rare events. Moreover, META-DDIE helps to identify the key factors that may cause DDI events and reveal the relationship among different events. Xinran Xu, Shichao Liu 0002, Zhongfei Zhang, Shanfeng Zhu, Wen Zhang 0008 |
Briefings Bioinform. | 6 |
| 2022 | HPODNets: deep graph convolutional networks for predicting human protein-phenotype associationsabstractMOTIVATION: Deciphering the relationship between human genes/proteins and abnormal phenotypes is of great importance in the prevention, diagnosis and treatment against diseases. The Human Phenotype Ontology (HPO) is a standardized vocabulary that describes the phenotype abnormalities encountered in human disorders. However, the current HPO annotations are still incomplete. Thus, it is necessary to computationally predict human protein-phenotype associations. In terms of current, cutting-edge computational methods for annotating proteins (such as functional annotation), three important features are (i) multiple network input, (ii) semi-supervised learning and (iii) deep graph convolutional network (GCN), whereas there are no methods with all these features for predicting HPO annotations of human protein. RESULTS: We develop HPODNets with all above three features for predicting human protein-phenotype associations. HPODNets adopts a deep GCN with eight layers which allows to capture high-order topological information from multiple interaction networks. Empirical results with both cross-validation and temporal validation demonstrate that HPODNets outperforms seven competing state-of-the-art methods for protein function prediction. HPODNets with the architecture of deep GCNs is confirmed to be effective for predicting HPO annotations of human protein and, more generally, node label ranking problem with multiple biomolecular networks input in bioinformatics. AVAILABILITY AND IMPLEMENTATION: https://github.com/liulizhi1996/HPODNets. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 3 |
| 2022 | DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity predictionabstractMOTIVATION: Computationally predicting major histocompatibility complex (MHC)-peptide binding affinity is an important problem in immunological bioinformatics. Recent cutting-edge deep learning-based methods for this problem are unable to achieve satisfactory performance for MHC class II molecules. This is because such methods generate the input by simply concatenating the two given sequences: (the estimated binding core of) a peptide and (the pseudo sequence of) an MHC class II molecule, ignoring biological knowledge behind the interactions of the two molecules. We thus propose a binding core-aware deep learning-based model, DeepMHCII, with a binding interaction convolution layer, which allows to integrate all potential binding cores (in a given peptide) with the MHC pseudo (binding) sequence, through modeling the interaction with multiple convolutional kernels. RESULTS: Extensive empirical experiments with four large-scale datasets demonstrate that DeepMHCII significantly outperformed four state-of-the-art methods under numerous settings, such as 5-fold cross-validation, leave one molecule out, validation with independent testing sets and binding core prediction. All these results and visualization of the predicted binding cores indicate the effectiveness of our model, DeepMHCII, and the importance of properly modeling biological facts in deep learning for high predictive performance and efficient knowledge discovery. AVAILABILITY AND IMPLEMENTATION: DeepMHCII is publicly available at https://github.com/yourh/DeepMHCII. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronghui You, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 4 |
| 2021 | Drug3D-DTI: Improved Drug-target Interaction Prediction by Incorporating Spatial Information of Small MoleculesabstractA number of machine learning (ML) approaches for drug discovery have been available that rely only on sequential (1D) and planar (2D) information without effectively using the 3D information for generating features of drugs. However, 3D information of small molecules can reflect relative position of atoms more directly, which affects molecular properties. In this work, we present a new deep learning model called Drug3D-DTI for drug-target interaction prediction. Drug3D-DTI takes advantage of molecular spatial information, i.e., atom proximity in three-dimensional (3D) structures. We comprehensively evaluated the performance of Drug3D-DTI on two datasets with two tasks of regression and classification. In particular, we compared Drug3D-DTI with several existing methods including the two cutting-edge methods for compound-protein interaction prediction. From the experimental results, Drug3D-DTI clearly outperformed other methods under all settings. Further, this performance improvement was validated by ablation experiments and a case study. The implementation of Drug3D-DTI is available at (https://github.com/zhiruiliao/Drug3D-DTI). Zhirui Liao, Xiaodi Huang 0001, Hiroshi Mamitsuka, Shanfeng Zhu |
BIBM | 4 |
| 2021 | Corrigendum to: Assessment of metagenomic assemblers based on hybrid reads of real and simulated metagenomic sequencesabstractThe first version of this article did not include Shanfeng Zhu's funding from Project 111. This has now been corrected. The authors regret the error. Ying Wang 0005, Jed A. Fuhrman, Fengzhu Sun, Shanfeng Zhu |
Briefings Bioinform. | 5 |
| 2021 | HPOFiller: identifying missing protein-phenotype associations by graph convolutional networkabstractMOTIVATION: Exploring the relationship between human proteins and abnormal phenotypes is of great importance in the prevention, diagnosis and treatment of diseases. The human phenotype ontology (HPO) is a standardized vocabulary that describes the phenotype abnormalities encountered in human diseases. However, the current HPO annotations of proteins are not complete. Thus, it is important to identify missing protein-phenotype associations. RESULTS: We propose HPOFiller, a graph convolutional network (GCN)-based approach, for predicting missing HPO annotations. HPOFiller has two key GCN components for capturing embeddings from complex network structures: (i) S-GCN for both protein-protein interaction network and HPO semantic similarity network to utilize network weights; (ii) Bi-GCN for the protein-phenotype bipartite graph to conduct message passing between proteins and phenotypes. The core idea of HPOFiller is to repeat run these two GCN modules consecutively over the three networks, to refine the embeddings. Empirical results of extremely stringent evaluation avoiding potential information leakage including cross-validation and temporal validation demonstrates that HPOFiller significantly outperforms all other state-of-the-art methods. In particular, the ablation study shows that batch normalization contributes the most to the performance. The further examination offers literature evidence for highly ranked predictions. Finally using known disease-HPO term associations, HPOFiller could suggest promising, unknown disease-gene associations, presenting possible genetic causes of human disorders. AVAILABILITYAND IMPLEMENTATION: https://github.com/liulizhi1996/HPOFiller. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 3 |
| 2021 | BERTMeSH: deep contextual representation learning for large-scale high-performance MeSH indexing with full textabstractMOTIVATION: With the rapid increase of biomedical articles, large-scale automatic Medical Subject Headings (MeSH) indexing has become increasingly important. FullMeSH, the only method for large-scale MeSH indexing with full text, suffers from three major drawbacks: FullMeSH (i) uses Learning To Rank, which is time-consuming, (ii) can capture some pre-defined sections only in full text and (iii) ignores the whole MEDLINE database. RESULTS: We propose a computationally lighter, full text and deep-learning-based MeSH indexing method, BERTMeSH, which is flexible for section organization in full text. BERTMeSH has two technologies: (i) the state-of-the-art pre-trained deep contextual representation, Bidirectional Encoder Representations from Transformers (BERT), which makes BERTMeSH capture deep semantics of full text. (ii) A transfer learning strategy for using both full text in PubMed Central (PMC) and title and abstract (only and no full text) in MEDLINE, to take advantages of both. In our experiments, BERTMeSH was pre-trained with 3 million MEDLINE citations and trained on ∼1.5 million full texts in PMC. BERTMeSH outperformed various cutting-edge baselines. For example, for 20 K test articles of PMC, BERTMeSH achieved a Micro F-measure of 69.2%, which was 6.3% higher than FullMeSH with the difference being statistically significant. Also prediction of 20 K test articles needed 5 min by BERTMeSH, while it took more than 10 h by FullMeSH, proving the computational efficiency of BERTMeSH. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronghui You, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 4 |
| 2021 | DeepGraphGO: graph neural network for large-scale, multispecies protein function predictionabstractMOTIVATION: Automated function prediction (AFP) of proteins is a large-scale multi-label classification problem. Two limitations of most network-based methods for AFP are (i) a single model must be trained for each species and (ii) protein sequence information is totally ignored. These limitations cause weaker performance than sequence-based methods. Thus, the challenge is how to develop a powerful network-based method for AFP to overcome these limitations. RESULTS: We propose DeepGraphGO, an end-to-end, multispecies graph neural network-based method for AFP, which makes the most of both protein sequence and high-order protein network information. Our multispecies strategy allows one single model to be trained for all species, indicating a larger number of training samples than existing methods. Extensive experiments with a large-scale dataset show that DeepGraphGO outperforms a number of competing state-of-the-art methods significantly, including DeepGOPlus and three representative network-based methods: GeneMANIA, deepNF and clusDCA. We further confirm the effectiveness of our multispecies strategy and the advantage of DeepGraphGO over so-called difficult proteins. Finally, we integrate DeepGraphGO into the state-of-the-art ensemble method, NetGO, as a component and achieve a further performance improvement. AVAILABILITY AND IMPLEMENTATION: https://github.com/yourh/DeepGraphGO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ronghui You, Shuwei Yao, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 4 |
| 2021 | GrantExtractor: Accurate Grant Support Information Extraction from Biomedical Fulltext Based on Bi-LSTM-CRFabstractGrant support (GS) in the MEDLINE database refers to funding agencies and contract numbers. It is important for funding organizations to track their funding outcomes from the GS information. As such, how to accurately and automatically extract funding information from biomedical literature is challenging. In this paper, we present a pipeline system called GrantExtractor that is able to accurately extract GS information from fulltext biomedical literature. GrantExtractor effectively integrates several advanced machine learning techniques. In particular, we use a sentence classifier to identify funding sentences from articles first. A bi-directional LSTM and the CRF layer (BiLSTM-CRF), and pattern matching are then used to extract entities of grant numbers and agencies from these identified funding sentences. After removing noisy numbers by a multi-class model, we finally match each grant number with its corresponding agency. Experimental results on benchmark datasets have demonstrated that GrantExtractor clearly outperforms all baseline methods. It is further evident that GrantExtractor won the first place in Task 5C of 2017 BioASQ challenge, with achieving the Micro-recall of 0.9526 for 22,610 articles. Moreover, GrantExtractor has achieved the Micro F-measure score as high as 0.90 in extracting grant pairs. Suyang Dai, Yuxia Ding, Wenxuan Zuo, Xiaodi Huang 0001, Shanfeng Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2020 | Constrained Multi-scale Dense Connections for Accurate Biomedical Image SegmentationabstractBiomedical image segmentation plays a critical role in clinical diagnosis and medical intervention. Recently, a variety of deep neural networks have boosted the biomedical image segmentation performance with a large margin, which adopts dense connections to explore rich representations in multiple scales. In multi-scale dense connections, features from all or most scales are fused or iteratively aggregated. In this paper, we propose constrained multi-scale dense connections (CMDC) for accurate biomedical image segmentation, which only fuse features from the nearest scales containing the most relevant appearance or semantic information. Based on CMDC, we further construct constraint multi-scale dense networks (CMD-Net) by applying CMDC to existing segmentation networks. Experiments across various architectures (including FCN-8s, U-Net, and DeepLabV3) and datasets (including GlaS, CRAG, KID, and ECS) demonstrate that CMD-Net not only outperforms existing schemes on both accuracy and efficiency but also can be easily generalized to a variety of segmentation networks. In addition, CMD-Net achieves state-of-the-art performance on two instance segmentation datasets, GlaS and CRAG. Yanchun Zhang, Shanfeng Zhu, Xiaowei Xu 0004 |
BIBM | 3 |
| 2020 | Assessment of metagenomic assemblers based on hybrid reads of real and simulated metagenomic sequencesabstractIn metagenomic studies of microbial communities, the short reads come from mixtures of genomes. Read assembly is usually an essential first step for the follow-up studies in metagenomic research. Understanding the power and limitations of various read assembly programs in practice is important for researchers to choose which programs to use in their investigations. Many studies evaluating different assembly programs used either simulated metagenomes or real metagenomes with unknown genome compositions. However, the simulated datasets may not reflect the real complexities of metagenomic samples and the estimated assembly accuracy could be misleading due to the unknown genomes in real metagenomes. Therefore, hybrid strategies are required to evaluate the various read assemblers for metagenomic studies. In this paper, we benchmark the metagenomic read assemblers by mixing reads from real metagenomic datasets with reads from known genomes and evaluating the integrity, contiguity and accuracy of the assembly using the reads from the known genomes. We selected four advanced metagenome assemblers, MEGAHIT, MetaSPAdes, IDBA-UD and Faucet, for evaluation. We showed the strengths and weaknesses of these assemblers in terms of integrity, contiguity and accuracy for different variables, including the genetic difference of the real genomes with the genome sequences in the real metagenomic datasets and the sequencing depth of the simulated datasets. Overall, MetaSPAdes performs best in terms of integrity and continuity at the species-level, followed by MEGAHIT. Faucet performs best in terms of accuracy at the cost of worst integrity and continuity, especially at low sequencing depth. MEGAHIT has the highest genome fractions at the strain-level and MetaSPAdes has the overall best performance at the strain-level. MEGAHIT is the most efficient in our experiments. Availability: The source code is available at https://github.com/ziyewang/MetaAssemblyEval. Ying Wang 0005, Jed A. Fuhrman, Fengzhu Sun, Shanfeng Zhu |
Briefings Bioinform. | 5 |
| 2020 | FullMeSH: improving large-scale MeSH indexing with full textabstractMOTIVATION: With the rapidly growing biomedical literature, automatically indexing biomedical articles by Medical Subject Heading (MeSH), namely MeSH indexing, has become increasingly important for facilitating hypothesis generation and knowledge discovery. Over the past years, many large-scale MeSH indexing approaches have been proposed, such as Medical Text Indexer, MeSHLabeler, DeepMeSH and MeSHProbeNet. However, the performance of these methods is hampered by using limited information, i.e. only the title and abstract of biomedical articles. RESULTS: We propose FullMeSH, a large-scale MeSH indexing method taking advantage of the recent increase in the availability of full text articles. Compared to DeepMeSH and other state-of-the-art methods, FullMeSH has three novelties: (i) Instead of using a full text as a whole, FullMeSH segments it into several sections with their normalized titles in order to distinguish their contributions to the overall performance. (ii) FullMeSH integrates the evidence from different sections in a 'learning to rank' framework by combining the sparse and deep semantic representations. (iii) FullMeSH trains an Attention-based Convolutional Neural Network for each section, which achieves better performance on infrequent MeSH headings. FullMeSH has been developed and empirically trained on the entire set of 1.4 million full-text articles in the PubMed Central Open Access subset. It achieved a Micro F-measure of 66.76% on a test set of 10 000 articles, which was 3.3% and 6.4% higher than DeepMeSH and MeSHLabeler, respectively. Furthermore, FullMeSH demonstrated an average improvement of 4.7% over DeepMeSH for indexing Check Tags, a set of most frequently indexed MeSH headings. AVAILABILITY AND IMPLEMENTATION: The software is available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Suyang Dai, Ronghui You, Zhiyong Lu, Xiaodi Huang 0001, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 6 |
| 2020 | HPOLabeler: improving prediction of human protein-phenotype associations by learning to rankabstractMOTIVATION: Annotating human proteins by abnormal phenotypes has become an important topic. Human Phenotype Ontology (HPO) is a standardized vocabulary of phenotypic abnormalities encountered in human diseases. As of November 2019, only <4000 proteins have been annotated with HPO. Thus, a computational approach for accurately predicting protein-HPO associations would be important, whereas no methods have outperformed a simple Naive approach in the second Critical Assessment of Functional Annotation, 2013-2014 (CAFA2). RESULTS: We present HPOLabeler, which is able to use a wide variety of evidence, such as protein-protein interaction (PPI) networks, Gene Ontology, InterPro, trigram frequency and HPO term frequency, in the framework of learning to rank (LTR). LTR has been proved to be powerful for solving large-scale, multi-label ranking problems in bioinformatics. Given an input protein, LTR outputs the ranked list of HPO terms from a series of input scores given to the candidate HPO terms by component learning models (logistic regression, nearest neighbor and a Naive method), which are trained from given multiple evidence. We empirically evaluate HPOLabeler extensively through mainly two experiments of cross validation and temporal validation, for which HPOLabeler significantly outperformed all component models and competing methods including the current state-of-the-art method. We further found that (i) PPI is most informative for prediction among diverse data sources and (ii) low prediction performance of temporal validation might be caused by incomplete annotation of new proteins. AVAILABILITY AND IMPLEMENTATION: http://issubmission.sjtu.edu.cn/hpolabeler/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaodi Huang 0001, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 4 |
| 2019 | DeepDock: Enhancing Ligand-protein Interaction Prediction by a Combination of Ligand and Structure InformationabstractThe prediction of precise protein-ligand binding activities can accelerate drug discovery by virtual screening-a computational technique that predicts whether a small molecule ligand is able to bind to a specific target. Thus, it is crucial to improve the performance of virtual screening. However, previous models for solving this problem are either ligand-based or structure-based. In this paper, we propose a universal deep neural network model called DeepDock that predicts protein-ligand interaction by using both ligand and structure information. Using the combination of two types of information, our model consists of embedding, convolution, max pooling, and fully-connected layers. In particular, different types of inputs are concatenated before being fed into the fully-connected layers. In the experiments, we compare our approach to the competing methods against two benchmark datasets under different settings. The experiment results have demonstrated that DeepDock can improve predictive performance by more than 4% on both DUD-E and MUV datasets in terms of AUPR. Zhirui Liao, Ronghui You, Xiaodi Huang 0001, Shanfeng Zhu |
BIBM | 6 |
| 2019 | AttentionXML: Label Tree-based Attention-Aware Deep Model for High-Performance Extreme Multi-Label Text ClassificationabstractExtreme multi-label text classification (XMTC) is an important problem in the era of {\it big data}, for tagging a given text with the most relevant multiple labels from an extremely large-scale label set. XMTC can be found in many applications, such as item categorization, web page tagging, and news annotation. Traditionally most methods used bag-of-words (BOW) as inputs, ignoring word context as well as deep semantic information. Recent attempts to overcome the problems of BOW by deep learning still suffer from 1) failing to capture the important subtext for each label and 2) lack of scalability against the huge number of labels. We propose a new label tree-based deep learning model for XMTC, called AttentionXML, with two unique features: 1) a multi-label attention mechanism with raw text as input, which allows to capture the most relevant part of text to each label; and 2) a shallow and wide probabilistic label tree (PLT), which allows to handle millions of labels, especially for "tail labels". We empirically compared the performance of AttentionXML with those of eight state-of-the-art methods over six benchmark datasets, including Amazon-3M with around 3 million labels. AttentionXML outperformed all competing methods under all experimental settings. Experimental results also show that AttentionXML achieved the best performance against tail labels among label tree-based methods. The code and datasets are available at \url{http://github.com/yourh/AttentionXML} . Ronghui You, Suyang Dai, Hiroshi Mamitsuka, Shanfeng Zhu |
NeurIPS | 6 |
| 2019 | SolidBin: improving metagenome binning with semi-supervised normalized cutabstractMOTIVATION: Metagenomic contig binning is an important computational problem in metagenomic research, which aims to cluster contigs from the same genome into the same group. Unlike classical clustering problem, contig binning can utilize known relationships among some of the contigs or the taxonomic identity of some contigs. However, the current state-of-the-art contig binning methods do not make full use of the additional biological information except the coverage and sequence composition of the contigs. RESULTS: We developed a novel contig binning method, Semi-supervised Spectral Normalized Cut for Binning (SolidBin), based on semi-supervised spectral clustering. Using sequence feature similarity and/or additional biological information, such as the reliable taxonomy assignments of some contigs, SolidBin constructs two types of prior information: must-link and cannot-link constraints. Must-link constraints mean that the pair of contigs should be clustered into the same group, while cannot-link constraints mean that the pair of contigs should be clustered in different groups. These constraints are then integrated into a classical spectral clustering approach, normalized cut, for improved contig binning. The performance of SolidBin is compared with five state-of-the-art genome binners, CONCOCT, COCACOLA, MaxBin, MetaBAT and BMC3C on five next-generation sequencing benchmark datasets including simulated multi- and single-sample datasets and real multi-sample datasets. The experimental results show that, SolidBin has achieved the best performance in terms of F-score, Adjusted Rand Index and Normalized Mutual Information, especially while using the real datasets and the single-sample dataset. AVAILABILITY AND IMPLEMENTATION: https://github.com/sufforest/SolidBin. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Young Lu, Fengzhu Sun, Shanfeng Zhu |
Bioinform. | 5 |
| 2018 | GrantExtractor: A Winning System for Extracting Grant Support Information from Biomedical Literature
Suyang Dai, Wenxuan Zuo, Xiaodi Huang 0001, Shanfeng Zhu |
BIBM | 5 |
| 2018 | AiProAnnotator: Low-rank Approximation with network side information for high-performance, large-scale human Protein abnormality Annotator
Junning Gao, Shuwei Yao, Hiroshi Mamitsuka, Shanfeng Zhu |
BIBM | 4 |
| 2018 | GOLabeler: improving sequence-based large-scale protein function prediction by learning to rankabstractMotivation: Gene Ontology (GO) has been widely used to annotate functions of proteins and understand their biological roles. Currently only <1% of >70 million proteins in UniProtKB have experimental GO annotations, implying the strong necessity of automated function prediction (AFP) of proteins, where AFP is a hard multilabel classification problem due to one protein with a diverse number of GO terms. Most of these proteins have only sequences as input information, indicating the importance of sequence-based AFP (SAFP: sequences are the only input). Furthermore, homology-based SAFP tools are competitive in AFP competitions, while they do not necessarily work well for so-called difficult proteins, which have <60% sequence identity to proteins with annotations already. Thus, the vital and challenging problem now is how to develop a method for SAFP, particularly for difficult proteins. Methods: The key of this method is to extract not only homology information but also diverse, deep-rooted information/evidence from sequence inputs and integrate them into a predictor in a both effective and efficient manner. We propose GOLabeler, which integrates five component classifiers, trained from different features, including GO term frequency, sequence alignment, amino acid trigram, domains and motifs, and biophysical properties, etc., in the framework of learning to rank (LTR), a paradigm of machine learning, especially powerful for multilabel classification. Results: The empirical results obtained by examining GOLabeler extensively and thoroughly by using large-scale datasets revealed numerous favorable aspects of GOLabeler, including significant performance advantage over state-of-the-art AFP methods. Availability and implementation: http://datamining-iip.fudan.edu.cn/golabeler. Supplementary information: Supplementary data are available at Bioinformatics online. Ronghui You, Yi Xiong 0002, Fengzhu Sun, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 6 |
| 2017 | THCluster: Herb supplements categorization for precision traditional Chinese medicineabstractThere has been a continuing demand for traditional and complementary medicine worldwide. A fundamental and important topic in Traditional Chinese Medicine (TCM) is to optimize the prescription and to detect herb regularities from TCM data. In this paper, we propose a novel clustering model to solve this general problem of herb categorization, a pivotal task of prescription optimization and herb regularities. The model utilizes Random Walks method, Bayesian rules and Expectation Maximization(EM) models to complete a clustering analysis effectively on a heterogeneous information network. We performed extensive experiments on the real-world datasets and compared our method with other algorithms and experts. Experimental results have demonstrated the effectiveness of the proposed model for discovering useful categorization of herbs and its potential clinical manifestations. Chunyang Ruan, Ye Wang 0015, Yanchun Zhang, Jiangang Ma, Huijuan Chen, Uwe Aickelin, Shanfeng Zhu |
BIBM | 7 |
| 2017 | DeepText2Go: Improving large-scale protein function prediction with deep semantic text representationabstractUniProtKB has collected more than 88 million protein sequences by July 2017. Less than 0.2% of these proteins, however, have added experimental GO annotations. To reduce this huge gap, automatic protein function prediction (AFP) becomes increasingly important. Results on CAFA (the Critical Assessment of protein Function Annotation algorithms) benchmark demonstrates that sequence homology based methods are highly competitive in AFP. One imperative issues will be incorporating other information sources other than sequence for AFP. In contrast to using BOW (bag of words) representation in traditional text-based AFP, we proposed a new method called DeepText2GO to improve large-scale AFP by using deep semantic text representation instead. Furthermore, DeepText2GO integrates both text-based and sequence homology-based methods through a consensus approach. Extensive experiments on the benchmark dataset extracted from UniProt/SwissProt have demonstrated that DeepText2GO significantly outperformed both text-based and sequence homology-based methods, validating its superiority. Ronghui You, Shanfeng Zhu |
BIBM | 2 |
| 2016 | A Robust Convex Formulation for Ensemble Clustering
Junning Gao, Makoto Yamada, Samuel Kaski, Hiroshi Mamitsuka, Shanfeng Zhu |
IJCAI | 5 |
| 2016 | DeepMeSH: deep semantic representation for improving large-scale MeSH indexingabstractMOTIVATION: Medical Subject Headings (MeSH) indexing, which is to assign a set of MeSH main headings to citations, is crucial for many important tasks in biomedical text mining and information retrieval. Large-scale MeSH indexing has two challenging aspects: the citation side and MeSH side. For the citation side, all existing methods, including Medical Text Indexer (MTI) by National Library of Medicine and the state-of-the-art method, MeSHLabeler, deal with text by bag-of-words, which cannot capture semantic and context-dependent information well. METHODS: We propose DeepMeSH that incorporates deep semantic information for large-scale MeSH indexing. It addresses the two challenges in both citation and MeSH sides. The citation side challenge is solved by a new deep semantic representation, D2V-TFIDF, which concatenates both sparse and dense semantic representations. The MeSH side challenge is solved by using the 'learning to rank' framework of MeSHLabeler, which integrates various types of evidence generated from the new semantic representation. RESULTS: DeepMeSH achieved a Micro F-measure of 0.6323, 2% higher than 0.6218 of MeSHLabeler and 12% higher than 0.5637 of MTI, for BioASQ3 challenge data with 6000 citations. AVAILABILITY AND IMPLEMENTATION: The software is available upon request. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shengwen Peng, Ronghui You, Hongning Wang, ChengXiang Zhai, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 6 |
| 2016 | DrugE-Rank: improving drug-target interaction prediction of new candidate drugs or targets by ensemble learning to rankabstractMOTIVATION: Identifying drug-target interactions is an important task in drug discovery. To reduce heavy time and financial cost in experimental way, many computational approaches have been proposed. Although these approaches have used many different principles, their performance is far from satisfactory, especially in predicting drug-target interactions of new candidate drugs or targets. METHODS: Approaches based on machine learning for this problem can be divided into two types: feature-based and similarity-based methods. Learning to rank is the most powerful technique in the feature-based methods. Similarity-based methods are well accepted, due to their idea of connecting the chemical and genomic spaces, represented by drug and target similarities, respectively. We propose a new method, DrugE-Rank, to improve the prediction performance by nicely combining the advantages of the two different types of methods. That is, DrugE-Rank uses LTR, for which multiple well-known similarity-based methods can be used as components of ensemble learning. RESULTS: The performance of DrugE-Rank is thoroughly examined by three main experiments using data from DrugBank: (i) cross-validation on FDA (US Food and Drug Administration) approved drugs before March 2014; (ii) independent test on FDA approved drugs after March 2014; and (iii) independent test on FDA experimental drugs. Experimental results show that DrugE-Rank outperforms competing methods significantly, especially achieving more than 30% improvement in Area under Prediction Recall curve for FDA approved new drugs and FDA experimental drugs. AVAILABILITY: http://datamining-iip.fudan.edu.cn/service/DrugE-Rank CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qingjun Yuan, Junning Gao, Dongliang Wu, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 6 |
| 2015 | Instance-Wise Weighted Nonnegative Matrix Factorization for Aggregating Partitions with Locally Reliable Clusters
Shanfeng Zhu, Junning Gao, Hiroshi Mamitsuka |
IJCAI | 2 |
| 2015 | MeSHLabeler: improving the accuracy of large-scale MeSH indexing by integrating diverse evidenceabstractMOTIVATION: Medical Subject Headings (MeSHs) are used by National Library of Medicine (NLM) to index almost all citations in MEDLINE, which greatly facilitates the applications of biomedical information retrieval and text mining. To reduce the time and financial cost of manual annotation, NLM has developed a software package, Medical Text Indexer (MTI), for assisting MeSH annotation, which uses k-nearest neighbors (KNN), pattern matching and indexing rules. Other types of information, such as prediction by MeSH classifiers (trained separately), can also be used for automatic MeSH annotation. However, existing methods cannot effectively integrate multiple evidence for MeSH annotation. METHODS: We propose a novel framework, MeSHLabeler, to integrate multiple evidence for accurate MeSH annotation by using 'learning to rank'. Evidence includes numerous predictions from MeSH classifiers, KNN, pattern matching, MTI and the correlation between different MeSH terms, etc. Each MeSH classifier is trained independently, and thus prediction scores from different classifiers are incomparable. To address this issue, we have developed an effective score normalization procedure to improve the prediction accuracy. RESULTS: MeSHLabeler won the first place in Task 2A of 2014 BioASQ challenge, achieving the Micro F-measure of 0.6248 for 9,040 citations provided by the BioASQ challenge. Note that this accuracy is around 9.15% higher than 0.5724, obtained by MTI. AVAILABILITY AND IMPLEMENTATION: The software is available upon request. Ke Liu 0002, Shengwen Peng, Junqiu Wu, ChengXiang Zhai, Hiroshi Mamitsuka, Shanfeng Zhu |
Bioinform. | 6 |
| 2015 | Enhancing Time Series Clustering by Incorporating Multiple Distance Measures with Semi-Supervised Learning
Shanfeng Zhu, Xiaodi Huang 0001, Yanchun Zhang |
J. Comput. Sci. Technol. | 2 |
| 2015 | BMExpert: Mining MEDLINE for Finding Experts in Biomedical Domains Based on Language ModelabstractWith the rapid development of biomedical sciences, a great number of documents have been published to report new scientific findings and advance the process of knowledge discovery. By the end of 2013, the largest biomedical literature database, MEDLINE, has indexed over 23 million abstracts. It is thus not easy for scientific professionals to find experts on a certain topic in the biomedical domain. In contrast to the existing services that use some ad hoc approaches, we developed a novel solution to biomedical expert finding, BMExpert, based on the language model. For finding biomedical experts, who are the most relevant to a specific topic query, BMExpert mines MEDLINE documents by considering three important factors: relevance of documents to the query topic, importance of documents, and associations between documents and experts. The performance of BMExpert was evaluated on a benchmark dataset, which was built by collecting the program committee members of ISMB in the past three years (2012-2014) on 14 different topics. Experimental results show that BMExpert outperformed three existing biomedical expert finding services: JANE, GoPubMed, and eTBLAST, with respect to both MAP (mean average precision) and P@50 (Precision). BMExpert is freely accessed at http://datamining-iip.fudan.edu.cn/service/BMExpert/. Beichen Wang, Hiroshi Mamitsuka, Shanfeng Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2014 | Similarity-based machine learning methods for predicting drug-target interactions: a brief reviewabstractComputationally predicting drug-target interactions is useful to select possible drug (or target) candidates for further biochemical verification. We focus on machine learning-based approaches, particularly similarity-based methods that use drug and target similarities, which show relationships among drugs and those among targets, respectively. These two similarities represent two emerging concepts, the chemical space and the genomic space. Typically, the methods combine these two types of similarities to generate models for predicting new drug-target interactions. This process is also closely related to a lot of work in pharmacogenomics or chemical biology that attempt to understand the relationships between the chemical and genomic spaces. This background makes the similarity-based approaches attractive and promising. This article reviews the similarity-based machine learning methods for predicting drug-target interactions, which are state-of-the-art and have aroused great interest in bioinformatics. We describe each of these methods briefly, and empirically compare these methods under a uniform experimental setting to explore their advantages and limitations. Ichigaku Takigawa, Hiroshi Mamitsuka, Shanfeng Zhu |
Briefings Bioinform. | 4 |
| 2014 | An intelligent market making strategy in algorithmic trading
Xiaodong Li 0007, Xiaotie Deng, Shanfeng Zhu, Feng Wang 0048, Haoran Xie 0001 |
Frontiers Comput. Sci. | 3 |
| 2014 | Enhancing quantitative intra-day stock return prediction by integrating both market news and stock prices information
Xiaodong Li 0007, Xiaodi Huang 0001, Xiaotie Deng, Shanfeng Zhu |
Neurocomputing | 4 |
| 2013 | Document Summarization via Self-Present Sentence Relevance Model
Xiaodong Li 0007, Shanfeng Zhu, Haoran Xie 0001, Qing Li 0001 |
DASFAA (2) | 2 |
| 2013 | Collaborative matrix factorization with multiple similarities for predicting drug-target interactionsabstractWe address the problem of predicting new drug-target interactions from three inputs: known interactions, similarities over drugs and those over targets. This setting has been considered by many methods, which however have a common problem of allowing to have only one similarity matrix over drugs and that over targets. The key idea of our approach is to use more than one similarity matrices over drugs as well as those over targets, where weights over the multiple similarity matrices are estimated from data to automatically select similarities, which are effective for improving the performance of predicting drug-target interactions. We propose a factor model, named Multiple Similarities Collaborative Matrix Factorization(MSCMF), which projects drugs and targets into a common low-rank feature space, which is further consistent with weighted similarity matrices over drugs and those over targets. These two low-rank matrices and weights over similarity matrices are estimated by an alternating least squares algorithm. Our approach allows to predict drug-target interactions by the two low-rank matrices collaboratively and to detect similarities which are important for predicting drug-target interactions. This approach is general and applicable to any binary relations with similarities over elements, being found in many applications, such as recommender systems. In fact, MSCMF is an extension of weighted low-rank approximation for one-class collaborative filtering. We extensively evaluated the performance of MSCMF by using both synthetic and real datasets. Experimental results showed nice properties of MSCMF on selecting similarities useful in improving the predictive performance and the performance advantage of MSCMF over six state-of-the-art methods for predicting drug-target interactions. Hiroshi Mamitsuka, Shanfeng Zhu |
KDD | 4 |
| 2013 | Presenting XML Schema Mapping with Conjunctive-Disjunctive Views
Xuhui Li 0001, Shanfeng Zhu, Mengchi Liu, Ming Zhong 0002 |
WAIM | 2 |
| 2013 | Efficient Semisupervised MEDLINE Document Clustering With MeSH-Semantic and Global-Content ConstraintsabstractFor clustering biomedical documents, we can consider three different types of information: the local-content (LC) information from documents, the global-content (GC) information from the whole MEDLINE collections, and the medical subject heading (MeSH)-semantic (MS) information. Previous methods for clustering biomedical documents are not necessarily effective for integrating different types of information, by which only one or two types of information have been used. Recently, the performance of MEDLINE document clustering has been enhanced by linearly combining both the LC and MS information. However, the simple linear combination could be ineffective because of the limitation of the representation space for combining different types of information (similarities) with different reliability. To overcome the limitation, we propose a new semisupervised spectral clustering method, i.e., SSNCut, for clustering over the LC similarities, with two types of constraints: must-link (ML) constraints on document pairs with high MS (or GC) similarities and cannot-link (CL) constraints on those with low similarities. We empirically demonstrate the performance of SSNCut on MEDLINE document clustering, by using 100 data sets of MEDLINE records. Experimental results show that SSNCut outperformed a linear combination method and several well-known semisupervised clustering methods, being statistically significant. Furthermore, the performance of SSNCut with constraints from both MS and GC similarities outperformed that from only one type of similarities. Another interesting finding was that ML constraints more effectively worked than CL constraints, since CL constraints include around 10% incorrect ones, whereas this number was only 1% for ML constraints. Wei Feng 0005, Hiroshi Mamitsuka, Shanfeng Zhu |
IEEE Trans. Cybern. | 5 |
| 2012 | Toward more accurate pan-specific MHC-peptide binding prediction: a review of current methods and toolsabstractBinding of short antigenic peptides to major histocompatibility complex (MHC) molecules is a core step in adaptive immune response. Precise identification of MHC-restricted peptides is of great significance for understanding the mechanism of immune response and promoting the discovery of immunogenic epitopes. However, due to the extremely high MHC polymorphism and huge cost of biochemical experiments, there is no experimentally measured binding data for most MHC molecules. To address the problem of predicting peptides binding to these MHC molecules, recently computational approaches, called pan-specific methods, have received keen interest. Pan-specific methods make use of experimentally obtained binding data of multiple alleles, by which binding peptides (binders) of not only these alleles but also those alleles with no known binders can be predicted. To investigate the possibility of further improvement in performance and usability of pan-specific methods, this article extensively reviews existing pan-specific methods and their web servers. We first present a general framework of pan-specific methods. Then, the strategies and performance as well as utilities of web servers are compared. Finally, we discuss the future direction to improve pan-specific methods for MHC-peptide binding prediction. Lianming Zhang, Keiko Udaka, Hiroshi Mamitsuka, Shanfeng Zhu |
Briefings Bioinform. | 4 |
| 2011 | Improving Stock Market Prediction by Integrating Both Market News and Stock Prices
Xiaodong Li 0007, Feng Wang 0048, Xiaotie Deng, Shanfeng Zhu |
DEXA (2) | 6 |
| 2011 | Enhanced clustering of biomedical documents using ensemble non-negative matrix factorization
Xiaodi Huang 0001, Shanfeng Zhu |
Inf. Sci. | 5 |
| 2010 | Multiconstrained gene clustering based on generalized projectionsabstractBACKGROUND: Gene clustering for annotating gene functions is one of the fundamental issues in bioinformatics. The best clustering solution is often regularized by multiple constraints such as gene expressions, Gene Ontology (GO) annotations and gene network structures. How to integrate multiple pieces of constraints for an optimal clustering solution still remains an unsolved problem. RESULTS: We propose a novel multiconstrained gene clustering (MGC) method within the generalized projection onto convex sets (POCS) framework used widely in image reconstruction. Each constraint is formulated as a corresponding set. The generalized projector iteratively projects the clustering solution onto these sets in order to find a consistent solution included in the intersection set that satisfies all constraints. Compared with previous MGC methods, POCS can integrate multiple constraints from different nature without distorting the original constraints. To evaluate the clustering solution, we also propose a new performance measure referred to as Gene Log Likelihood (GLL) that considers genes having more than one function and hence in more than one cluster. Comparative experimental results show that our POCS-based gene clustering method outperforms current state-of-the-art MGC methods. CONCLUSIONS: The POCS-based MGC method can successfully combine multiple constraints from different nature for gene clustering. Also, the proposed GLL is an effective performance measure for the soft clustering solutions. Shanfeng Zhu, Alan Wee-Chung Liew, Hong Yan 0001 |
BMC Bioinform. | 2 |
| 2009 | Towards accurate human promoter recognition: a review of currently used sequence features and classification methodsabstractThis review describes important advances that have been made during the past decade for genome-wide human promoter recognition. Interest in promoter recognition algorithms on a genome-wide scale is worldwide and touches on a number of practical systems that are important in analysis of gene regulation and in genome annotation without experimental support of ESTs, cDNAs or mRNAs. The main focus of this review is on feature extraction and model selection for accurate human promoter recognition, with descriptions of what they are, what has been accomplished, and what remains to be done. Shanfeng Zhu, Hong Yan 0001 |
Briefings Bioinform. | 2 |
| 2009 | Enhancing MEDLINE document clustering by incorporating MeSH semantic similarityabstractMOTIVATION: Clustering MEDLINE documents is usually conducted by the vector space model, which computes the content similarity between two documents by basically using the inner-product of their word vectors. Recently, the semantic information of MeSH (Medical Subject Headings) thesaurus is being applied to clustering MEDLINE documents by mapping documents into MeSH concept vectors to be clustered. However, current approaches of using MeSH thesaurus have two serious limitations: first, important semantic information may be lost when generating MeSH concept vectors, and second, the content information of the original text has been discarded. METHODS: Our new strategy includes three key points. First, we develop a sound method for measuring the semantic similarity between two documents over the MeSH thesaurus. Second, we combine both the semantic and content similarities to generate the integrated similarity matrix between documents. Third, we apply a spectral approach to clustering documents over the integrated similarity matrix. RESULTS: Using various 100 datasets of MEDLINE records, we conduct extensive experiments with changing alternative measures and parameters. Experimental results show that integrating the semantic and content similarities outperforms the case of using only one of the two similarities, being statistically significant. We further find the best parameter setting that is consistent over all experimental conditions conducted. We finally show a typical example of resultant clusters, confirming the effectiveness of our strategy in improving MEDLINE document clustering. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shanfeng Zhu, Hiroshi Mamitsuka |
Bioinform. | 1 |
| 2009 | Field independent probabilistic model for clustering multi-field documents
Shanfeng Zhu, Ichigaku Takigawa, Hiroshi Mamitsuka |
Inf. Process. Manag. | 1 |
| 2007 | A Probabilistic Model for Clustering Text Documents with Multiple Fields
Shanfeng Zhu, Ichigaku Takigawa, Shuqin Zhang, Hiroshi Mamitsuka |
ECIR | 1 |
| 2006 | Improving MHC binding peptide prediction by incorporating binding data of auxiliary MHC moleculesabstractMOTIVATION: Various computational methods have been proposed to tackle the problem of predicting the peptide binding ability for a specific MHC molecule. These methods are based on known binding peptide sequences. However, current available peptide databases do not have very abundant amounts of examples and are highly redundant. Existing studies show that MHC molecules can be classified into supertypes in terms of peptide-binding specificities. Therefore, we first give a method for reducing the redundancy in a given dataset based on information entropy, then present a novel approach for prediction by learning a predictive model from a dataset of binders for not only the molecule of interest but also for other MHC molecules. RESULTS: We experimented on the HLA-A family with the binding nonamers of A1 supertype (HLA-A*0101, A*2601, A*2902, A*3002), A2 supertype (A*0201, A*0202, A*0203, A*0206, A*6802), A3 supertype (A*0301, A*1101, A*3101, A*3301, A*6801) and A24 supertype (A*2301 and A*2402), whose data were collected from six publicly available peptide databases and two private sources. The results show that our approach significantly improves the prediction accuracy of peptides that bind a specific HLA molecule when we combine binding data of HLA molecules in the same supertype. Our approach can thus be used to help find new binders for MHC molecules. Shanfeng Zhu, Keiko Udaka, John Sidney, Alessandro Sette, Kiyoko F. Aoki-Kinoshita, Hiroshi Mamitsuka |
Bioinform. | 1 |
| 2004 | Approximate and dynamic rank aggregation
Francis Y. L. Chin, Xiaotie Deng, Qizhi Fang, Shanfeng Zhu |
Theor. Comput. Sci. | 4 |
| 2003 | Approximate Rank Aggregation (Preliminary Version)
Xiaotie Deng, Qizhi Fang, Shanfeng Zhu |
COCOON | 3 |
| 2003 | Metasearch via Voting
Shanfeng Zhu, Qizhi Fang, Xiaotie Deng |
IDEAL | 1 |
| 2002 | Text Distinguishers Used in an Interactive Meta Search Engine
Kang Chen 0001, Xiaotie Deng, Haodi Feng, Shanfeng Zhu |
WAIM | 5 |
| 2001 | Membership for Core of LP Games and Other Games
Qizhi Fang, Shanfeng Zhu, Mao-cheng Cai, Xiaotie Deng |
COCOON | 2 |
| 2001 | On-Line Selection Of Distinguishing Elements For Focused Information RetrievalabstractThe internet can be viewed as a database with a huge amount of information. Powerful search engines present to users web pages that contain useful information, ranked according to the relevance to a particular query, and the quality of the web pages. However, users may have different interests even for the same query. In this work, we develop an on-line approach to learn to rank the web pages according to user preference. Kang Chen 0001, Hung Chim, Xiaotie Deng, Haodi Feng, Shanfeng Zhu |
ICME | 6 |
| 2001 | Using Online Relevance Feedback to Build Effective Personalized Metasearch EngineabstractMetasearch Engine is popular for facilitating users' queries over multiple search engines and increasing the coverage of the WWW. How to rank the merged results becomes crucial for the success of metasearch engines. Many current metasearch engines have poor precision, for one or more of selected source search engine returns irrelevant results. On the other hand, users with different interests may prefer distinct ranking order even for the same query. In this work, we try to use online relevance feedback to improve precision of the search results. At the same time, Users' preferences are recorded during the process of feedback for future ranking. Our elementary experiment shows that it is effective in improving precision of the metasearch engine. Shanfeng Zhu, Xiaotie Deng, Kang Chen 0001 |
WISE (1) | 1 |