EDBT 2026 Demo / reviewers in the wild / expert
Fei He 0003
dblp:13/6794-3
· DBLP profile ↗
9ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-3284-9506ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | pathCLIP: Detection of Genes and Gene Relations From Biological Pathway Figures Through Image-Text Contrastive LearningabstractIn biomedical literature, biological pathways are commonly described through a combination of images and text. These pathways contain valuable information, including genes and their relationships, which provide insight into biological mechanisms and precision medicine. Curating pathway information across the literature enables the integration of this information to build a comprehensive knowledge base. While some studies have extracted pathway information from images and text independently, they often overlook the correspondence between the two modalities. In this paper, we present a pathway figure curation system named pathCLIP for identifying genes and gene relations from pathway figures. Our key innovation is the use of an image-text contrastive learning model to learn coordinated embeddings of image snippets and text descriptions of genes and gene relations, thereby improving curation. Our validation results, using pathway figures from PubMed, showed that our multimodal model outperforms models using only a single modality. Additionally, our system effectively curates genes and gene relations from multiple literature sources. Two case studies on extracting pathway information from literature of non-small cell lung cancer and Alzheimer's disease further demonstrate the usefulness of our curated pathway information in enhancing related pathways in the KEGG database. Fei He 0003, Richard D. Hammer, Dong Xu 0002, Mihail Popescu |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Comprehensive Assessment of OCR Tools for Gene Name Recognition in Biological Pathway FiguresabstractOptical Character Recognition (OCR) is becoming more and more effective in text detection in images. However, OCR’s performance in special applications may vary. In particular, OCR in visual representations of complex processes known as pathway figures in the biomedical literature is challenging. The information depicted in a pathway graphic usually represents the article’s most important conclusions. Still, the huge number of pathway figures cannot be automatically processed for large-scale search, data mining, and downstream analysis. Assisted by recent developments in OCR, we have developed a method to extract gene names from pathway images. For usage in the method, we thoroughly evaluated and compared major available OCR tools using 563 genes from 45 pathway images, 1000 images of alphanumeric characters and gene names from HUGO, and KEGG data with 20 random pathway genes curated routes. Our study showed that Google Cloud Vision and MMOCR are best suitable for gene name recognition in pathway figures. Stuart Aldrich, Micheal Olaolu Arowolo, Fei He 0003, Mihail Popescu, Dong Xu 0002 |
BIBM | 3 |
| 2022 | Extraction of Gene Regulatory Relation Using BioBERTabstractRelation Extraction (RE) is a critical task typically carried out after Named Entity recognition for identifying gene-gene association from scientific publication. Current state-of the-art tools have limited capacity as most of them only extract entity relations from abstract texts. The retrieved gene-gene relations typically do not cover gene regulatory relations. There still exists a lot of room for improvement. In this work, we propose GeREx, a transformer based RE tool for identifying gene regulatory relations from full texts by incorporating labeling from pathway figures. GeREx achieved an F1-Score of 83.45% when evaluated on an independent test dataset. Clement Essien, Fei He 0003, Mark Hannink, Mihail Popescu, Dong Xu 0002 |
BIBM | 2 |
| 2022 | Evaluating template-based and template-free protein-peptide complex structure prediction using AlphaFold2abstractProtein-peptide interactions play a crucial role in a wide range of biological processes, such as cellular regulation, immune responses, and signal transduction. Understanding the details of these interactions and predicting their complex structures is a vital key to peptide-based drug design. Predicting protein-peptide docking structure has experienced impressive scientific momentum over the past few years, which plays a crucial role in designing and developing peptide drugs. A wide range of computational algorithms has been developed for protein-peptide-docking predictions, but almost all of them require experimental protein structures, which are expensive and time-consuming. Benefiting from Alphafold2 and RoseTTAFold tools, most protein structures can be accurately deciphered based on protein sequences only. Formulating protein-peptide binding/docking as a protein complex folding problem, we can use Alphafold2 to generate a series of bounded protein-peptide conformations. We have designed a pipeline for predicting protein-peptide complexes and scoring the predicted models using AI-based tools. We benchmarked the pipeline by a set of non-redundant protein-peptide complex structures derived from the most recent-released complexes in PDB during the 2021 and 2022 years. Both template-based and template-free methods of Alphafold2 were specifically analyzed in protein-peptide complex structure predictions, and the strengths and weaknesses of each method were identified. We showed that the near-native complex structures were often not obtained by both of them and suggested that integrating the results of both template-based and template-free methods could be a good strategy. We compared the results with several protein-peptide tools, such as RoseTTAFold. We also evaluated several scoring schemes, including our in-house method based on a graph neural network, in ranking protein-peptide binding conformations. Negin Manshour, Wenyuan Qin, Fei He 0003, Duolin Wang, Dong Xu 0002 |
BIBM | 4 |
| 2021 | Discover the Binding Domain of Transmembrane Proteins Based on Structural UniversalityabstractTransmembrane proteins (TMPs) serve as drug targets for more than half of the drugs currently available in the market. However, it had not been clearly explained how they realize their drug effects through multiple complex molecules bindings actions. Research into TMPs bindings and corresponding structural basis will provide key information for drug research and new drug development. In this study, we defined the binding domain of TMPs according to the binding region investigation of multiple conjugate types. A 3D deep learning model was architected to discover the structural universality inside those domains. The experimental results proved such binding domains existing on the surface of TMPs, and they are structural specific distinguishing to the surface regions without any binding activities. This work provides a new theoretical basis for TMPs binding research and can greatly boost the development of the drug industry. Yihang Bao, Fei He 0003, Weixi Wang, Han Wang 0028, Minglong Dong |
BIBM | 2 |
| 2021 | An Ensemble Deep Learning based Predictor for Simultaneously Identifying Protein Ubiquitylation and SUMOylation SitesabstractBACKGROUND: Several computational tools for predicting protein Ubiquitylation and SUMOylation sites have been proposed to study their regulatory roles in gene location, gene expression, and genome replication. However, existing methods generally rely on feature engineering, and ignore the natural similarity between the two types of protein translational modification. This study is the first all-in-one deep network to predict protein Ubiquitylation and SUMOylation sites from protein sequences as well as their crosstalk sites simultaneously. Our deep learning architecture integrates several meta classifiers that apply deep neural networks to protein sequence information and physico-chemical properties, which were trained on multi-label classification mode for simultaneously identifying protein Ubiquitylation and SUMOylation as well as their crosstalk sites. RESULTS: The promising AUCs of our method on Ubiquitylation, SUMOylation and crosstalk sites achieved 0.838, 0.888, and 0.862 respectively on tenfold cross-validation. The corresponding APs reached 0.683, 0.804 and 0.552, which also validated our effectiveness. CONCLUSIONS: The proposed architecture managed to classify ubiquitylated and SUMOylated lysine residues along with their crosstalk sites, and outperformed other well-known Ubiquitylation and SUMOylation site prediction tools. Fei He 0003, Xiaowei Zhao 0004 |
BMC Bioinform. | 1 |
| 2019 | Protein Ubiquitylation and Sumoylation Site Prediction Based on Ensemble and Transfer LearningabstractUbiquitylation, a typical post-translational modification (PTM), plays an important role in signal transduction, apoptosis and cell proliferation. A ubiquitylation like PTM, sumoylation also may affect gene mapping, expression and genomic replication. Over the past two decades, machine learning has been widely employed in protein ubiquitylation and sumoylation site prediction tools. These existing tools require feature engineering, but failed to provide general interpretable features and probably underutilized the growing amount of data. This prompted us to propose a deep learning-based model that integrates multiple convolution and fully-connected layers of seven supervised learning sub-models to extract deep representations from protein sequences and physico-chemical properties (PCPs). Especially, we divided PCPs into 6 clusters and customized deep networks accordingly for handling the high correlations among one cluster. A stacking ensemble strategy was applied to combine these deep representations to make prediction. Furthermore, with the advantage of transfer learning, our deep learning model can work well on protein sumoylation site prediction as well after fine-tuning. On the high-quality annotated database Swiss-Prot, our model outperformed several well-known ubiquitylation and sumoylation site prediction tools. Our code is freely available at https://github.com/ruiwcoding/DeepUbiSumoPre. Fei He 0003, Yanxin Gao, Duolin Wang, Dong Xu 0002, Xiaowei Zhao 0004 |
BIBM | 1 |
| 2017 | Effective small interfering RNA design based on convolutional neural networkabstractIn functional genomics, small interfering RNA (siRNA) can be used to knockdown gene expression. Usually, a target gene has numerous potential siRNAs, but their efficiencies of gene silencing often varies. Thus, for a successful RNA interference (RNAi), selecting the most effective siRNA is a critical step. Despite various computational algorithms have been developed, the efficacy prediction accuracy is not so satisfactory. In this paper, to explore the effect of different motifs on gene silencing and further improve the prediction accuracy, we developed a new powerful predictor by using a deep learning algorithm - Convolutional Neural Network (CNN). The comparison results showed that the Pearson Correlation Coefficient (PCC) of our model is 0.717, which is 13.81%, 16.78% and 5.91% higher than Biopredsi, i-Score, ThermoComposition21 and DSIR. In addition, the area under the ROC curve (AUC) of our model is 0.894, which is 10.10%, 12.59% and 7.07% higher than those four algorithms. The results show that our model is stable and efficient to predict siRNA silencing efficacy. Fei He 0003, Xian Tan, Helong Yu |
BIBM | 2 |
| 2017 | A multimodal deep architecture for large-scale protein ubiquitylation site predictionabstractIn eukaryotes, protein ubiquitylation is an important type of post-translation modification, in which the ubiquitin conjugates to a substrate protein. To have a better insight of the mechanisms underlying ubiquitylation, a key step is to identify protein ubiquitylation sites. Many existing computational methods are based on feature engineering, which may lead to biased and incomplete features. Deep learning provides multiple-layer networks and non-linear mapping operations to detect potential complex patterns in a data-driven way, especially for large-scale data. It provides a promising new method to predict ubiquitylation sites. In this paper, we proposed a multimodal deep architecture for protein ubiquitylation sites prediction. First, we designed different multiple layers to extract hidden informative patterns from three modalities, namely protein fragments, physico-chemical properties, and sequence profiles. Then, the deep representations corresponding to three modalities were merged to implement the classification. On the available largest scale protein ubiquitylation site database PLMD, the performance of our proposed method was measured with 66.7% sensitivity, 66.4% specificity, 66.43% accuracy, and 0.221 MCC value. A range of comparative experiments also showed that our proposed architecture outperformed several popular protein ubiquitylation site prediction tools. Our source code is freely available at https://github.com/jiagenlee/deepUbiquitylation. Fei He 0003, Lingling Bao, Jiagen Li, Dong Xu 0002, Xiaowei Zhao 0004 |
BIBM | 1 |