Yingwen Zhao

dblp:210/9184 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 6 since 2021
YearPublicationVenuePosition
2025 A Novel Method of Stacking InSAR Interferograms for Measuring Localized Earthquake Deformation
abstract
Surface displacements caused by small-to-moderate earthquakes are generally localized within small regions with magnitudes of only several centimeters. Interferometric synthetic aperture radar (InSAR) measurements of such deformation are prone to be contaminated by localized turbulent atmospheric delays, which are difficult to be mitigated by current methods. Here we develop a new method to generate an optimal combination of coseismic interferograms for stacking to measure the deformation. The interferograms are weighted during the stacking process based on their signal-to-noise ratios. To verify the performance of mitigating the tropospheric delays, we conduct both synthetic data experiments and practical applications with three different types of small-magnitude earthquakes to compare our method with current methods of differential InSAR, time series analysis and multi-interferogram stacking. The experiment results demonstrate that our method can mitigate not only turbulent atmospheric delays but also other non-seismic deformation signals better than previous methods. The optimal combinations for measuring localized earthquake deformation consist of interferograms with both small and large standard deviation values. Further inversion tests reveal that accurate measurements of surface deformation obtained by our method are favorable to determine source faults and slip distributions of small-magnitude earthquakes without surface rupture. Although the experiments are conducted with only consideration of earthquake deformation, our method is also expected to be applicable for other tectonic and non-tectonic activities, like volcanoes and landslides, that occur within localized regions and over short time periods around several days.
Yingwen Zhao, Guoyan Jiang, Kefeng He, Caijun Xu, Renqi Lu
IEEE Trans. Geosci. Remote. Sens.1
2024 Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasks
abstract
OBJECTIVE: Most existing fine-tuned biomedical large language models (LLMs) focus on enhancing performance in monolingual biomedical question answering and conversation tasks. To investigate the effectiveness of the fine-tuned LLMs on diverse biomedical natural language processing (NLP) tasks in different languages, we present Taiyi, a bilingual fine-tuned LLM for diverse biomedical NLP tasks. MATERIALS AND METHODS: We first curated a comprehensive collection of 140 existing biomedical text mining datasets (102 English and 38 Chinese datasets) across over 10 task types. Subsequently, these corpora were converted to the instruction data used to fine-tune the general LLM. During the supervised fine-tuning phase, a 2-stage strategy is proposed to optimize the model performance across various tasks. RESULTS: Experimental results on 13 test sets, which include named entity recognition, relation extraction, text classification, and question answering tasks, demonstrate that Taiyi achieves superior performance compared to general LLMs. The case study involving additional biomedical NLP tasks further shows Taiyi's considerable potential for bilingual biomedical multitasking. CONCLUSION: Leveraging rich high-quality biomedical corpora and developing effective fine-tuning strategies can significantly improve the performance of LLMs within the biomedical domain. Taiyi shows the bilingual multitasking capability through supervised fine-tuning. However, those tasks such as information extraction that are not generation tasks in nature remain challenging for LLM-based generative approaches, and they still underperform the conventional discriminative approaches using smaller language models.
Ling Luo 0001, Jinzhong Ning, Yingwen Zhao, Zeyuan Ding, Weiru Fu, Qinyu Han, Guangtao Xu, Yunzhi Qiu, Dinghao Pan, Jiru Li, Wenduo Feng, Senbo Tu, Jian Wang 0021, Yuanyuan Sun 0002, Hongfei Lin
J. Am. Medical Informatics Assoc.3
2024 Predicting Protein Functions Based on Heterogeneous Graph Attention Technique
abstract
In bioinformatics, protein function prediction stands as a fundamental area of research and plays a crucial role in addressing various biological challenges, such as the identification of potential targets for drug discovery and the elucidation of disease mechanisms. However, known functional annotation databases usually provide positive experimental annotations that proteins carry out a given function, and rarely record negative experimental annotations that proteins do not carry out a given function. Therefore, existing computational methods based on deep learning models focus on these positive annotations for prediction and ignore these scarce but informative negative annotations, leading to an underestimation of precision. To address this issue, we introduce a deep learning method that utilizes a heterogeneous graph attention technique. The method first constructs a heterogeneous graph that covers the protein-protein interaction network, ontology structure, and positive and negative annotation information. Then, it learns embedding representations of proteins and ontology terms by using the heterogeneous graph attention technique. Finally, it leverages these learned representations to reconstruct the positive protein-term associations and score unobserved functional annotations. It can enhance the predictive performance by incorporating these known limited negative annotations into the constructed heterogeneous graph. Experimental results on three species (i.e., Human, Mouse, and Arabidopsis) demonstrate that our method can achieve better performance in predicting new protein annotations than state-of-the-art methods.
Yingwen Zhao, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021
IEEE J. Biomed. Health Informatics1
2023 Improving Protein Function Prediction by Adaptively Fusing Information From Protein Sequences and Biomedical Literature
abstract
Proteins are the main undertakers of life activities, and accurately predicting their biological functions can help human better understand life mechanism and promote the development of themselves. With the rapid development of high-throughput technologies, an abundance of proteins are discovered. However, the gap between proteins and function annotations is still huge. To accelerate the process of protein function prediction, some computational methods taking advantage of multiple data have been proposed. Among these methods, the deep-learning-based methods are currently the most popular for their capability of learning information automatically from raw data. However, due to the diversity and scale difference between data, it is challenging for existing deep learning methods to capture related information from different data effectively. In this paper, we introduce a deep learning method that can adaptively learn information from protein sequences and biomedical literature, namely DeepAF. DeepAF first extracts the two kinds of information by using different extractors, which are built based on pre-trained language models and can capture rudimentary biological knowledge. Then, to integrate those information, it performs an adaptive fusion layer based on a Cross-attention mechanism that considers the knowledge of mutual interactions between two information. Finally, based on the mixed information, DeepAF utilizes logistic regression to obtain prediction scores. The experimental results on the datasets of two species (i.e., Human and Yeast) show that DeepAF outperforms other state-of-the-art approaches.
Yingwen Zhao, Yongkai Hong, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021
IEEE J. Biomed. Health Informatics1
2022 Adaptive Multi-view Graph Convolutional Network for Gene Ontology Annotations of Proteins
abstract
Gene Ontology (GO) containing a set of standard concepts (or terms) is launched to unify the functional descriptions of proteins. Developing computational models based on GO to automatically annotate protein functions has been a longstanding active research area. In this paper, we propose a novel method to adaptively fuse functional and topological information between GO Terms. Our method is composed of a pre-trained language model for encoding protein sequences and an adaptive multi-view graph convolutional network (Multi-view GCN) for representing GO terms. Particularly, the Multi-view GCN considers multiple views from functional information, topological structures, and their combinations, and extracts multiple corresponding representations of GO terms. Then, an attention mechanism is applied to adaptively learn the importance weights of these representations. Finally, the predicted scores are calculated by using a dot product between protein sequence features and GO term representations. Experimental results on the datasets of two species (i.e., Human and Yeast) show that our method outperforms other state-of-the-art methods. The code of our proposed method is available at: https://github.com/Candyperfect/Master.
Yingwen Zhao, Yongkai Hong, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021
BIBM1
2021 Cross-Species Protein Function Prediction with Asynchronous-Random Walk
abstract
Protein function prediction is a fundamental task in the post-genomic era. Available functional annotations of proteins are incomplete and the annotations of two homologous species are complementary to each other. However, how to effectively leverage mutually complementary annotations of different species to further boost the prediction performance is still not well studied. In this paper, we propose a cross-species protein function prediction approach by performing Asynchronous Random Walk on a heterogeneous network (AsyRW). AsyRW first constructs a heterogeneous network to integrate multiple functional association networks derived from different biological data, established homology-relationships between proteins from different species, known annotations of proteins and Gene Ontology (GO). To account for the intrinsic structures of intra- and inter-species of proteins and that of GO, AsyRW quantifies the individual walk lengths of each network node using the gravity-like theory, and then performs asynchronous-random walk with the individual length to predict associations between proteins and GO terms. Experiments on annotations archived in different years show that individual walk length and asynchronous-random walk can effectively leverage the complementary annotations of different species, AsyRW has a significantly improved performance to other related and competitive methods. The codes of AsyRW are available at: http://mlda.swu.edu.cn/codes.php?name=AsyRW.
Yingwen Zhao, Jun Wang 0035, Maozu Guo 0001, Xiangliang Zhang 0001, Guoxian Yu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 NewGOA: Predicting New GO Annotations of Proteins by Bi-Random Walks on a Hybrid Graph
abstract
A remaining key challenge of modern biology is annotating the functional roles of proteins. Various computational models have been proposed for this challenge. Most of them assume the annotations of annotated proteins are complete. But in fact, many of them are incomplete. We proposed a method called NewGOA to predict new Gene Ontology (GO) annotations for incompletely annotated proteins and for completely un-annotated ones. NewGOA employs a hybrid graph, composed of two types of nodes (proteins and GO terms), to encode interactions between proteins, hierarchical relationships between terms and available annotations of proteins. To account for structural difference between GO terms subgraph and proteins subgraph, NewGOA applies a bi-random walks algorithm, which executes asynchronous random walks on the hybrid graph, to predict new GO annotations of proteins. Experimental study on archived GO annotations of two model species (H. Sapiens and S. cerevisiae) shows that NewGOA can more accurately and efficiently predict new annotations of proteins than other related methods. Experimental results also indicate the bi-random walks can explore and further exploit the structural difference between GO terms subgraph and proteins subgraph. The supplementary files and codes of NewGOA are available at: http://mlda.swu.edu.cn/codes.php?name=NewGOA.
Guoxian Yu, Guangyuan Fu, Jun Wang 0035, Yingwen Zhao
IEEE ACM Trans. Comput. Biol. Bioinform.4