EDBT 2026 Demo / reviewers in the wild / expert
Xiangwen Zheng
dblp:278/8400
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0001-7940-0514ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | BioEGRE: a linguistic topology enhanced method for biomedical relation extraction based on BioELECTRA and graph pointer neural networkabstractBACKGROUND: Automatic and accurate extraction of diverse biomedical relations from literature is a crucial component of bio-medical text mining. Currently, stacking various classification networks on pre-trained language models to perform fine-tuning is a common framework to end-to-end solve the biomedical relation extraction (BioRE) problem. However, the sequence-based pre-trained language models underutilize the graphical topology of language to some extent. In addition, sequence-oriented deep neural networks have limitations in processing graphical features. RESULTS: In this paper, we propose a novel method for sentence-level BioRE task, BioEGRE (BioELECTRA and Graph pointer neural net-work for Relation Extraction), aimed at leveraging the linguistic topological features. First, the biomedical literature is preprocessed to retain sentences involving pre-defined entity pairs. Secondly, SciSpaCy is employed to conduct dependency parsing; sentences are modeled as graphs based on the parsing results; BioELECTRA is utilized to generate token-level representations, which are modeled as attributes of nodes in the sentence graphs; a graph pointer neural network layer is employed to select the most relevant multi-hop neighbors to optimize representations; a fully-connected neural network layer is employed to generate the sentence-level representation. Finally, the Softmax function is employed to calculate the probabilities. Our proposed method is evaluated on three BioRE tasks: a multi-class (CHEMPROT) and two binary tasks (GAD and EU-ADR). The results show that our method achieves F1-scores of 79.97% (CHEMPROT), 83.31% (GAD), and 83.51% (EU-ADR), surpassing the performance of existing state-of-the-art models. CONCLUSION: The experimental results on 3 biomedical benchmark datasets demonstrate the effectiveness and generalization of BioEGRE, which indicates that linguistic topology and a graph pointer neural network layer explicitly improve performance for BioRE tasks. Xiangwen Zheng, Xuan-Ze Wang, Fan Tong |
BMC Bioinform. | 1 |
| 2022 | An Efficient and User-Friendly Software for PCR Primer Design for Detection of Highly Variable Bacteria
Dongzheng Hu, Wubin Qu, Fan Tong, Xiangwen Zheng, Jiangyu Li |
ISBRA | 4 |
| 2022 | BioByGANS: biomedical named entity recognition by fusing contextual and syntactic features through graph attention network in node classification frameworkabstractBACKGROUND: Automatic and accurate recognition of various biomedical named entities from literature is an important task of biomedical text mining, which is the foundation of extracting biomedical knowledge from unstructured texts into structured formats. Using the sequence labeling framework and deep neural networks to implement biomedical named entity recognition (BioNER) is a common method at present. However, the above method often underutilizes syntactic features such as dependencies and topology of sentences. Therefore, it is an urgent problem to be solved to integrate semantic and syntactic features into the BioNER model. RESULTS: In this paper, we propose a novel biomedical named entity recognition model, named BioByGANS (BioBERT/SpaCy-Graph Attention Network-Softmax), which uses a graph to model the dependencies and topology of a sentence and formulate the BioNER task as a node classification problem. This formulation can introduce more topological features of language and no longer be only concerned about the distance between words in the sequence. First, we use periods to segment sentences and spaces and symbols to segment words. Second, contextual features are encoded by BioBERT, and syntactic features such as part of speeches, dependencies and topology are preprocessed by SpaCy respectively. A graph attention network is then used to generate a fusing representation considering both the contextual features and syntactic features. Last, a softmax function is used to calculate the probabilities and get the results. We conduct experiments on 8 benchmark datasets, and our proposed model outperforms existing BioNER state-of-the-art methods on the BC2GM, JNLPBA, BC4CHEMD, BC5CDR-chem, BC5CDR-disease, NCBI-disease, Species-800, and LINNAEUS datasets, and achieves F1-scores of 85.15%, 78.16%, 92.97%, 94.74%, 87.74%, 91.57%, 75.01%, 90.99%, respectively. CONCLUSION: The experimental results on 8 biomedical benchmark datasets demonstrate the effectiveness of our model, and indicate that formulating the BioNER task into a node classification problem and combining syntactic features into the graph attention networks can significantly improve model performance. Xiangwen Zheng, Haijian Du, Fan Tong |
BMC Bioinform. | 1 |
| 2021 | Combining Query Reformulation and Re-ranking to Improve Query Expansion in Chinese EMR RetrievalabstractThe methods of query expansion (QE) have achieved significant performance in information retrieval of electronic medical records (EMR). It is a pity that the direct addition of expansion terms may cause query drift, which decreases the precision of EMR retrieval. To solve the issue, we combine the methods of query reformulation and re-ranking to improve the performance of QE in Chinese EMR retrieval. First, the synonyms and hyponyms are extracted from a Chinese medical knowledge graph as expansion terms, and the weights of expansion terms are calculated with the combination of semantic similarities, category weights, and co-occurrence frequencies. Then the query is reformulated by these expansion terms, and the EMR documents are retrieved by the expanded query. Second, four categories of medical terms in top retrieval results are selected, and the performances of all combinations in re-ranking are tested to select the most proper terms for re-ranking. Then the original retrieval results are re-ranked by these terms. Experiments show that the algorithm can promote the effectiveness of EMR retrieval compared with baselines, which shows that the combination of query reformulation and re-ranking can differentiate expansion terms of different categories and maximize their effects. Songchun Yang, Xiangwen Zheng |
BIBM | 2 |
| 2021 | COVID19-OBKG: An Ontology-Based Knowledge Graph and Web Service for COVID-19abstractGiven the huge amount of data from diverse sources and involving various conceptual fields in heterogeneous formats, researchers have encountered challenges in their effort to process, search for, and access knowledge about coronavirus disease 2019 (COVID-19). In this paper, we built COVID19-OBKG, an ontology-based knowledge graph and web service for COVID-19, to enable the access and retrieval of knowledge. First, we built the schema of COVID19-OBKG based on biomedical ontologies to guide the construction of the instance layer of COVID19-OBKG from top to bottom. Secondly, we collected data sources related to COVID-19, including structured databases and web pages. We acquired entities and relationships from data sources through named entity recognition and relation extraction algorithms and merged them with knowledge in biomedical ontologies. Thirdly, we modeled our data in the form of an attribute graph and stored it in Dgraph. Finally, we built a web service to support the retrieval and visualization of COVID19-OBKG, which verified the effectiveness of our approach to constructing a knowledge graph, and the usability of COVID19-OBKG. Xiangwen Zheng, Fan Tong |
BIBM | 1 |
| 2021 | Improving Chinese electronic medical record retrieval by field weight assignment, negation detection, and re-ranking
Songchun Yang, Xiangwen Zheng, Xiangfei Yin, Jianfei Pang, Huajian Mao, Wenqin Zhang |
J. Biomed. Informatics | 2 |
| 2020 | A Novel Algorithm of Expansion Term Selection and Weight Assignment for Query Expansion of Chinese EMR RetrievalabstractThe techniques of information retrieval eliminate the difficulty of finding specific information in electronic medical records (EMR), and the methods of query expansion (QE) improve the recall of EMR retrieval. However, most existing QE methods of EMR retrieval can't get high-quality expansion terms and corresponding weights for Chinese EMR retrieval because of low-quality sources and improper calculation methods. In this paper, we propose a novel algorithm of expansion term selection and weight assignment for QE of Chinese EMR retrieval based on the clinical needs and unique characteristics of Chinese medical terms. The algorithm first selects expansion terms from a high quality Chinese medical knowledge graph and standard medical term sets, which can ensure the quality of expansion terms. Then it assigns weights to the selected expansion terms based on semantic similarities and manually designed expansion categories that reflect the opinions of medical experts. Experiment results show that our algorithm gets higher-quality expansion terms and more rational weights compared with four benchmark algorithms, using Precision at 10, Recall, Mean Average Precision and Binary Preference-based Measure as evaluation metrics, which implies that our algorithm can significantly improve the effectiveness of QE. Songchun Yang, Xiangwen Zheng, Fan Tong, Huajian Mao |
BIBM | 2 |