VLDB 2026 Research / reviewers in the wild / expert
Di Zhao 0003
dblp:58/4162-3
· DBLP profile ↗
26ranked-venue papers
6as first author
24since 2021 · last 2027
0000-0002-0876-5126ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | SEL-PMF:Semantic enhancement via LLMs and progressive multimodal fusion for fake news detection
Linlin Gao, Jiana Meng, Di Zhao 0003 |
Expert Syst. Appl. | 3 |
| 2026 | A Multi-Agent LLM Framework for Multi-Domain Low-Resource In-Context NER via Knowledge Retrieval, Disambiguation and Reflective AnalysisabstractIn-context learning (ICL) with large language models (LLMs) has emerged as a promising paradigm for named entity recognition (NER) in low-resource scenarios. However, existing ICL-based NER methods suffer from three key limitations: (1) reliance on dynamic retrieval of annotated examples, which is problematic when annotated data is scarce; (2) limited generalization to unseen domains due to the LLM's insufficient internal domain knowledge; and (3) failure to incorporate external knowledge or resolve entity ambiguities. To address these challenges, we propose KDR-Agent, a novel multi-agent framework for multi-domain low-resource in-context NER that integrates Knowledge retrieval, Disambiguation, and Reflective analysis. KDR-Agent leverages natural-language type definitions and a static set of entity-level contrastive demonstrations to reduce dependency on large annotated corpora. A central planner coordinates specialized agents to (i) retrieve factual knowledge from Wikipedia for domain-specific mentions, (ii) resolve ambiguous entities via contextualized reasoning, and (iii) reflect on and correct model predictions through structured self-assessment. Experiments across ten datasets from five domains demonstrate that KDR-Agent significantly outperforms existing zero-shot and few-shot ICL baselines across multiple LLM backbones. Wenxuan Mu, Jinzhong Ning, Di Zhao 0003, Yi-Jia Zhang 0001 |
AAAI | 3 |
| 2026 | Data Augmentation for Few-Shot Biomedical NER Using ChatGPT
Wenxuan Mu, Di Zhao 0003, Jiana Meng, Shichang Sun, Jian Wang 0021, Hongfei Lin |
Artif. Intell. Medicine | 2 |
| 2025 | Bio-R3: Entity Aware and Domain Adaptive Enhanced Framework for Biomedical Document Level Relation ExtractionabstractBiomedical document-level relation extraction plays a crucial role in mining structured knowledge from biomedical texts. However, existing large language model-based methods struggle with information redundancy, noise interference, and insufficient domain-specific knowledge when processing complex documents. To address these challenges, we propose Bio-$\mathbf{R}^{\mathbf{3}}$, a novel framework that enhances document-level relation extraction through entity aware rewriting and domain-adaptive reasoning. Our Bio-R${ }^{3}$follows a “Rewrite-Retrieval-Reason” pipeline: (1) Entity aware rewriting refocuses the document content on target entity pairs, reducing irrelevant context; (2) Retrieval-augmented generation supplements the document with external biomedical evidence to support relation inference; and (3) Chain-of-Thought prompting and few-shot learning improve reasoning accuracy, while Low-Rank Adaptation fine-tuning enhances domain adaptation. Experiments on the CDR and GDA benchmarks demonstrate that Bio-$\mathbf{R}^{\mathbf{3}}$significantly outperforms existing methods, demonstrating its effectiveness in the extraction of biomedical document relations. Di Zhao 0003, Jiana Meng, Hongfei Lin |
BIBM | 1 |
| 2025 | SAVe-Vis: An End-to-End Multimodal Framework for Schema-Aware and Self-Validating Data Visualization Generation
Jiana Meng, Di Zhao 0003 |
IEEE Big Data | 3 |
| 2025 | MEAN:Multi-Modal Explainability Analysis Network for Fake News DetectionabstractWith the rapid development of the internet, the news domain is flooded with an increasing amount of multi-modal fake news, which poses a great threat to today’s society. Although some methods can automatically detect fake news recently, the lack of explainability remains a challenge. This paper first pre-defines three discriminative patterns: image tampering, text forgery, and inconsistency between image and text. Revealing these three patterns not only lead to better prediction results but also provide clear and concise explanations. Therefore, we propose Multi-modal Explainability Analysis Network (MEAN) for fake news detection. This model first extracts discriminative patterns using both uni-modal and multi-modal methods, then employs a multi-branch network to obtain fine-grained representations for these patterns, and finally learns the discriminative patterns through logical rules to obtain the predicted labels of news and the explainability analysis of the model. Extensive experiments on two real-world datasets demonstrate the superiority of the proposed MEAN model in detecting fake news while providing explanations for the discriminative patterns. Linlin Gao, Jiana Meng, Di Zhao 0003 |
IJCNN | 3 |
| 2025 | LCDL: Classification of ICD codes based on disease label co-occurrence dependency and LongFormer with medical knowledge
Hongfei Lin, Yi-Jia Zhang 0001, Di Zhao 0003, Ling Luo 0001 |
Artif. Intell. Medicine | 5 |
| 2025 | SyRACT: zero-shot biomedical document-level relation extraction with synergistic RAG and CoTabstractMOTIVATION: With the advancement of large language models (LLMs), the field of biomedical document-level relation extraction (BioDocRE) has encountered new opportunities. However, LLMs often face challenges such as hallucinated generation, insufficient reasoning capabilities, and a lack of interpretability when performing relation extraction tasks. RESULTS: To address these issues, we propose the SyRACT (Synergistic Retrieval Augmented Generation and Chain of Thought) framework for high precision relation extraction in biomedical documents. This framework is built around three core strategies: (i) reframing the relation extraction task as a question answering problem to better align with the processing logic of LLMs; (ii) leveraging an external database constructed from PubMed to provide LLMs with rich and reliable contextual information, thus mitigating hallucination generation; and (iii) construct a specific Chain of Thought for BioDocRE tasks, thereby enhancing the model's reasoning ability and the interpretability of its output. We validated this approach on three biomedical relation extraction datasets: CDR, GDA, and ADE. Experimental results show that the SyRACT model improves F1 scores by 11.04%, 9.10%, and 41.00% on three datasets, respectively, compared to the DocRE method, which uses standard prompts for LLMs. AVAILABILITY AND IMPLEMENTATION: Our source code and data are available at https://github.com/donggggxin/SyRACT. Di Zhao 0003, Jiana Meng, Bocheng Guo, Hongfei Lin |
Bioinform. | 2 |
| 2025 | Exploring biomedical relation extraction by combining natural language inference and dual dependency trees
Xueying Li 0002, Di Zhao 0003, Jiana Meng, Hongfei Lin |
Expert Syst. Appl. | 2 |
| 2025 | Domain feature transfer-based multi-domain fake news detection
Xuan Meng, Di Zhao 0003, Jiana Meng, Xiaopei Wang |
Knowl. Inf. Syst. | 2 |
| 2025 | Fake news detection based on multi-modal domain adaptation
Xiaopei Wang, Jiana Meng, Di Zhao 0003, Xuan Meng, Hewen Sun |
Neural Comput. Appl. | 3 |
| 2024 | Few-shot Biomedical NER via Multi-task Learning and More Fine-grained Grid-tagging StrategyabstractBiomedical Named Entity Recognition (NER) serves as a crucial task in biomedical information extraction, aiming to identify and classify entities from unstructured biomedical texts. However, obtaining high-quality annotated biomedical texts is scarce due to their high privacy and the specialized expertise required. Therefore, more researchers are focusing on few-shot NER in the biomedical field. Recent methods mainly fall into three categories: 1) transferring knowledge from high-resource data and fine-tuning with low-resource data, 2) modifying data using various methods to generate new samples, and 3) decompose the NER task into two subtasks with multi-task learning. However, these methods either suffer from domain shift, generate low-quality synthetic data, or encounter error propagation issues. To address these limitations, we have investigated an few-shot NER method, which firstly employs specific prompt templates to guide Large Language Models in generating high-quality new samples. Then, we propose a novel more fine-grained grid-tagging strategy and single-end sensitive multi-task learning framework. This method sets multiple losses to enable the model to learn from multiple task objectives: what constitutes a fully correct, partially correct, and incorrect entity. Additionally, through the new grid-tagging strategy, the model can decode entities based on more tagging clues. To better align with real-world scenarios, we trained in a few-shot scenario and evaluated on the full test set. Extensive experimental results on the NCBI, BC5CDR, BioNLP11EPI, and BioNLP13GE datasets confirm that our method outperforms previous state-of-the-art methods in most scenarios.1 Wenxuan Mu, Di Zhao 0003, Jiana Meng, Shuang Liu 0010, Hongfei Lin |
BIBM | 2 |
| 2024 | Few-shot biomedical relation extraction using data augmentation and domain information
Bocheng Guo, Di Zhao 0003, Jiana Meng, Hongfei Lin |
Neurocomputing | 2 |
| 2024 | Integrating graph convolutional networks to enhance prompt learning for biomedical relation extraction
Bocheng Guo, Jiana Meng, Di Zhao 0003, Xiangxing Jia, Yonghe Chu, Hongfei Lin |
J. Biomed. Informatics | 3 |
| 2023 | Few-shot biomedical named entity recognition via knowledge-guided instance generation and prompt contrastive learningabstractMOTIVATION: Few-shot learning that can effectively perform named entity recognition in low-resource scenarios has raised growing attention, but it has not been widely studied yet in the biomedical field. In contrast to high-resource domains, biomedical named entity recognition (BioNER) often encounters limited human-labeled data in real-world scenarios, leading to poor generalization performance when training only a few labeled instances. Recent approaches either leverage cross-domain high-resource data or fine-tune the pre-trained masked language model using limited labeled samples to generate new synthetic data, which is easily stuck in domain shift problems or yields low-quality synthetic data. Therefore, in this article, we study a more realistic scenario, i.e. few-shot learning for BioNER. RESULTS: Leveraging the domain knowledge graph, we propose knowledge-guided instance generation for few-shot BioNER, which generates diverse and novel entities based on similar semantic relations of neighbor nodes. In addition, by introducing question prompt, we cast BioNER as question-answering task and propose prompt contrastive learning to improve the robustness of the model by measuring the mutual information between query-answer pairs. Extensive experiments conducted on various few-shot settings show that the proposed framework achieves superior performance. Particularly, in a low-resource scenario with only 20 samples, our approach substantially outperforms recent state-of-the-art models on four benchmark datasets, achieving an average improvement of up to 7.1% F1. AVAILABILITY AND IMPLEMENTATION: Our source code and data are available at https://github.com/cpmss521/KGPC. Jian Wang 0021, Hongfei Lin, Di Zhao 0003 |
Bioinform. | 4 |
| 2023 | ADPG: Biomedical entity recognition based on Automatic Dependency Parsing Graph
Hongfei Lin, Yi-Jia Zhang 0001, Di Zhao 0003, Shuaiheng Huai |
J. Biomed. Informatics | 5 |
| 2023 | Biomedical document relation extraction with prompt learning and KNN
Di Zhao 0003, Jiana Meng, Shichang Sun, Jian Wang 0021, Hongfei Lin |
J. Biomed. Informatics | 1 |
| 2022 | Syntactic Type-aware Graph Attention Network for Drug-drug Interactions and their Adverse Effects ExtractionabstractAutomatic extraction of drug-drug interactions and their adverse effects can promote the research of pharmacovigilance and thus attracts attention from both academia and industry. Recent efforts focus on span-based approaches and show more promising results. However, span-based methods enumerate all possible candidate entity spans while ignoring boundary information of spans. Meanwhile, lacking sufficient interactions in intra-span and inter-span further hinders the performance of the nested entity and overlapping relation extraction. To this end, we propose a syntactic type-aware graph attention network for drug-drug interactions and their adverse effects extraction. Specifically, a boundary heuristic module is designed firstly to generate the boundary of linguistically legitimate entity spans. And then, different from the general syntactic graph (i.e., only considering dependency edges), we construct a syntactic type-aware graph attention network (STG) to capture interactions in intra-span and inter-span by considering syntactic edges and types simultaneously. Results1achieved on two biomedical benchmark datasets, including drug-drug interaction (DDI) and adverse drug effect (ADE), indicate that our model obtains significantly more performance than the state-of-the-art methods, achieving improvements in the relation F1 score of 1.63% on ADE and 2.07% on DDI dataset, respectively. Jian Wang 0021, Hongfei Lin, Di Zhao 0003, Yi-Jia Zhang 0001 |
BIBM | 5 |
| 2022 | Refining electronic medical records representation in manifold subspaceabstractBACKGROUND: Electronic medical records (EMR) contain detailed information about patient health. Developing an effective representation model is of great significance for the downstream applications of EMR. However, processing data directly is difficult because EMR data has such characteristics as incompleteness, unstructure and redundancy. Therefore, preprocess of the original data is the key step of EMR data mining. The classic distributed word representations ignore the geometric feature of the word vectors for the representation of EMR data, which often underestimate the similarities between similar words and overestimate the similarities between distant words. This results in word similarity obtained from embedding models being inconsistent with human judgment and much valuable medical information being lost. RESULTS: In this study, we propose a biomedical word embedding framework based on manifold subspace. Our proposed model first obtains the word vector representations of the EMR data, and then re-embeds the word vector in the manifold subspace. We develop an efficient optimization algorithm with neighborhood preserving embedding based on manifold optimization. To verify the algorithm presented in this study, we perform experiments on intrinsic evaluation and external classification tasks, and the experimental results demonstrate its advantages over other baseline methods. CONCLUSIONS: Manifold learning subspace embedding can enhance the representation of distributed word representations in electronic medical record texts. Reduce the difficulty for researchers to process unstructured electronic medical record text data, which has certain biomedical research value. Yuanyuan Sun 0002, Yonghe Chu, Di Zhao 0003, Jian Wang 0021 |
BMC Bioinform. | 4 |
| 2022 | Manifold biomedical text sentence embedding
Yuanyuan Sun 0002, Yonghe Chu, Hongfei Lin, Di Zhao 0003, Liang Yang 0003, Chen Shen 0001, Jian Wang 0021 |
Neurocomputing | 5 |
| 2022 | A Scalable Embedding Based Neural Network Method for Discovering Knowledge From Biomedical LiteratureabstractNowadays, the amount of biomedical literatures is growing at an explosive speed, and much useful knowledge is yet undiscovered in the literature. Classical information retrieval techniques allow to access explicit information from a given collection of information, but are not able to recognize implicit connections. Literature-based discovery (LBD) is characterized by uncovering hidden associations in non-interacting literature. It could significantly support scientific research by identifying new connections between biomedical entities. However, most of the existing approaches to LBD are not scalable and may not be sufficient to detect complex associations in non-directly-connected literature. In this article, we present a model which incorporates biomedical knowledge graph, graph embedding, and deep learning methods for literature-based discovery. First, the relations between biomedical entities are extracted from biomedical abstracts and then a knowledge graph is constructed by using these obtained relations. Second, the graph embedding technologies are applied to convert the entities and relations in the knowledge graph into a low-dimensional vector space. Third, a bidirectional Long Short-Term Memory (BLSTM) network is trained based on the entity associations represented by the pre-trained graph embeddings. Finally, the learned model is used for open and closed literature-based discovery tasks. The experimental results show that our method could not only effectively discover hidden associations between entities, but also reveal the corresponding mechanism of interactions. It suggests that incorporating knowledge graph and deep learning methods is an effective way for capturing the underlying complex associations between entities hidden in the literature. Shengtian Sang, Di Zhao 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Co-Attentive Span Network with Multi-task learning for Biomedical Named Entity RecognitionabstractBiomedical Named Entity Recognition (BioNER) is often modeled as a sequence labeling task, which assigns the predefined label to each token in given input sequence. Although these sequential labeling models achieve significant achievements, they often fail to give precise boundaries of the named entity. In addition, a vast amount of work focuses more on textual sequence representation but ignores label information. To tackle these problems, in this paper, we directly model span-level named entity recognition, specifically, we treat the BioNER as a joint task of boundary detection and span classification under a multitask framework. In order to enhance boundary supervision, we introduce an entity type label as an additional guide and propose a co-attentive interactive mechanism to improve the span representation. Extensive experiments1on four benchmark datasets demonstrate that our proposed method obtains competitive results, achieving 90.26%, 78.04%, 90.21%, and 86.58% on BC5CDR, JNLPBA, NCBI, and BC2GM datasets, respectively, in terms of F1 score. Jian Wang 0021, Hongfei Lin, Yi-Jia Zhang 0001, Di Zhao 0003, Hui Ma 0011 |
BIBM | 6 |
| 2021 | Improving biomedical word representation with locally linear embedding
Di Zhao 0003, Jian Wang 0021, Yonghe Chu, Yi-Jia Zhang 0001, Hongfei Lin |
Neurocomputing | 1 |
| 2021 | Sentence representation with manifold learning for biomedical texts
Di Zhao 0003, Jian Wang 0021, Hongfei Lin, Yonghe Chu, Yi-Jia Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2020 | Incorporating representation learning and multihead attention to improve biomedical cross-sentence n-ary relation extractionabstractBACKGROUND: Most biomedical information extraction focuses on binary relations within single sentences. However, extracting n-ary relations that span multiple sentences is in huge demand. At present, in the cross-sentence n-ary relation extraction task, the mainstream method not only relies heavily on syntactic parsing but also ignores prior knowledge. RESULTS: In this paper, we propose a novel cross-sentence n-ary relation extraction method that utilizes the multihead attention and knowledge representation that is learned from the knowledge graph. Our model is built on self-attention, which can directly capture the relations between two words regardless of their syntactic relation. In addition, our method makes use of entity and relation information from the knowledge base to impose assistance while predicting the relation. Experiments on n-ary relation extraction show that combining context and knowledge representations can significantly improve the n-ary relation extraction performance. Meanwhile, we achieve comparable results with state-of-the-art methods. CONCLUSIONS: We explored a novel method for cross-sentence n-ary relation extraction. Unlike previous approaches, our methods operate directly on the sequence and learn how to model the internal structures of sentences. In addition, we introduce the knowledge representations learned from the knowledge graph into the cross-sentence n-ary relation extraction. Experiments based on knowledge representation learning show that entities and relations can be extracted in the knowledge graph, and coding this knowledge can provide consistent benefits. Di Zhao 0003, Jian Wang 0021, Yi-Jia Zhang 0001, Hongfei Lin |
BMC Bioinform. | 1 |
| 2019 | Extracting drug-drug interactions with hybrid bidirectional gated recurrent unit and graph convolutional network
Di Zhao 0003, Jian Wang 0021, Hongfei Lin, Yi-Jia Zhang 0001 |
J. Biomed. Informatics | 1 |