VLDB 2026 Research / reviewers in the wild / expert
Weiru Fu
dblp:361/2188
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0004-9511-5923ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CAKADE: Improving ADE Detection on Social Media with LLMs via Counterfactual Augmentation and Knowledge-Enhanced Instruction TuningabstractAutomatic detection of Adverse Drug Events (ADEs) from social media has become increasingly important for post-market drug safety surveillance and pharmacovigilance. Although existing social media-based ADE detection methods have effectively addressed some challenges such as data sparsity and class imbalance, they still suffer from spurious correlations, where models tend to learn co-occurrence patterns between drugs and symptoms, leading them to incorrectly identify drug inefficacy or therapeutic intent as ADEs. To address this issue, we propose a structured medical knowledge-guided counterfactual generation method that leverages authoritative medical databases and large language models to construct clinically plausible counterfactual samples, thereby mitigating spurious correlations at the data level. Furthermore, we propose a knowledge-enhanced instruction-tuning strategy that injects drug indications as causal cues into model inputs to enhance its causal reasoning capabilities. Experimental results on two benchmark datasets demonstrate that our method consistently outperforms state-of-the-art models, effectively alleviating spurious correlations and exhibiting superior capabilities in drug-symptom relationship identification. Weiru Fu, Yunzhi Qiu, Ling Luo 0001, Jian Wang 0021, Hongfei Lin |
BIBM | 1 |
| 2025 | Can Herpes Zoster Vaccine Reduce Alzheimer's Disease Risk? A KG and LLMS Synergistic Integration ApproachabstractCurrent research reveals that the herpes zoster vaccine(HZV) can effectively prevent and mitigate Alzheimer's disease(AD). However, the underlying mechanisms linking the HZV and AD remain incompletely understood. We employ a literature-based discovery (LBD) paradigm to investigate these mechanisms. Traditional knowledge graph-based approaches demonstrate strong performance in implicit knowledge discovery tasks. However, existing methods still face critical challenges: (1) The diverse representations of biomedical entities often degrade knowledge graph quality, such as introducing path redundancy; (2) Current path-ranking mechanisms lack rigorous biological plausibility assessment, resulting in a high number of false-positive paths being retained. To address the aforementioned challenges, this paper proposes a Synergistic Integration framework of Knowledge graphs and Large language models (SIKL). First, the PubTator3.0 tool and manual rule-based methods are employed to extract entities and relational triples from abstract content. A knowledge graph is constructed with medical subject headings as nodes, integrating multiple attributes and relations, effectively mitigating path redundancy caused by diverse entity expressions. Next, a twostage LLM-driven path filtering module is designed. In the first stage, a retrieval-augmented large language model assesses whether candidate paths meet causality conditions, performing preliminary filtering to eliminate false positives. The second stage leverages a chain-of-thought large language model to conduct step-by-step reasoning on remaining paths, further evaluating their biological plausibility. Finally, a comprehensive path ranking strategy combines literature support counts and LLM-generated scores to output Top-p high-confidence hypothetical paths. Experimental results reveals that HZV may delay AD progression through pathways such as neuroinflammatory regulation, with partial mechanisms supported by literature. Yunzhi Qiu, Weiru Fu, Ling Luo 0001, Hongfei Lin |
BIBM | 2 |
| 2025 | ADENER: A syntax-augmented grid-tagging model for Adverse Drug Event extraction in social media
Weiru Fu, Ling Luo 0001, Hongfei Lin |
J. Biomed. Informatics | 1 |
| 2024 | CFAH: A Chinese Dataset for Detecting False Advertising in Healthcare
Weiru Fu, Junyu Lu 0001, Youlin Wu, Guangtao Xu, Liang Yang 0003, Hongfei Lin, Jian Wang 0021, Ruiyuan Wang |
BIBM | 1 |
| 2024 | Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasksabstractOBJECTIVE: Most existing fine-tuned biomedical large language models (LLMs) focus on enhancing performance in monolingual biomedical question answering and conversation tasks. To investigate the effectiveness of the fine-tuned LLMs on diverse biomedical natural language processing (NLP) tasks in different languages, we present Taiyi, a bilingual fine-tuned LLM for diverse biomedical NLP tasks. MATERIALS AND METHODS: We first curated a comprehensive collection of 140 existing biomedical text mining datasets (102 English and 38 Chinese datasets) across over 10 task types. Subsequently, these corpora were converted to the instruction data used to fine-tune the general LLM. During the supervised fine-tuning phase, a 2-stage strategy is proposed to optimize the model performance across various tasks. RESULTS: Experimental results on 13 test sets, which include named entity recognition, relation extraction, text classification, and question answering tasks, demonstrate that Taiyi achieves superior performance compared to general LLMs. The case study involving additional biomedical NLP tasks further shows Taiyi's considerable potential for bilingual biomedical multitasking. CONCLUSION: Leveraging rich high-quality biomedical corpora and developing effective fine-tuning strategies can significantly improve the performance of LLMs within the biomedical domain. Taiyi shows the bilingual multitasking capability through supervised fine-tuning. However, those tasks such as information extraction that are not generation tasks in nature remain challenging for LLM-based generative approaches, and they still underperform the conventional discriminative approaches using smaller language models. Ling Luo 0001, Jinzhong Ning, Yingwen Zhao, Zeyuan Ding, Weiru Fu, Qinyu Han, Guangtao Xu, Yunzhi Qiu, Dinghao Pan, Jiru Li, Wenduo Feng, Senbo Tu, Jian Wang 0021, Yuanyuan Sun 0002, Hongfei Lin |
J. Am. Medical Informatics Assoc. | 7 |