VLDB 2026 Research / reviewers in the wild / expert
Yuzhang Xie
dblp:337/4335
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0001-3241-9418ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Utilizing Large Language Models for Zero-Shot Medical Ontology Extension from Clinical NotesabstractIntegrating novel medical concepts and relationships into existing ontologies can significantly enhance their coverage and utility for both biomedical research and clinical applications. Clinical notes, as unstructured documents rich with detailed patient observations, offer valuable context-specific insights and represent a promising yet underutilized source for ontology extension. Despite this potential, directly leveraging clinical notes for ontology extension remains largely unexplored. To address this gap, we propose CLOZE, a novel framework that uses large language models (LLMs) to automatically extract medical entities from clinical notes and integrate them into hierarchical medical ontologies. By capitalizing on the strong language understanding and extensive biomedical knowledge of pre-trained LLMs, CLOZE effectively identifies disease-related concepts and captures complex hierarchical relationships. The zero-shot framework requires no additional training or labeled data, making it a cost-efficient solution. Furthermore, CLOZE ensures patient privacy through automated removal of protected health information (PHI). Experimental results demonstrate that CLOZE provides an accurate, scalable, and privacy-preserving ontology extension framework, with strong potential to support a wide range of downstream applications in biomedical research and clinical informatics. Guanchen Wu, Yuzhang Xie, Huanwei Wu, Zhe He 0001, Xiao Hu 0002, Carl Yang 0001 |
BIBM | 2 |
| 2025 | HypKG: Hypergraph-Based Knowledge Graph Contextualization for Precision Healthcare
Yuzhang Xie, Ran Xu 0002, Xiao Hu 0002, Jiaying Lu 0001, Carl Yang 0001 |
ISWC (1) | 1 |
| 2024 | TACCO: Task-guided Co-clustering of Clinical Concepts and Patient Visits for Disease Subtyping based on EHR DataabstractThe growing availability of well-organized Electronic Health Records (EHR) data has enabled the development of various machine learning models towards disease risk prediction. However, existing risk prediction methods overlook the heterogeneity of complex diseases, failing to model the potential disease subtypes regarding their corresponding patient visits and clinical concept subgroups. In this work, we introduce TACCO, a novel framework that jointly discovers clusters of clinical concepts and patient visits based on a hypergraph modeling of EHR data. Specifically, we develop a novel self-supervised co-clustering framework that can be guided by the risk prediction task of specific diseases. Furthermore, we enhance the hypergraph model of EHR data with textual embeddings and enforce the alignment between the clusters of clinical concepts and patient visits through a contrastive objective. Comprehensive experiments conducted on the public MIMIC-III dataset and Emory internal CRADLE dataset over the downstream clinical tasks of phenotype classification and cardiovascular risk prediction demonstrate an average 31.25% performance improvement compared to traditional ML baselines and a 5.26% improvement on top of the vanilla hypergraph model without our co-clustering mechanism. In-depth model analysis, clustering results analysis, and clinical case studies further validate the improved utilities and insightful interpretations delivered by TACCO. Code is available at https://github.com/PericlesHat/TACCO. Hejie Cui, Ran Xu 0002, Yuzhang Xie, Joyce C. Ho, Carl Yang 0001 |
KDD | 4 |
| 2024 | PromptLink: Leveraging Large Language Models for Cross-Source Biomedical Concept LinkingabstractLinking (aligning) biomedical concepts across diverse data sources enables various integrative analyses, but it is challenging due to the discrepancies in concept naming conventions. Various strategies have been developed to overcome this challenge, such as those based on string-matching rules, manually crafted thesauri, and machine learning models. However, these methods are constrained by limited prior biomedical knowledge and can hardly generalize beyond the limited amounts of rules, thesauri, or training samples. Recently, large language models (LLMs) have exhibited impressive results in diverse biomedical NLP tasks due to their unprecedentedly rich prior knowledge and strong zero-shot prediction abilities. However, LLMs suffer from issues including high costs, limited context length, and unreliable predictions. In this research, we propose PromptLink, a novel biomedical concept linking framework that leverages LLMs. It first employs a biomedical-specialized pre-trained language model to generate candidate concepts that can fit in the LLM context windows. Then it utilizes an LLM to link concepts through two-stage prompts, where the first-stage prompt aims to elicit the biomedical prior knowledge from the LLM for the concept linking task and the second-stage prompt enforces the LLM to reflect on its own predictions to further enhance their reliability. Empirical results on the concept linking task between two EHR datasets and an external biomedical KG demonstrate the effectiveness of PromptLink. Furthermore, PromptLink is a generic framework without reliance on additional prior knowledge, context, or training data, making it well-suited for concept linking across various types of data sources. The source code of this study is available at https://github.com/constantjxyz/PromptLink. Yuzhang Xie, Jiaying Lu 0001, Joyce C. Ho, Fadi B. Nahab, Xiao Hu 0002, Carl Yang 0001 |
SIGIR | 1 |
| 2024 | Improving diagnosis and outcome prediction of gastric cancer via multimodal learning using whole slide pathological images and gene expression
Yuzhang Xie, Qingqing Sang, Qian Da, Guoshuai Niu, Yunqin Chen, Bing-Ya Liu, Yang Yang 0030, Wentao Dai |
Artif. Intell. Medicine | 1 |
| 2022 | Survival Prediction for Gastric Cancer via Multimodal Learning of Whole Slide Images and Gene ExpressionabstractGastric cancer (GC) is one of the most common malignancies worldwide. As histopathology tissue analysis is considered as the gold standard in cancer studies, whole slide images (WSIs) have been widely used for GC diagnosis and prognosis, while multimodal studies for GC patients have been very few. Especially, WSIs and gene expression are complementary modalities of data, thus fusion of these two modalities has great potential in the prediction of survival outcomes and other computer-aided tasks, like the mechanism study and clinical treatment for GC patients. However, multimodal learning requires good data fusion strategies and also suffers from the missing data issue. To address these issues, we propose GC-SPLeM, to predict risk scores for patients, which consists of three parts, WSI feature extraction, modal-fusing network, and GNN-based predictor. We conduct experiments on a GC dataset built by ourselves and a public dataset for survival prediction. For both datasets, GC-SPLeM outperforms the state-of-the-art single-modality learning method and multimodal learning method by large margins (over 5% on C-index). We find that the GNN plays an important role in performance enhancement. Through learning the graph of patients, topological structure and neighborhood clinical information are encoded into feature representations of patients. GC-SPLeM not only improves the survival prediction results but also has advantages in dealing with incomplete data over other methods. The Source code, sample data, and gene list of this study are available at https://github.com/constantjxyz/GC-SPLeM. Yuzhang Xie, Guoshuai Niu, Qian Da, Wentao Dai, Yang Yang 0030 |
BIBM | 1 |