Jiaying Lu 0001

dblp:61/9803-1 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0001-9052-6951ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Integrating Epigenetic and Phenotypic Features for Biological Age Estimation in Cancer Patients via Multimodal Learning
abstract
Biological age, which may be older or younger than chronological age due to factors such as genetic predisposition, environmental exposures, serves as a meaningful biomarker of aging processes and can inform risk stratification, treatment planning, and survivorship care in cancer patients. We propose EpiCAge, a multimodal framework that integrates epigenetic and phenotypic data to improve biological age prediction. Evaluated on eight internal and four external cancer cohorts, EpiCAge consistently outperforms existing epigenetic and phenotypic age clocks. Our analyses show that EpiCAge identifies biologically relevant markers, and its derived age acceleration is significantly associated with mortality risk. These results highlight EpiCAge as a promising multimodal machine learning tool for biological age assessment in oncology.
Shuyue Jiang, Wenjing Ma, Shaojun Yu, Runze Yan, Jiaying Lu 0001
BIBM6
2025 ScPanKD: Distilling Pan-Cancer Knowledge for Enhanced T Cell Subtypes Annotation in Single-Cell Transcriptomics Data
abstract
Single-cell RNA sequencing (scRNA-seq) enables high-resolution characterization of cellular heterogeneity, and annotating major cell types has become a standard practice in scRNA -seq analysis pipelines. However, accurately identifying fine-grained subtypes within major cell types remains challenging, particularly in heterogeneous tissues such as cancer samples. Here, we present ScPanKD, a computational frame-work for accurate and robust classification of fine-grained T cell subtypes in cancer samples. Unlike existing methods that suffer from cell type mismatches between reference and query datasets due to cancer heterogeneity, ScPanKD leverages knowledge distillation (KD) to accurately identify T cell subtypes even when the reference dataset contains more subtype diversity than the query. ScPanKD learns a cancer-invariant feature space and employs a two-step strategy, anchor cell selection followed by KD, to mitigate distribution shifts between reference and query datasets. Across extensive experiments using pan-cancer level CD4+ or CD8+ T cell atlases as references, we demonstrate that ScPanKD outperforms conventional annotation methods and single-cell foundation models, achieving more accurate and robust T cell subtype classification. ScPanKD and all reproducible scripts are available at https://github.com/marvinquiet/ScPanKD.
Wenjing Ma, Xiaoqing Yu, Jiaying Lu 0001
BIBM3
2025 HypKG: Hypergraph-Based Knowledge Graph Contextualization for Precision Healthcare
Yuzhang Xie, Ran Xu 0002, Xiao Hu 0002, Jiaying Lu 0001, Carl Yang 0001
ISWC (1)5
2025 A review on knowledge graphs for healthcare: Resources, applications, and promises
Hejie Cui, Jiaying Lu 0001, Ran Xu 0002, Shiyu Wang 0002, Wenjing Ma, Yue Yu 0001, Shaojun Yu, Xuan Kan, Chen Ling 0003, Liang Zhao 0002, Zhaohui S. Qin, Joyce C. Ho, Tianfan Fu, Jing Ma 0005, Mengdi Huai, Carl Yang 0001
J. Biomed. Informatics2
2024 PromptLink: Leveraging Large Language Models for Cross-Source Biomedical Concept Linking
abstract
Linking (aligning) biomedical concepts across diverse data sources enables various integrative analyses, but it is challenging due to the discrepancies in concept naming conventions. Various strategies have been developed to overcome this challenge, such as those based on string-matching rules, manually crafted thesauri, and machine learning models. However, these methods are constrained by limited prior biomedical knowledge and can hardly generalize beyond the limited amounts of rules, thesauri, or training samples. Recently, large language models (LLMs) have exhibited impressive results in diverse biomedical NLP tasks due to their unprecedentedly rich prior knowledge and strong zero-shot prediction abilities. However, LLMs suffer from issues including high costs, limited context length, and unreliable predictions. In this research, we propose PromptLink, a novel biomedical concept linking framework that leverages LLMs. It first employs a biomedical-specialized pre-trained language model to generate candidate concepts that can fit in the LLM context windows. Then it utilizes an LLM to link concepts through two-stage prompts, where the first-stage prompt aims to elicit the biomedical prior knowledge from the LLM for the concept linking task and the second-stage prompt enforces the LLM to reflect on its own predictions to further enhance their reliability. Empirical results on the concept linking task between two EHR datasets and an external biomedical KG demonstrate the effectiveness of PromptLink. Furthermore, PromptLink is a generic framework without reliance on additional prior knowledge, context, or training data, making it well-suited for concept linking across various types of data sources. The source code of this study is available at https://github.com/constantjxyz/PromptLink.
Yuzhang Xie, Jiaying Lu 0001, Joyce C. Ho, Fadi B. Nahab, Xiao Hu 0002, Carl Yang 0001
SIGIR2
2023 Closed-book Question Generation via Contrastive Learning
abstract
Question Generation (QG) is a fundamental NLP task for many downstream applications.Recent studies on open-book QG, where supportive answer-context pairs are provided to models, have achieved promising progress.However, generating natural questions under a more practical closed-book setting that lacks these supporting documents still remains a challenge.In this work, we propose a new QG model for this closed-book setting that is designed to better understand the semantics of long-form abstractive answers and store more information in its parameters through contrastive learning and an answer reconstruction module.Through experiments, we validate the proposed QG model on both public datasets and a new WikiCQA dataset.Empirical results show that the proposed QG model outperforms baselines in both automatic evaluation and human evaluation.In addition, we show how to leverage the proposed model to improve existing question-answering systems.These results further indicate the effectiveness of our QG model for enhancing closed-book questionanswering tasks.
Xiangjue Dong, Jiaying Lu 0001, Jianling Wang, James Caverlee
EACL2
2023 HiPrompt: Few-Shot Biomedical Knowledge Fusion via Hierarchy-Oriented Prompting
abstract
Medical decision-making processes can be enhanced by comprehensive biomedical knowledge bases, which require fusing knowledge graphs constructed from different sources via a uniform index system. The index system often organizes biomedical terms in a hierarchy to provide the aligned entities with fine-grained granularity. To address the challenge of scarce supervision in the biomedical knowledge fusion (BKF) task, researchers have proposed various unsupervised methods. However, these methods heavily rely on ad-hoc lexical and structural matching algorithms, which fail to capture the rich semantics conveyed by biomedical entities and terms. Recently, neural embedding models have proved effective in semantic-rich tasks, but they rely on sufficient labeled data to be adequately trained. To bridge the gap between the scarce-labeled BKF and neural embedding models, we propose HiPrompt, a supervision-efficient knowledge fusion framework that elicits the few-shot reasoning ability of large language models through hierarchy-oriented prompts. Empirical results on the collected KG-Hi-BKF benchmark datasets demonstrate the effectiveness of HiPrompt.
Jiaying Lu 0001, Bo Xiong 0001, Wenjing Ma, Steffen Staab, Carl Yang 0001
SIGIR1
2023 Weakly Supervised Concept Map Generation Through Task-Guided Graph Translation
abstract
Recent years have witnessed the rapid development of concept map generation techniques due to their advantages in providing well-structured summarization of knowledge from free texts. Traditional unsupervised methods do not generate task-oriented concept maps, whereas deep generative models require large amounts of training data. In this work, we presentGT-D2G(Graph Translation-based Document To Graph), an automatic concept map generation framework that leverages generalized NLP pipelines to derive semantic-rich initial graphs, and translates them into more concise structures under the weak supervision of downstream task labels. The concept maps generated byGT-D2Gcan provide interpretable summarization of structured knowledge for the input texts, which are demonstrated through human evaluation and case studies on three real-world corpora. Further experiments on the downstream task of document classification show thatGT-D2Gbeats other concept map generation methods. Moreover, we specifically validate the labeling efficiency ofGT-D2Gin the label-efficient learning setting and the flexibility of generated graph sizes in controlled hyper-parameter studies.
Jiaying Lu 0001, Xiangjue Dong, Carl Yang 0001
IEEE Trans. Knowl. Data Eng.1
2022 How Can Graph Neural Networks Help Document Retrieval: A Case Study on CORD19 with Concept Map Generation
Hejie Cui, Jiaying Lu 0001, Yao Ge 0003, Carl Yang 0001
ECIR (2)2