EDBT 2026 Demo / reviewers in the wild / expert
Hang Lv 0010
dblp:369/5929
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
0009-0007-2566-390XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Awaken the Giant: Activating LLMs via Deep Model Guidance for Boundary-aware Medication Recommendation
Hang Lv 0010, Yanchao Tan, Wanzi Shao, Hengyu Zhang 0005, Carl Yang 0001 |
KDD (1) | 1 |
| 2026 | EviCare: Enhancing Diagnosis Prediction with Deep Model-Guided Evidence for In-Context ReasoningabstractRecent advances in large language models (LLMs) have enabled promising progress in diagnosis prediction from electronic health records (EHRs). However, existing LLM-based approaches tend to overfit to historically observed diagnoses, often overlooking novel yet clinically important conditions that are critical for early intervention. To address this, we propose EviCare, an in-context reasoning framework that integrates deep model guidance into LLM-based diagnosis prediction. Rather than prompting LLMs directly with raw EHR inputs, EviCare performs (1) deep model inference for candidate selection, (2) evidential prioritization for set-based EHRs, and (3) relational evidence construction for novel diagnosis prediction. These signals are then composed into an adaptive in-context prompt to guide LLM reasoning in an accurate and interpretable manner. Extensive experiments on two real-world EHR benchmarks (MIMIC-III and MIMIC-IV) demonstrate that EviCare achieves significant performance gains, which consistently outperforms both LLM-only and deep model-only baselines by an average of 20.65% across precision and accuracy metrics. The improvements are particularly notable in challenging novel diagnosis prediction, yielding average improvements of 30.97%. Hengyu Zhang 0005, Xuyun Zhang, Pengxiang Zhan, Linhao Luo, Hang Lv 0010, Yanchao Tan, Shirui Pan, Carl Yang 0001 |
KDD (1) | 5 |
| 2026 | Towards Efficient and Interpretable Medical Concept Representation via Ontology-driven Residual Vector QuantizationabstractMedical concepts, the core entities in Electronic Health Records (EHRs), provide essential inputs for clinical decision-making systems. However, most existing healthcare models still rely on massive concept-specific embedding tables, resulting in substantial memory overhead. Recent studies compress medical concepts into discrete code sequences for memory efficiency, but their flat semantic quantization fails to explicitly encode the hierarchical structure of medical ontologies, thereby limiting clinical interpretability. To this end, we propose MedRQ, an ontology-driven residual vector quantization framework that aligns discrete codes with multi-level clinical ontologies. By incorporating hierarchical supervision into the quantization process, MedRQ generates compact and ontology-consistent concept representations that generalize seamlessly across healthcare prediction tasks. Experiments on two real-world EHR datasets demonstrate that MedRQ significantly outperforms state-of-the-art baselines while reducing memory usage. Hang Lv 0010, Kaisong Zhang, Yanchao Tan, Xing Chen 0002 |
WWW | 1 |
| 2026 | Graph-Oriented Instruction Tuning of Large Language Models for Generic Graph MiningabstractGraphs with abundant attributes are essential in modeling interconnected entities and enhancing predictions across various real-world applications. Traditional Graph Neural Networks (GNNs) often require re-training for different graph tasks and datasets. Although the emergence of Large Language Models (LLMs) has introduced new paradigms in natural language processing, their potential for generic graph mining-training a single model to simultaneously handle diverse tasks and datasets-remains under-explored. To this end, our novel framework ${\sf MuseGraph}$MuseGraph, seamlessly integrates the strengths of GNNs and LLMs into one foundation model for graph mining across tasks and datasets. This framework first features a compact graph description to encapsulate key graph information within language token limitations. Then, we propose a diverse instruction generation mechanism with Chain-of-Thought (CoT)-based instruction packages to distill the reasoning capabilities from advanced LLMs like GPT-4. Finally, we design a graph-aware instruction tuning strategy to facilitate mutual enhancement across multiple tasks and datasets while preventing catastrophic forgetting of LLMs' generative abilities. Our experimental results demonstrate significant improvements in five graph tasks and ten datasets, showcasing the potential of our ${\sf MuseGraph}$MuseGraph in enhancing the accuracy of graph-oriented downstream tasks while improving the generation abilities of LLMs. Yanchao Tan, Hang Lv 0010, Pengxiang Zhan, Shiping Wang, Carl Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | DMAP: A Deep-Model Driven Multi-Agent Framework for Reliable Diagnosis PredictionabstractDeep learning based diagnosis prediction from Electronic Health Record (EHR) data has demonstrated strong performance but remains hindered by limited interpretability, reducing its clinical applicability. Existing multi-agent frameworks attempt to address this challenge; however, they often rely on superficial, semantics-only predictions from a large language model (LLM) or adopt multi-classification heads, thereby inheriting the opacity of conventional deep learning approaches. We introduce DMAP, a multi-agent framework designed to enhance EHR modeling through structured collaboration and interpretable decision-making. Inspired by multidisciplinary clinical workflows, DMAP employs three types of agents: DeepL Agents, a Leader Agent, and a Critical Agent, which work together to analyze patient data and generate reliable reports. DeepL Agents process structured EHR inputs, model outputs, and interpretability signals to provide initial clinical evaluations. Leader Agent coordinates these insights through iterative discussions, synthesizing consensus-driven disease predictions and draft report. Critical Agent evaluates the medical soundness of the draft report, identifying gaps or inconsistencies. When uncertainty arises, a retrieval-augmented generation (RAG) module integrates medical knowledge to support final decisions. By simulating expert collaboration and integrating structured reasoning with external knowledge, DMAP delivers reliable, interpretable predictions and reports. Extensive experiments on two EHR datasets demonstrate its superior performance in disease prediction and report generation, highlighting its potential to advance clinical decision support through explainable and adaptive multi-agent collaboration. Wenjing Yu, Guofang Ma, Hang Lv 0010, Zhigang Lin, Xiping Chen, Yanchao Tan |
BIBM | 3 |
| 2025 | RQCare: A Residual Quantization Model for Disease Representation and Diagnosis Prediction in Healthcare DataabstractElectronic Health Records (EHRs) offer rich longitudinal data for clinical prediction, but existing deep models struggle to jointly capture semantic richness and hierarchical structure. We propose RQCare, a novel framework that integrates semantic and structural information for interpretable hierarchical disease representation learning. RQCare consists of three components: (1) a Disease Embedding Residual Quantization module that learns discrete, interpretable hierarchies from embeddings; (2) a Dual-Graph Structure Learning module that refines representations using both patient-disease interactions and ICD hierarchy; (3) a GRU-based temporal model with attention for next-visit prediction. Evaluated on MIMIC-III and MIMIC-IV, RQCare outperforms state-of-the-art baselines, achieving up to 6.06% and 2.84% gains in Precision@10, respectively. Xusheng Yu, Kaisong Zhang, Hang Lv 0010, Guofang Ma, Zhigang Lin, Xiping Chen, Yanchao Tan |
BIBM | 3 |
| 2025 | $\mathrm{D}^{2}$ KGMed: Dynamic Diagnostic Knowledge Graphs for Medical Diagnosis PredictionabstractAccurate diagnosis prediction using Electronic Health Records (EHRs) is essential for personalized healthcare. Clinical knowledge graphs (KGs) can enrich EHRs by structuring medical knowledge, and recent work integrates large language models (LLMs) with KGs to enhance reasoning. However, these approaches often depends on static, expensive global graph construction and one-time retrieval, yielding noisy or irrelevant subgraphs that hinder effective diagnosis prediction in real-world clinical scenarios. To this end, we propose$\mathrm{D}^{2}$KGMed, a diagnosis prediction framework that constructs a patient-specific Dynamic Diagnostic Knowledge Graph guided by LLMs. It consists of two stages: constructing an initial graph from diagnostic entities and multi-source medical knowledge; refining its construction via supervised fine-tuning to better align with the ideal graph for conciseness and relevance, and subsequently leveraging it for interpretable predictions. This design reduces graph construction costs and retrieval noise common in KG+LLM methods, enabling more accurate diagnosis prediction. Extensive experiments on two real-world EHR datasets demonstrate that D2KGMed outperforms state-of-the-art baselines, especially in few-shot learning scenarios, showcasing its practical utility in real-world clinical settings. Jie Zhang 0166, Gaoyang Zheng, Hang Lv 0010, Linhao Luo, Guofang Ma, Zhigang Lin, Xiping Chen, Yanchao Tan |
BIBM | 3 |
| 2025 | Higher-order Structure and Semantics-enhanced User Profiling for RecommendationabstractAccurate user profiles are crucial for personalized recommendation systems to mitigate information overload on large-scale online platforms. While recent advances in large language models have enhanced semantic understanding for profile construction through textual artifacts, existing methods often neglect the higher-order structural patterns inherent in user-item interaction graphs-a key limitation for achieving accurate and diverse recommendations. In this paper, we propose SSPRec, a Higher-order Structure and Semantics-enhanced User Profiling for Recommendation. Specifically, we first introduce a multi-hop proximity matrix over item-item transitions, followed by low-rank approximation and clustering to group users based on behavioral similarity. Group-level user profiles are then distilled via representative keywords extracted from co-interacted items, and collaborative embeddings are concurrently learned from the interaction graph. To integrate collaborative signals with language-based profiles, we introduce a cross-view contrastive objective that encourages coherence between structural and semantic representations. Final recommendations are made using a fused user-item similarity score. Extensive experiments on four real-world datasets show that SSPRec not only outperforms baselines in accuracy (with 46.35% improvements), but also remains diverse and robust, even under incomplete interactions. Yanchao Tan, Xinyi Huang 0010, Hang Lv 0010, Hengyu Zhang 0005, Wei Huang 0037, Guofang Ma |
CIKM | 4 |
| 2025 | BoxLM: Unifying Structures and Semantics of Medical Concepts for Diagnosis Prediction in HealthcareabstractLanguage Models (LMs) have advanced diagnosis prediction by leveraging the semantic understanding of medical concepts in Electronic Health Records (EHRs). Despite these advancements, existing LM-based methods often fail to capture the structures of medical concepts (e.g., hierarchy structure from domain knowledge). In this paper, we propose BoxLM, a novel framework that unifies the structures and semantics of medical concepts for diagnosis prediction. Specifically, we propose a structure-semantic fusion mechanism via box embeddings, which integrates both ontology-driven and EHR-driven hierarchical structures with LM-based semantic embeddings, enabling interpretable medical concept representations. Furthermore, in the box-aware diagnosis prediction module, an evolve-and-memorize patient box learning mechanism is proposed to model the temporal dynamics of patient visits, and a volume-based similarity measurement is proposed to enable accurate diagnosis prediction. Extensive experiments demonstrate that BoxLM consistently outperforms state-of-the-art baselines, especially achieving strong performance in few-shot learning scenarios, showcasing its practical utility in real-world clinical settings. Yanchao Tan, Hang Lv 0010, Yunfei Zhan, Guofang Ma, Bo Xiong 0001, Carl Yang 0001 |
ICML | 2 |
| 2025 | MedAlign: Enhancing Combinatorial Medication Recommendation with Multi-modality AlignmentabstractCombinatorial Medication Recommendation (CMR) based on multimodal Electronic Health Records (EHRs) is a promising yet challenging frontier in AI-driven healthcare. Existing approaches usually rely on feature extraction from individual modalities without explicitly aligning information across different data sources. As a result, they may ignore complementary information from other modalities, leading to suboptimal representations for CMR. To this end, we propose MedAlign, a novel combinatorial Medication recommendation framework with multi-modality Alignment. Specifically, we first design a distribution-aware multimodal medication alignment module. This aligns distinct modality distributions of medications within a unified latent space, generating consistent medication representations. Furthermore, we introduce a longitudinal multi-view patient aggregation module, which aggregates the historical visits of patients with multi-view information to form informative patient representations. Finally, we propose a combinatorial medication recommendation module, enabling an accurate and safe medication recommendation combination for each patient. Extensive experiments on two real-world multimodal EHR datasets demonstrate the effectiveness of our MedAlign. Hang Lv 0010, Yanchao Tan, Guofang Ma, Zhigang Lin, Xiping Chen, Hong Cheng 0001, Carl Yang 0001 |
ACM Multimedia | 1 |
| 2024 | Logical Relation Modeling and Mining in Hyperbolic Space for RecommendationabstractThe sparse interactions between users and items have aggravated the difficulty of their representations in recommender systems. Existing methods leverage tags to alleviate the sparsity problem but ignore prevalent logical relations among items and tags (e.g., membership, hierarchy, and exclusion), which can be leveraged to enhance the accuracy of modeling user preferences and conducting recommendations. To this end, we propose to extract logical relations among item tags from existing tag taxonomies and exploit the individual strengths of the Poincaré and the Lorentz models in hyperbolic space for logical relation modeling towards enhanced recommendations. Moreover, we find that the logical relations directly extracted from existing tag taxonomies can be inaccurate and coarse. Therefore, we further devise innovative consistency-based and granularity- based weighting mechanisms based on user behavior patterns for data-driven logical relation mining that can be jointly optimized along with recommendations in an end-to-end fashion. Extensive experiments on four real-world benchmark datasets show drastic performance gains brought by our proposed framework, which constantly achieves an average of 8.25% improvement over state-of-the-art competitors regarding both Recall and NDCG metrics. Insightful case studies further demonstrate that our automatically refined logical relations are highly accurate and interpretable. Yanchao Tan, Hang Lv 0010, Wenzhong Guo, Bo Xiong 0001, Weiming Liu 0005, Chaochao Chen 0001, Shiping Wang, Carl Yang 0001 |
ICDE | 2 |
| 2024 | ExpertODE: Continuous Diagnosis Prediction with Expert Enhanced Neural Ordinary Differential EquationsabstractContinuous diagnosis prediction based on multi-modal electronic health records (EHRs) of patients is a promising yet challenging task for AI in healthcare. Existing studies ignore abundant domain knowledge of diseases (e.g., specific medical terms and their interrelations) in textual EHRs, which fails to accurately predict disease progression and assist in sequential diagnosis prediction. To this end, we first propose an Expert enhanced neural Ordinary Differential Equations (ExpertODE) framework for continuous diagnosis prediction. In particular, we first propose a novel Mixture of Language Experts (MoLE) module to enhance disease embeddings with domain knowledge. Furthermore, we propose a Contrastive Neural Ordinary Differential Equation (CNODE) module to continuously model temporal correlations of disease progression, and implement a unified contrastive learning framework to jointly optimize the domain-based MoLE module and the temporal-based CNODE module. Extensive experiments on two real-world textual EHR datasets show significant performance gains brought by our ExpertODE, yielding average improvements of 3.91% for diagnosis prediction over state-of-the-art competitors. Hengyu Zhang 0006, Hang Lv 0010, Yanchao Tan, Guofang Ma, Fan Wang 0020, Carl Yang 0001 |
ICME | 2 |
| 2023 | WalkLM: A Uniform Language Model Fine-tuning Framework for Attributed Graph EmbeddingabstractGraphs are widely used to model interconnected entities and improve downstream predictions in various real-world applications. However, real-world graphs nowadays are often associated with complex attributes on multiple types of nodes and even links that are hard to model uniformly, while the widely used graph neural networks (GNNs) often require sufficient training toward specific downstream predictions to achieve strong performance. In this work, we take a fundamentally different approach than GNNs, to simultaneously achieve deep joint modeling of complex attributes and flexible structures of real-world graphs and obtain unsupervised generic graph representations that are not limited to specific downstream predictions. Our framework, built on a natural integration of language models (LMs) and random walks (RWs), is straightforward, powerful and data-efficient. Specifically, we first perform attributed RWs on the graph and design an automated program to compose roughly meaningful textual sequences directly from the attributed RWs; then we fine-tune an LM using the RW-based textual sequences and extract embedding vectors from the LM, which encapsulates both attribute semantics and graph structures. In our experiments, we evaluate the learned node embeddings towards different downstream prediction tasks on multiple real-world attributed graph datasets and observe significant improvements over a comprehensive set of state-of-the-art unsupervised node embedding methods. We believe this work opens a door for more sophisticated technical designs and empirical evaluations toward the leverage of LMs for the modeling of real-world graphs. Yanchao Tan, Hang Lv 0010, Weiming Liu 0005, Carl Yang 0001 |
NeurIPS | 3 |