EDBT 2026 Demo / reviewers in the wild / expert
Qiuhao Lu
dblp:213/8514
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-2368-8410ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
entity linking |
0.5 | 1 | 2021 | LNN-EL: A Neuro-Symbolic Approach to Short-text Entity Linking · ACL/IJCNLP (1) 2021 |
Medical and health informatics
clinical text processing |
0.5 | 1 | 2021 | Predicting Patient Readmission Risk from Medical Text via Knowledge Graph Enhanced Multiview Graph Convolution · SIGIR 2021 |
Medical and health informatics
electronic health records |
0.5 | 1 | 2021 | Predicting Patient Readmission Risk from Medical Text via Knowledge Graph Enhanced Multiview Graph Convolution · SIGIR 2021 |
Medical and health informatics › clinical prediction
readmission prediction |
0.5 | 1 | 2021 | Predicting Patient Readmission Risk from Medical Text via Knowledge Graph Enhanced Multiview Graph Convolution · SIGIR 2021 |
Information retrieval
knowledge-enhanced representation learning |
0.5 | 1 | 2021 | Predicting Patient Readmission Risk from Medical Text via Knowledge Graph Enhanced Multiview Graph Convolution · SIGIR 2021 |
Methods — techniques the papers use, named apart from their topics
knowledge graph · 1.0graph convolutional network · 1.0neuro-symbolic reasoning · 0.5logic neural network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Clinical document metadata extraction: A scoping reviewabstractOBJECTIVES: Clinical document metadata, such as document type, structure, author role, medical specialty, and encounter setting, is essential for accurate interpretation of information captured in clinical documents. However, vast documentation heterogeneity and drift over time challenge harmonization of document metadata. Automated extraction methods have emerged to coalesce metadata from disparate practices into target schema. This scoping review aims to catalog research on clinical document metadata extraction, identify methodological trends and applications, and highlight gaps warranting further investigation. METHODS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines to identify articles from Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science and external sources that perform clinical document metadata extraction, either primarily as a methodology study, secondarily as a feature for a downstream application, or for analysis. We initially identified and screened 342 articles published between 2011 and 2025, then comprehensively reviewed 77 we deemed relevant to our study. RESULTS: Among the 77 articles included in our full text review, 49 were methodological, 22 used document metadata as features in a downstream application, and 6 analyzed document metadata composition. We observe myriad purposes for methodological study and application types. Available labelled public data remains sparse except for structural section datasets. Methods for extracting document metadata have progressed from largely rule-based and traditional machine learning with ample feature engineering to transformer-based architectures with minimal feature engineering. DISCUSSION AND CONCLUSION: Clinical document metadata extraction research has accelerated over recent years. The emergence of large language models has enabled broader exploration of generalizability across tasks and datasets, allowing the possibility of advanced clinical text processing systems. We anticipate that research will continue to expand into richer document metadata representations and integrate further into clinical applications and workflows. Kurt Miller, Qiuhao Lu, William R. Hersh, Kirk Roberts, Steven Bedrick, Andrew Wen |
J. Biomed. Informatics | 2 |
| 2025 | Dynamic few-shot prompting for clinical note section classification using lightweight, open-source large language modelsabstractOBJECTIVE: Unlocking clinical information embedded in clinical notes has been hindered to a significant degree by domain-specific and context-sensitive language. Identification of note sections and structural document elements has been shown to improve information extraction and dependent downstream clinical natural language processing (NLP) tasks and applications. This study investigates the viability of a dynamic example selection prompting method to section classification using lightweight, open-source large language models (LLMs) as a practical solution for real-world healthcare clinical NLP systems. MATERIALS AND METHODS: We develop a dynamic few-shot prompting approach to classifying sections where section samples are first embedded using a transformer-based model and deposited in a vector store. During inference, the embedded samples with the most similar contextual embeddings to a given input section text are retrieved from the vector store and inserted into the LLM prompt. We evaluate this technique on two datasets comprising two section schemas, including varying levels of context. We compare the performance to baseline zero-shot and randomly selected few-shot scenarios. RESULTS: The dynamic few-shot prompting experiments yielded the highest F1 scores in each of the classification tasks and datasets for all seven of the LLMs included in the evaluation, averaging a macro F1 increase of 39.3% and 21.1% in our primary section classification task over the zero-shot and static few-shot baselines, respectively. DISCUSSION AND CONCLUSION: Our results showcase substantial performance improvements imparted by dynamically selecting examples for few-shot LLM prompting, and further improvement by including section context, demonstrating compelling potential for clinical applications. Kurt Miller, Steven Bedrick, Qiuhao Lu, Andrew Wen, William R. Hersh, Kirk Roberts |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Discovering signature disease trajectories in pancreatic cancer and soft-tissue sarcoma from longitudinal patient records
Liwei Wang 0010, Andrew Wen, Qiuhao Lu, Jinlian Wang, Xiaoyang Ruan, Adriana Gamboa, Neha Malik, Christina L. Roland, Matthew H. G. Katz, Heather Lyu |
J. Biomed. Informatics | 4 |
| 2025 | A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and TrustworthinessabstractLarge language models (LLMs) have demonstrated emergent abilities in text generation, question answering, and reasoning, facilitating various tasks and domains. Despite their proficiency in various tasks, LLMs like PaLM 540B and Llama-3.1 405B face limitations due to large parameter sizes and computational demands, often requiring cloud API use, which raises privacy concerns, limits real-time applications on edge devices, and increases fine-tuning costs. Additionally, LLMs often underperform in specialized domains such as healthcare and law due to insufficient domain-specific knowledge, necessitating specialized models. Therefore, Small Language Models (SLMs) are increasingly favored for their low inference latency, cost-effectiveness, efficient development, and easy customization and adaptability. These models are particularly well-suited for resource-limited environments and domain knowledge acquisition, addressing LLMs’ challenges and proving ideal for applications that require localized data handling for privacy, minimal inference latency for efficiency, and domain knowledge acquisition through lightweight fine-tuning. The rising demand for SLMs has spurred extensive research and development. However, a comprehensive survey investigating issues related to the definition, acquisition, application, enhancement, and reliability of SLM remains lacking, prompting us to conduct a detailed survey on these topics. The definition of SLMs varies widely; thus, to standardize, we propose defining SLMs by their capability to perform specialized tasks and suitability for resource-constrained settings, setting boundaries based on the minimal size for emergent abilities and the maximum size sustainable under resource constraints. For other aspects, we provide a taxonomy of relevant models/methods and develop general frameworks for each category to enhance and utilize SLMs effectively. We have compiled the collected SLM models and related methods on GitHub: https://github.com/FairyFali/SLMs-Survey . Fali Wang, Zhiwei Zhang 0028, Xianren Zhang, Zongyu Wu 0001, Tzuhao Mo, Qiuhao Lu, Wanjing Wang, Xianfeng Tang, Qi He 0002, Yao Ma 0001, Ming Huang 0006, Suhang Wang |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2021 | LNN-EL: A Neuro-Symbolic Approach to Short-text Entity LinkingabstractHang Jiang, Sairam Gurajada, Qiuhao Lu, Sumit Neelam, Lucian Popa, Prithviraj Sen, Yunyao Li, Alexander Gray. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sairam Gurajada, Qiuhao Lu, Sumit Neelam, Lucian Popa 0001, Prithviraj Sen, Yunyao Li 0001, Alexander G. Gray |
ACL/IJCNLP (1) | 3 |
| 2021 | Textual Data Augmentation for Patient Outcomes PredictionabstractDeep learning models have demonstrated superior performance in various healthcare applications. However, the major limitation of these deep models is usually the lack of high-quality training data due to the private and sensitive nature of this field. In this study, we propose a novel textual data augmentation method to generate artificial clinical notes in patients’ Electronic Health Records (EHRs) that can be used as additional training data for patient outcomes prediction. Essentially, we fine-tune the generative language model GPT-2 to synthesize labeled text with the original training data. More specifically, We propose a teacher-student framework where we first pre-train a teacher model on the 0riginal data, and then train a student model on the GPT-augmented data under the guidance of the teacher. We evaluate our method on the most common patient outcome, i.e., the 30-day readmission rate. The experimental results show that deep models can improve their predictive performance with the augmented data, indicating the effectiveness of the proposed architecture. Qiuhao Lu, Dejing Dou, Thien Huu Nguyen |
BIBM | 1 |
| 2021 | Predicting Patient Readmission Risk from Medical Text via Knowledge Graph Enhanced Multiview Graph ConvolutionabstractUnplanned intensive care unit (ICU) readmission rate is an important metric for evaluating the quality of hospital care. Efficient and accurate prediction of ICU readmission risk can not only help prevent patients from inappropriate discharge and potential dangers, but also reduce associated costs of healthcare. In this paper, we propose a new method that uses medical text of Electronic Health Records (EHRs) for prediction, which provides an alternative perspective to previous studies that heavily depend on numerical and time-series features of patients. More specifically, we extract discharge summaries of patients from their EHRs, and represent them with multiview graphs enhanced by an external knowledge graph. Graph convolutional networks are then used for representation learning. Experimental results prove the effectiveness of our method, yielding state-of-the-art performance for this task. Qiuhao Lu, Thien Huu Nguyen, Dejing Dou |
SIGIR | 1 |
| 2020 | Exploiting Node Content for Multiview Graph Convolutional Network and Adversarial RegularizationabstractNetwork representation learning (NRL) is crucial in the area of graph learning.Recently, graph autoencoders and its variants have gained much attention and popularity among various types of node embedding approaches.Most existing graph autoencoder-based methods aim to minimize the reconstruction errors of the input network while not explicitly considering the semantic relatedness between nodes.In this paper, we propose a novel network embedding method which models the consistency across different views of networks.More specifically, we create a second view from the input network which captures the relation between nodes based on node content and enforce the latent representations from the two views to be consistent by incorporating a multiview adversarial regularization module.The experimental studies on benchmark datasets prove the effectiveness of this method, and demonstrate that our method compares favorably with the state-of-the-art algorithms on challenging tasks such as link prediction and node clustering.We also evaluate our method on a real-world application, i.e., 30-day unplanned ICU readmission prediction, and achieve promising results compared with several baseline methods. Qiuhao Lu, Nisansa de Silva, Dejing Dou, Thien Huu Nguyen, Prithviraj Sen, Berthold Reinwald, Yunyao Li 0001 |
COLING | 1 |
| 2017 | Wikipedia-Based Entity Semantifying in Open Information ExtractionabstractIn the recent years, Open Information Extraction (OIE), an unsupervised strategy which extracts open-domain facts of knowledge from massive heterogeneous text corpora, has achieved impressive improvements. However, the facts (generally represented by a triple) extracted by OIE systems are in lack of clear semantics and then difficult for computer systems to understand. In this paper, we present a new method to semantify the facts by mapping the string arguments in the triples to the corresponding real-world entities based on the existing knowledge base Wikipedia. First, for each query of string argument, we consider a set of its most likely mapping entities and assign each candidate a fused prior probability. Then we calculate the graph-based similarity between candidates as the contextual evidence by propagating semantics on the neighborhood graph of candidates. Finally, we transform the mapping task into an optimization problem and find the maximum a posteriori (MAP) mapping by combining the prior information and contextual evidence through Bayes' theorem. Due to the fusion of multiple cues and the semantics propagation over the graph, our approach improves the performance of the entity semantifying. Experimental results demonstrate the effectiveness of our approach. Qiuhao Lu, Youtian Du |
ICDAR | 1 |