VLDB 2026 Research / reviewers in the wild / expert
Su-Yuan Peng
dblp:157/0852 · also Suyuan Peng
· DBLP profile ↗
9ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-8221-7574ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AcuKG: a comprehensive knowledge graph for medical acupunctureabstractBACKGROUND: Acupuncture, a key modality in traditional Chinese medicine, is gaining global recognition as a complementary therapy and a subject of increasing scientific interest. However, fragmented and unstructured acupuncture knowledge spread across diverse sources poses challenges for semantic retrieval, reasoning, and in-depth analysis. To address this gap, we developed AcuKG, a comprehensive knowledge graph that systematically organizes acupuncture-related knowledge to support sharing, discovery, and artificial intelligence-driven innovation in the field. METHODS: AcuKG integrates data from multiple sources, including online resources, guidelines, PubMed literature, ClinicalTrials.gov, and multiple ontologies (SNOMED CT, UBERON, and MeSH). We employed entity recognition, relation extraction, and ontology mapping to establish AcuKG, with human-in-the-loop to ensure data quality. Two cases evaluated AcuKG's usability: (1) how AcuKG advances acupuncture research for obesity and (2) how AcuKG enhances large language model (LLM) application on acupuncture question-answering. RESULTS: AcuKG comprises 1839 entities and 11 527 relations, mapped to 1836 standard concepts in 3 ontologies. Two use cases demonstrated AcuKG's effectiveness and potential in advancing acupuncture research and supporting LLM applications. In the obesity use case, AcuKG identified highly relevant acupoints (eg, ST25, ST36) and uncovered novel research insights based on evidence from clinical trials and literature. When applied to LLMs in answering acupuncture-related questions, integrating AcuKG with GPT-4o and LLaMA 3 significantly improved accuracy (GPT-4o: 46% → 54%, P = .03; LLaMA 3: 17% → 28%, P = .01). CONCLUSION: AcuKG is an open dataset that provides a structured and computational framework for acupuncture applications, bridging traditional practices with acupuncture research and cutting-edge LLM technologies. Xueqing Peng, Su-Yuan Peng, Jianfu Li, Donghong Pei, Fang Li 0011, Yongqun He, Cui Tao, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Construction of a Clinical Data Standardization System Based on Medical Data Element and Large Language ModelabstractTo address challenges such as diverse clinical data sources, inconsistent standards, and difficulties in sharing, this study developed a clinical data standardization system based on medical data element and Large Language Model (LLM). 1. Customizable patient and medical record templates were created using medical data element and related terminology, enabling dynamic binding for data structuring and standardization. 2. LLM was leveraged to perform intelligent information extraction, semantic normalization, and standardization of unstructured or semi-structured clinical text. 3. User permissions were configured to manage medical record data across both record repositories and institutions, supporting cross-repository and cross-institutional export of standardized data. This study offers an alternative solution for integrating multi-source, heterogeneous clinical data while providing high-quality AI-ready data for applications such as data mining and LLM training. This advancement supports research in clinical science, medical artificial intelligence, and related fields. Yuzhu Li, Su-Yuan Peng, Yan Zhu 0021, Lihong Liu, Keyu Yao |
BIBM | 3 |
| 2025 | Applied LLM for the Construction of Traditional Chinese Medicine Tongue Diagnosis OntologyabstractOntology is essential for representing semantics and knowledge. This study compares three Large Language Models (LLM)-assisted ontology construction methods- no prompting, chain - of - thought prompting accompanied by a guiding table, and the combination of ISO standards and chain - of - thought prompting accompanied by a guiding table. And applies them to tongue diagnosis ontology in Traditional Chinese Medicine. The construction process systematically leverages LLMs' generative and analytical capabilities, following a structured workflow of domain definition, terminology collection, class hierarchy construction, and manual verification, with continuous expert involvement to refine definitions and validate outputs. Evaluation based on ISO standards entity alignment and semantic similarity shows that the combination of ISO standards and chain - of - thought prompting accompanied by a guiding table method performs best. The resulting ontology also demonstrate logical consistency, structural soundness, semantic accuracy, and inferential completeness. Keyu Yao, Peixin Ge, Yan Zhu 0021, Su-Yuan Peng |
BIBM | 7 |
| 2024 | Relation extraction using large language models: a case study on acupuncture point locationsabstractOBJECTIVE: In acupuncture therapy, the accurate location of acupoints is essential for its effectiveness. The advanced language understanding capabilities of large language models (LLMs) like Generative Pre-trained Transformers (GPTs) and Llama present a significant opportunity for extracting relations related to acupoint locations from textual knowledge sources. This study aims to explore the performance of LLMs in extracting acupoint-related location relations and assess the impact of fine-tuning on GPT's performance. MATERIALS AND METHODS: We utilized the World Health Organization Standard Acupuncture Point Locations in the Western Pacific Region (WHO Standard) as our corpus, which consists of descriptions of 361 acupoints. Five types of relations ("direction_of", "distance_of", "part_of", "near_acupoint", and "located_near") (n = 3174) between acupoints were annotated. Four models were compared: pre-trained GPT-3.5, fine-tuned GPT-3.5, pre-trained GPT-4, as well as pretrained Llama 3. Performance metrics included micro-average exact match precision, recall, and F1 scores. RESULTS: Our results demonstrate that fine-tuned GPT-3.5 consistently outperformed other models in F1 scores across all relation types. Overall, it achieved the highest micro-average F1 score of 0.92. DISCUSSION: The superior performance of the fine-tuned GPT-3.5 model, as shown by its F1 scores, underscores the importance of domain-specific fine-tuning in enhancing relation extraction capabilities for acupuncture-related tasks. In light of the findings from this study, it offers valuable insights into leveraging LLMs for developing clinical decision support and creating educational modules in acupuncture. CONCLUSION: This study underscores the effectiveness of LLMs like GPT and Llama in extracting relations related to acupoint locations, with implications for accurately modeling acupuncture knowledge and promoting standard implementation in acupuncture training and practice. The findings also contribute to advancing informatics applications in traditional and complementary medicine, showcasing the potential of LLMs in natural language processing. Xueqing Peng, Jianfu Li, Xu Zuo, Su-Yuan Peng, Donghong Pei, Cui Tao, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Development of the International Classification of Diseases Ontology (ICDO) and its application for COVID-19 diagnostic data analysisabstractBACKGROUND: The 10th and 9th revisions of the International Statistical Classification of Diseases and Related Health Problems (ICD10 and ICD9) have been adopted worldwide as a well-recognized norm to share codes for diseases, signs and symptoms, abnormal findings, etc. The international Consortium for Clinical Characterization of COVID-19 by EHR (4CE) website stores diagnosis COVID-19 disease data using ICD10 and ICD9 codes. However, the ICD systems are difficult to decode due to their many shortcomings, which can be addressed using ontology. METHODS: An ICD ontology (ICDO) was developed to logically and scientifically represent ICD terms and their relations among different ICD terms. ICDO is also aligned with the Basic Formal Ontology (BFO) and reuses terms from existing ontologies. As a use case, the ICD10 and ICD9 diagnosis data from the 4CE website were extracted, mapped to ICDO, and analyzed using ICDO. RESULTS: We have developed the ICDO to ontologize the ICD terms and relations. Different from existing disease ontologies, all ICD diseases in ICDO are defined as disease processes to describe their occurrence with other properties. The ICDO decomposes each disease term into different components, including anatomic entities, process profiles, etiological causes, output phenotype, etc. Over 900 ICD terms have been represented in ICDO. Many ICDO terms are presented in both English and Chinese. The ICD10/ICD9-based diagnosis data of over 27,000 COVID-19 patients from 5 countries were extracted from the 4CE. A total of 917 COVID-19-related disease codes, each of which were associated with 1 or more cases in the 4CE dataset, were mapped to ICDO and further analyzed using the ICDO logical annotations. Our study showed that COVID-19 targeted multiple systems and organs such as the lung, heart, and kidney. Different acute and chronic kidney phenotypes were identified. Some kidney diseases appeared to result from other diseases, such as diabetes. Some of the findings could only be easily found using ICDO instead of ICD9/10. CONCLUSIONS: ICDO was developed to ontologize ICD10/10 codes and applied to study COVID-19 patient diagnosis data. Our findings showed that ICDO provides a semantic platform for more accurate detection of disease profiles. Ling Wan, Justin Song, Virginia He, Jennifer Roman, Grace Whah, Su-Yuan Peng, Luxia Zhang, Yongqun He |
BMC Bioinform. | 6 |
| 2019 | HPO2Vec+: Leveraging heterogeneous knowledge resources to enrich node embeddings for the Human Phenotype Ontology
Feichen Shen, Su-Yuan Peng, Yadan Fan, Andrew Wen, Sijia Liu 0002, Yanshan Wang, Liwei Wang 0010 |
J. Biomed. Informatics | 2 |
| 2018 | Leveraging Association Rule Mining to Detect Pathophysiological Mechanisms of Chronic Kidney Disease Complicated by Metabolic Syndrome
Su-Yuan Peng, Yadan Fan, Liwei Wang 0010, Andrew Wen, Xu-Sheng Liu, Feichen Shen |
BIBM | 1 |
| 2018 | Chronic Kidney Disease Medication Adherence and its Influencing Factors: An Observation and Analysis
Dingjun Zhang, Su-Yuan Peng, Lizhe Fu, Jiaowang Tan, Meiqin Ye, Xu-Sheng Liu |
BIBM | 3 |
| 2014 | TCMISS-based analysis on the regularity of ancient TCM prescriptions for chronic renal failureabstractObjective: To explore composing principles of prescriptions for chronic renal failure recorded on the Encyclopedia of Traditional Chinese Medicines based on the traditional Chinese medicine inheritance support system (TCMISS). Methods: Prescriptions for chronic renal failure recorded on the Encyclopedia of Traditional Chinese Medicines were entered into the TCMISS. Composing principles of these prescriptions were analyzed by TCMISS-integrated unsupervised data mining algorithms (including Improved Mutual Information, Complex Systems Entropy Clustering, and Unsupervised Entropy Hierarchical Clustering). Results: Eighteen core herbal combinations and 9 new prescriptions were evolved and obtained by analyzing 221 selected prescriptions for chronic renal failure and determining frequencies of occurrence of Chinese medicinal herbs in these prescriptions. Su-Yuan Peng, Xu-Sheng Liu |
BIBM | 1 |