EDBT 2026 Demo / reviewers in the wild / expert
Jianfu Li
dblp:01/7732
· DBLP profile ↗
25ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AcuKG: a comprehensive knowledge graph for medical acupunctureabstractBACKGROUND: Acupuncture, a key modality in traditional Chinese medicine, is gaining global recognition as a complementary therapy and a subject of increasing scientific interest. However, fragmented and unstructured acupuncture knowledge spread across diverse sources poses challenges for semantic retrieval, reasoning, and in-depth analysis. To address this gap, we developed AcuKG, a comprehensive knowledge graph that systematically organizes acupuncture-related knowledge to support sharing, discovery, and artificial intelligence-driven innovation in the field. METHODS: AcuKG integrates data from multiple sources, including online resources, guidelines, PubMed literature, ClinicalTrials.gov, and multiple ontologies (SNOMED CT, UBERON, and MeSH). We employed entity recognition, relation extraction, and ontology mapping to establish AcuKG, with human-in-the-loop to ensure data quality. Two cases evaluated AcuKG's usability: (1) how AcuKG advances acupuncture research for obesity and (2) how AcuKG enhances large language model (LLM) application on acupuncture question-answering. RESULTS: AcuKG comprises 1839 entities and 11 527 relations, mapped to 1836 standard concepts in 3 ontologies. Two use cases demonstrated AcuKG's effectiveness and potential in advancing acupuncture research and supporting LLM applications. In the obesity use case, AcuKG identified highly relevant acupoints (eg, ST25, ST36) and uncovered novel research insights based on evidence from clinical trials and literature. When applied to LLMs in answering acupuncture-related questions, integrating AcuKG with GPT-4o and LLaMA 3 significantly improved accuracy (GPT-4o: 46% → 54%, P = .03; LLaMA 3: 17% → 28%, P = .01). CONCLUSION: AcuKG is an open dataset that provides a structured and computational framework for acupuncture applications, bridging traditional practices with acupuncture research and cutting-edge LLM technologies. Xueqing Peng, Su-Yuan Peng, Jianfu Li, Donghong Pei, Fang Li 0011, Yongqun He, Cui Tao, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 4 |
| 2026 | Exploring the role of reinforcement learning in vision-language models for cardiovascular disease decision support
Pengze Li, Jianfu Li, Shuteng Niu, Farris K. Timimi, Joseph Cheung, Clark Otley, Sonya Makhni, Fang Li 0011, Jingna Feng, Xinyue Hu 0002, Yue Yu 0012, Cui Tao |
J. Biomed. Informatics | 2 |
| 2025 | ParaCUP-NBR: Next-basket Recommendation via Parallel Modeling of Compatibility and User Preference with Multi-attribute Items
Jianfu Li, Yuying Yang |
IEEE Big Data | 1 |
| 2025 | SCMF: Long-Term Time Series Forecasting Model Based on Series Collapsing and Multiple Periods FusionabstractLong-term time series forecasting(LTSF) aims to predict future series based on historical series. Due to the overfitting problem that troubles complex models, lightweight models have become the mainstream trend for LTSF. The recently proposed SparseTSF has excellent performance with parameters below 1K, bringing hope to the research of LTSF. However, SparseTSF has two limitations: (1) The model structure limits its ability to capture complex temporal variations. (2) It can not capture the multiple periodicity of time series, whereas in practice long time series are often interspersed with multiple periods. In order to improve SparseTSF, this paper proposes a LTSF model named as SCMF based on series collapsing and multiple periods fusion. On the basis of SparseTSF, SCMF optimizes the order of learning the intra-period and inter-period relationships, and fuses the dependencies between multiple periods by collapsing series and cross attention. To validate the effectiveness of SCMF, we conducted experiments on seven real-world datasets and the results demonstrate that SCMF is more accurate compared with the state-of-the-art models. Jianfu Li, Chonghao Liu |
CSCWD | 1 |
| 2025 | Enhancing Long-Tailed Recognition with Skill-specialized Experts and Bootstrap Latent ConsistencyabstractLong-tailed recognition tasks face the critical challenge of extreme data imbalance, where overrepresented head categories dominate, while tail categories suffer from severe underrepresentation. This imbalance hinders the model’s ability to generalize to tail categories, a key bottleneck in real-world applications. Existing approaches, such as resampling strategies or loss function adjustments, often fail to balance the trade-off between head and tail categories. To tackle these challenges, we propose SELA-Net (Skill-specialized Experts with Bootstrap Latent Adaptation Network), a novel framework that synergizes a multi-expert design with adaptive layer-wise feature alignment to enhance representation learning for underrepresented categories. SELA-Net integrates two core modules: Skill-specialized Experts (SE) and Bootstrap Latent Consistency (BLC). The SE module leverages logit adjustment and Mixup techniques to optimize category-specific feature learning, achieving fine-grained representations for head, medium, and tail categories. Meanwhile, the BLC module enforces adaptive layer-wise consistency across feature representations, effectively enhancing the model’s generalization, particularly for tail categories. Experimental results on four widely used long-tailed datasets, including CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist2018, demonstrate that SELA-Net consistently outperforms state-of-the-art methods, achieving significant improvements in tail category recognition. Notably, SELA-Net excels in diverse test distributions and ablation studies, validating its robustness and adaptability in addressing long-tailed challenges. The code for this work is publicly available at: https://github.com/MuZiYuHui/SELA-Net. Jianfu Li |
IJCNN | 1 |
| 2025 | Exploring multimodal large language models on transthoracic Echocardiogram (TTE) tasks for cardiovascular decision support
Jianfu Li, Zenan Sun, Evan Yu, Ahmed M. Abdelhameed, Weiguo Cao, Jianping He 0002, Pengze Li, Jingna Feng, Yue Yu 0012, Xinyue Hu 0002, Manqi Li, Yifang Dang, Fang Li 0011, Shahyar M. Gharacholou, Cui Tao |
J. Biomed. Informatics | 1 |
| 2025 | Improving entity recognition using ensembles of deep learning and fine-tuned large language models: A case study on adverse event extraction from VAERS and social media
Deepthi Viswaroopan, William He, Jianfu Li, Xu Zuo, Hua Xu 0001, Cui Tao |
J. Biomed. Informatics | 4 |
| 2024 | Relation extraction using large language models: a case study on acupuncture point locationsabstractOBJECTIVE: In acupuncture therapy, the accurate location of acupoints is essential for its effectiveness. The advanced language understanding capabilities of large language models (LLMs) like Generative Pre-trained Transformers (GPTs) and Llama present a significant opportunity for extracting relations related to acupoint locations from textual knowledge sources. This study aims to explore the performance of LLMs in extracting acupoint-related location relations and assess the impact of fine-tuning on GPT's performance. MATERIALS AND METHODS: We utilized the World Health Organization Standard Acupuncture Point Locations in the Western Pacific Region (WHO Standard) as our corpus, which consists of descriptions of 361 acupoints. Five types of relations ("direction_of", "distance_of", "part_of", "near_acupoint", and "located_near") (n = 3174) between acupoints were annotated. Four models were compared: pre-trained GPT-3.5, fine-tuned GPT-3.5, pre-trained GPT-4, as well as pretrained Llama 3. Performance metrics included micro-average exact match precision, recall, and F1 scores. RESULTS: Our results demonstrate that fine-tuned GPT-3.5 consistently outperformed other models in F1 scores across all relation types. Overall, it achieved the highest micro-average F1 score of 0.92. DISCUSSION: The superior performance of the fine-tuned GPT-3.5 model, as shown by its F1 scores, underscores the importance of domain-specific fine-tuning in enhancing relation extraction capabilities for acupuncture-related tasks. In light of the findings from this study, it offers valuable insights into leveraging LLMs for developing clinical decision support and creating educational modules in acupuncture. CONCLUSION: This study underscores the effectiveness of LLMs like GPT and Llama in extracting relations related to acupoint locations, with implications for accurately modeling acupuncture knowledge and promoting standard implementation in acupuncture training and practice. The findings also contribute to advancing informatics applications in traditional and complementary medicine, showcasing the potential of LLMs in natural language processing. Xueqing Peng, Jianfu Li, Xu Zuo, Su-Yuan Peng, Donghong Pei, Cui Tao, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Ensemble pretrained language models to extract biomedical knowledge from literatureabstractOBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking. MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and taxonomy names and annotated 70 more abstracts as additional corpus. We further finetuned the T0pp model, with 11 billion parameters, to boost the performance on relation extraction and leveraged entites' location information (eg, title, background) to enhance novelty prediction performance in relation extraction (RE). RESULTS: Our pioneering NLP system designed for this challenge secured first place in Phase I-NER and second place in Phase II-relation extraction and novelty prediction, outpacing over 200 teams. We tested OpenAI ChatGPT 3.5 and ChatGPT 4 in a Zero-Shot setting using the same test set, revealing that our finetuned model considerably surpasses these broad-spectrum large language models. DISCUSSION AND CONCLUSION: Our outcomes depict a robust NLP system excelling in NER and RE across various biomedical entities, emphasizing that task-specific models remain superior to generic large ones. Such insights are valuable for endeavors like knowledge graph development and hypothesis formulation in biomedical research. Qiang Wei 0002, Liang-Chin Huang, Jianfu Li, Yao-Shun Chuang, Jianping He 0002, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S. Diala, Kirk Roberts, Cui Tao, Xiaoqian Jiang, W. Jim Zheng, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | RefAI: a GPT-powered retrieval-augmented generative tool for biomedical literature recommendation and summarizationabstractOBJECTIVES: Precise literature recommendation and summarization are crucial for biomedical professionals. While the latest iteration of generative pretrained transformer (GPT) incorporates 2 distinct modes-real-time search and pretrained model utilization-it encounters challenges in dealing with these tasks. Specifically, the real-time search can pinpoint some relevant articles but occasionally provides fabricated papers, whereas the pretrained model excels in generating well-structured summaries but struggles to cite specific sources. In response, this study introduces RefAI, an innovative retrieval-augmented generative tool designed to synergize the strengths of large language models (LLMs) while overcoming their limitations. MATERIALS AND METHODS: RefAI utilized PubMed for systematic literature retrieval, employed a novel multivariable algorithm for article recommendation, and leveraged GPT-4 turbo for summarization. Ten queries under 2 prevalent topics ("cancer immunotherapy and target therapy" and "LLMs in medicine") were chosen as use cases and 3 established counterparts (ChatGPT-4, ScholarAI, and Gemini) as our baselines. The evaluation was conducted by 10 domain experts through standard statistical analyses for performance comparison. RESULTS: The overall performance of RefAI surpassed that of the baselines across 5 evaluated dimensions-relevance and quality for literature recommendation, accuracy, comprehensiveness, and reference integration for summarization, with the majority exhibiting statistically significant improvements (P-values <.05). DISCUSSION: RefAI demonstrated substantial improvements in literature recommendation and summarization over existing tools, addressing issues like fabricated papers, metadata inaccuracies, restricted recommendations, and poor reference integration. CONCLUSION: By augmenting LLM with external resources and a novel ranking algorithm, RefAI is uniquely capable of recommending high-quality literature and generating well-structured summaries, holding the potential to meet the critical needs of biomedical professionals in navigating and synthesizing vast amounts of scientific literature. Jeff Zhao, Manqi Li, Yifang Dang, Evan Yu, Jianfu Li, Zenan Sun, Usama Hussein, Jianguo Wen, Ahmed M. Abdelhameed, Junhua Mai, Shenduo Li, Yue Yu 0012, Xinyue Hu 0002, Daowei Yang, Jingna Feng, Zehan Li, Jianping He 0002, Tiehang Duan, Yanyan Lou, Fang Li 0011, Cui Tao |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Improving tabular data extraction in scanned laboratory reports using deep learning models
Qiang Wei 0002, Xinghan Chen, Jianfu Li, Cui Tao, Hua Xu 0001 |
J. Biomed. Informatics | 4 |
| 2023 | POI Recommendation Based on Double-Level Spatio-Temporal Relationship in Locations and Categories
Jianfu Li |
ICONIP (13) | 1 |
| 2022 | Extracting Cancer Chemotherapy and Response Information from Clinical Notes following the RECIST Definition
Xu Zuo, Natalie Gregoriou, Jianfu Li, Jeremy Warner, Yang Ping |
AMIA | 4 |
| 2021 | Are synthetic clinical notes useful for real natural language processing tasks: A case study on clinical entity recognitionabstractOBJECTIVE: : Developing clinical natural language processing systems often requires access to many clinical documents, which are not widely available to the public due to privacy and security concerns. To address this challenge, we propose to develop methods to generate synthetic clinical notes and evaluate their utility in real clinical natural language processing tasks. MATERIALS AND METHODS: : We implemented 4 state-of-the-art text generation models, namely CharRNN, SegGAN, GPT-2, and CTRL, to generate clinical text for the History and Present Illness section. We then manually annotated clinical entities for randomly selected 500 History and Present Illness notes generated from the best-performing algorithm. To compare the utility of natural and synthetic corpora, we trained named entity recognition (NER) models from all 3 corpora and evaluated their performance on 2 independent natural corpora. RESULTS: : Our evaluation shows GPT-2 achieved the best BLEU (bilingual evaluation understudy) score (with a BLEU-2 of 0.92). NER models trained on synthetic corpus generated by GPT-2 showed slightly better performance on 2 independent corpora: strict F1 scores of 0.709 and 0.748, respectively, when compared with the NER models trained on natural corpus (F1 scores of 0.706 and 0.737, respectively), indicating the good utility of synthetic corpora in clinical NER model development. In addition, we also demonstrated that an augmented method that combines both natural and synthetic corpora achieved better performance than that uses the natural corpus only. CONCLUSIONS: : Recent advances in text generation have made it possible to generate synthetic clinical notes that could be useful for training NER models for information extraction from natural clinical notes, thus lowering the privacy concern and increasing data availability. Further investigation is needed to apply this technology to practice. Jianfu Li, Yujia Zhou 0003, Xiaoqian Jiang, Karthik Natarajan, Serguei V. S. Pakhomov, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Deep Learning Approach to Parse Eligibility Criteria in Dietary Supplements Clinical Trials Following OMOP Common Data Model
Anusha Bompelli, Jianfu Li, Yiqi Xu, Yanshan Wang, Terrence Adam, Zhe He 0001, Rui Zhang 0028 |
AMIA | 2 |
| 2020 | Normalizing Clinical Document Titles to LOINC Document Ontology: an Initial Study
Xu Zuo, Jianfu Li, Bo Zhao 0001, Yujia Zhou 0003, Jon D. Duke, Karthik Natarajan, George Hripcsak, Nigam H. Shah, Juan M. Banda, Ruth M. Reeves, Hua Xu 0001 |
AMIA | 2 |
| 2020 | Dynamic Temporospatial Patterns of Functional Connectivity and Alterations in Idiopathic Generalized EpilepsyabstractThe dynamic profile of brain function has received much attention in recent years and is also a focus in the study of epilepsy. The present study aims to integrate the dynamics of temporal and spatial characteristics to provide comprehensive and novel understanding of epileptic dynamics. Resting state fMRI data were collected from eighty-three patients with idiopathic generalized epilepsy (IGE) and 87 healthy controls (HC). Specifically, we explored the temporal and spatial variation of functional connectivity density (tvFCD and svFCD) in the whole brain. Using a sliding-window approach, for a given region, the standard variation of the FCD series was calculated as the tvFCD and the variation of voxel-wise spatial distribution was calculated as the svFCD. We found primary, high-level, and sub-cortical networks demonstrated distinct tvFCD and svFCD patterns in HC. In general, the high-level networks showed the highest variation, the subcortical and primary networks showed moderate variation, and the limbic system showed the lowest variation. Relative to HC, the patients with IGE showed weaken temporal and enhanced spatial variation in the default mode network and weaken temporospatial variation in the subcortical network. Besides, enhanced temporospatial variation in sensorimotor and high-level networks was also observed in patients. The hyper-synchronization of specific brain networks was inferred to be associated with the phenomenon responsible for the intrinsic propensity of generation and propagation of epileptic activities. The disrupted dynamic characteristics of sensorimotor and high-level networks might potentially contribute to the driven motion and cognition phenotypes in patients. In all, presently provided evidence from the temporospatial variation of functional interaction shed light on the dynamics underlying neuropathological profiles of epilepsy. Sisi Jiang, Haonan Pei, Linli Liu, Jianfu Li, Dezhong Yao 0001 |
Int. J. Neural Syst. | 6 |
| 2020 | Rhythmic Network Modulation to Thalamocortical Couplings in EpilepsyabstractThalamus interacts with cortical areas, generating oscillations characterized by their rhythm and levels of synchrony. However, little is known of what function the rhythmic dynamic may serve in thalamocortical couplings. This work introduced a general approach to investigate the modulatory contribution of rhythmic scalp network to the thalamo-frontal couplings in juvenile myoclonic epilepsy (JME) and frontal lobe epilepsy (FLE). Here, time-varying rhythmic network was constructed using the adapted directed transfer function between EEG electrodes, and then was applied as a modulator in fMRI-based thalamocortical functional couplings. Furthermore, the relationship between corticocortical connectivity and rhythm-dependent thalamocortical coupling was examined. The results revealed thalamocortical couplings modulated by EEG scalp network have frequency-dependent characteristics. Increased thalamus- sensorimotor network (SMN) and thalamus-default mode network (DMN) couplings in JME were strongly modulated by alpha band. These thalamus-SMN couplings demonstrated enhanced association with SMN-related corticocortical connectivity. In addition, altered theta-dependent and beta-dependent thalamus-frontoparietal network (FPN) couplings were found in FLE. The reduced theta-dependent thalamus-FPN couplings were associated with the decreased FPN-related corticocortical connectivity. This study proposed interactive links between the rhythmic modulation and thalamocortical coupling. The crucial role of SMN and FPN in subcortical-cortical circuit may have implications for intervention in generalized and focal epilepsy. Yun Qin, Xiaojun Zuo, Sisi Jiang, Xiaole Zhao, Li Dong 0003, Jianfu Li, Tao Zhang 0017, Dezhong Yao 0001 |
Int. J. Neural Syst. | 8 |
| 2020 | COVID-19 TestNorm: A tool to normalize COVID-19 testing names to LOINC codesabstractLarge observational data networks that leverage routine clinical practice data in electronic health records (EHRs) are critical resources for research on coronavirus disease 2019 (COVID-19). Data normalization is a key challenge for the secondary use of EHRs for COVID-19 research across institutions. In this study, we addressed the challenge of automating the normalization of COVID-19 diagnostic tests, which are critical data elements, but for which controlled terminology terms were published after clinical implementation. We developed a simple but effective rule-based tool called COVID-19 TestNorm to automatically normalize local COVID-19 testing names to standard LOINC (Logical Observation Identifiers Names and Codes) codes. COVID-19 TestNorm was developed and evaluated using 568 test names collected from 8 healthcare systems. Our results show that it could achieve an accuracy of 97.4% on an independent test set. COVID-19 TestNorm is available as an open-source package for developers and as an online Web application for end users (https://clamp.uth.edu/covid/loinc.php). We believe that it will be a useful tool to support secondary use of EHRs for research on COVID-19. Jianfu Li, Ekin Soysal, Jiang Bian 0001, Scott L. DuVall, Elizabeth Hanchrow, Kristine E. Lynch, Michael E. Matheny, Karthik Natarajan, Lucila Ohno-Machado, Serguei V. S. Pakhomov, Ruth M. Reeves, Amy M. Sitapati, Swapna Abhyankar, Theresa A. Cullen, Jami Deckard, Xiaoqian Jiang, Robert Murphy, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Context-guided fully convolutional networks for joint craniomaxillofacial bone segmentation and landmark digitization
Jun Zhang 0018, Mingxia Liu 0001, Li Wang 0026, Peng Yuan 0001, Jianfu Li, Steve G. Shen, Ken-Chung Chen, James J. Xia, Dinggang Shen |
Medical Image Anal. | 6 |
| 2017 | Joint Craniomaxillofacial Bone Segmentation and Landmark Digitization by Context-Guided Fully Convolutional Networks
Jun Zhang 0018, Mingxia Liu 0001, Li Wang 0026, Peng Yuan 0001, Jianfu Li, Steve G. Shen, Ken-Chung Chen, James J. Xia, Dinggang Shen |
MICCAI (2) | 6 |
| 2017 | Reconstruction-Based Digital Dental Occlusion of the Partially Edentulous DentitionabstractPartially edentulous dentition presents a challenging problem for the surgical planning of digital dental occlusion in the field of craniomaxillofacial surgery because of the incorrect maxillomandibular distance caused by missing teeth. We propose an innovative approach called Dental Reconstruction with Symmetrical Teeth (DRST) to achieve accurate dental occlusion for the partially edentulous cases. In this DRST approach, the rigid transformation between two symmetrical teeth existing on the left and right dental model is estimated through probabilistic point registration by matching the two shapes. With the estimated transformation, the partially edentulous space can be virtually filled with the teeth in its symmetrical position. Dental alignment is performed by digital dental occlusion reestablishment algorithm with the reconstructed complete dental model. Satisfactory reconstruction and occlusion results are demonstrated with the synthetic and real partially edentulous models. James J. Xia, Jianfu Li, Xiaobo Zhou 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2015 | Automated Three-Piece Digital Dental Articulation
Jianfu Li, Flavio Ferraz, Shunyao Shen, Yi-Fang Lo, Xiaoyan Zhang 0002, Peng Yuan 0001, Ken-Chung Chen, Jaime Gateno, Xiaobo Zhou 0001, James J. Xia |
MICCAI (1) | 1 |
| 2014 | Estimating Anatomically-Correct Reference Model for Craniomaxillofacial Deformity via Sparse Representation
Li Wang 0026, Yaozong Gao, Ken-Chung Chen, Jianfu Li, Steve G. Shen, Philip K. M. Lee, Ben Chow, James J. Xia, Dinggang Shen |
MICCAI (2) | 6 |
| 2012 | Key frame selection based on Jensen-Rényi divergence
Qing Xu 0002, Xiu Li 0005, Mateu Sbert, Jianfu Li |
ICPR | 6 |