VLDB 2026 Research / reviewers in the wild / expert
Jungwei Fan 0001
dblp:254/1131-1 · also Jungwei W. Fan 0001, Jungwei Wilfred Fan
· DBLP profile ↗
22ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0001-6349-3752ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Developing and sustaining inclusive language in biomedical informatics communications: an AMIA Board of Directors endorsed paper on the Inclusive Language and Context Style GuidelinesabstractOBJECTIVES: In 2023, AMIA's Inclusive Language and Context Style Guidelines (the "Guidelines") were approved by the Board of Directors and made a publicly available resource. This work began in 2021 through AMIA's DEI Task Force and subsequent DEI Committee; many members provided input, feedback, and time to create the Guidelines. In this paper, the authors provide a transparent account of the origin, development, contents, and dissemination of the Guidelines and share plans for their future development and use. MATERIALS AND METHODS: Our approach to drafting, refining, and distributing the Guidelines included consulting existing language guides, AMIA member reviews, external expert reviews, webinars, and workshops. Through an iterative approach to drafting and refining the Guidelines, the authors consulted relevant language guidelines and many experts throughout and beyond the AMIA community. RESULTS: The Inclusive Language Context Guidelines were formally approved by the AMIA Board of Directors on February 15, 2023. The Guidelines included four principles to be considered in scientific communications: Plurality, Precision, Transparency, and Destigmatization. DISCUSSION: A moment of vulnerability where an AMIA member raised concerns about the use of harmful language during a presentation resulted in the creation of a principled approach to support inclusive language within biomedical and health informatics communications. We envision that the Guidelines will support health equity by challenging dominant public narratives around health, fostering stronger interdisciplinary collaboration and critical thinking about the impact of language, and creating a more welcoming environment for the broader AMIA community. This work could not have been completed without the support of many AMIA members and other researchers in biomedical and health informatics. The Guidelines are a living document that will continue to be updated with input and feedback from the AMIA community into the future. Oliver J. Bear Don't Walk IV, Shefali Haldar, Duo Helen Wei, Hu Huang 0004, Rebecca L. Rivera, Jungwei Fan 0001, Vipina Kuttichi Keloth, Tiffany I. Leung, Pooja M. Desai, Diane M. Korngiebel, Lisa Grossman Liu, Adrienne Pichon, Vignesh Subbian, Tony Solomonides, Laura K. Wiley, Omolola Ogunyemi, Gretchen Purcell Jackson, Irene Dankwa-Mullan, Lisa Dirks, Avery Rose Everhart, Andrea G. Parker, Bradley E. Iott, Clair A. Kronk, Randi E. Foraker, Krista G. Martin, Tara Anand, Salvatore G. Volpe, Nathan Yung, Rubina F. Rizvi, Robert James Lucero, Tiffani J. Bright |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Automatic uncovering of patient primary concerns in portal messages using a fusion framework of pretrained language modelsabstractOBJECTIVES: The surge in patient portal messages (PPMs) with increasing needs and workloads for efficient PPM triage in healthcare settings has spurred the exploration of AI-driven solutions to streamline the healthcare workflow processes, ensuring timely responses to patients to satisfy their healthcare needs. However, there has been less focus on isolating and understanding patient primary concerns in PPMs-a practice which holds the potential to yield more nuanced insights and enhances the quality of healthcare delivery and patient-centered care. MATERIALS AND METHODS: We propose a fusion framework to leverage pretrained language models (LMs) with different language advantages via a Convolution Neural Network for precise identification of patient primary concerns via multi-class classification. We examined 3 traditional machine learning models, 9 BERT-based language models, 6 fusion models, and 2 ensemble models. RESULTS: The outcomes of our experimentation underscore the superior performance achieved by BERT-based models in comparison to traditional machine learning models. Remarkably, our fusion model emerges as the top-performing solution, delivering a notably improved accuracy score of 77.67 ± 2.74% and an F1 score of 74.37 ± 3.70% in macro-average. DISCUSSION: This study highlights the feasibility and effectiveness of multi-class classification for patient primary concern detection and the proposed fusion framework for enhancing primary concern detection. CONCLUSIONS: The use of multi-class classification enhanced by a fusion of multiple pretrained LMs not only improves the accuracy and efficiency of patient primary concern identification in PPMs but also aids in managing the rising volume of PPMs in healthcare, ensuring critical patient communications are addressed promptly and accurately. Jungwei Fan 0001, Aditya Khurana, Sunyang Fu, Dezhi Wu, Ming Huang 0006 |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | FedFSA: Hybrid and federated framework for functional status ascertainment across institutions
Sunyang Fu, Heling Jia, Maria Vassilaki, Vipina Kuttichi Keloth, Yifang Dang, Yujia Zhou 0003, Muskan Garg, Ronald C. Petersen, Jennifer L. St. Sauver, Sungrim Moon, Liwei Wang 0010, Andrew Wen, Fang Li 0011, Hua Xu 0001, Cui Tao, Jungwei Fan 0001, Sunghwan Sohn |
J. Biomed. Informatics | 16 |
| 2022 | Quality Assessment of Functional Status Documentation in EHR Across Institutions
Sunyang Fu, Maria Vassilaki, Omar A. Ibrahim, Ronald C. Petersen, Jennifer L. St. Sauver, Liwei Wang 0010, Jungwei Fan 0001, Sunghwan Sohn |
AMIA | 7 |
| 2021 | Patient Asynchronous Response to Coronavirus Disease 2019 (COVID-19): A Retrospective Analysis of Patient Portal Messages
Ming Huang 0006, Aditya Khurana, George M. Mastorakos, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Christi A. Patten, Jungwei Fan 0001 |
AMIA | 14 |
| 2021 | Disparity analysis of patient portal messaging use for COVID-19 in urban versus rural locality
Ming Huang 0006, Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Yanshan Wang, Nansu Zong, Yue Yu 0012, Julie E. Prigge, Brian Costello, Nilay D. Shah, Henry Ting, Chyke Doubeni, Jungwei Fan 0001, Christi A. Patten |
AMIA | 14 |
| 2021 | Development of a Clinical Question-Answering Corpus with Realistic Multi-Answer Challenges
Sungrim Moon, Jungwei Fan 0001 |
AMIA | 3 |
| 2021 | A Scoping Review of Informatics Research for Clinical Practice Variation
Sunghwan Sohn, Sungrim Moon, Larry J. Prokop, Victor M. Montori, Jungwei Fan 0001 |
AMIA | 5 |
| 2021 | An aberration detection-based approach for sentinel syndromic surveillance of COVID-19 and other novel influenza-like illnesses
Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Sunyang Fu, Sunghwan Sohn, Jacob A. Kugel, Vinod Kaggal, Ming Huang 0006, Yanshan Wang, Feichen Shen, Jungwei Fan 0001 |
J. Biomed. Informatics | 12 |
| 2020 | Deep Semantic Embeddings and Clustering to Facilitate Identification of Transportation Barriers in Patient Portal Messages
Ming Huang 0006, Jungwei Fan 0001 |
AMIA | 2 |
| 2020 | Annotating Chronic Pain Episodes in EHR Text: Guideline Development and Corpus Analysis
Luke A. Carlson, Molly M. Jeffery, Sunyang Fu, Rozalina G. McCoy, Yanshan Wang, W. M. Hooten, Jennifer L. St. Sauver, Jungwei Fan 0001 |
AMIA | 10 |
| 2020 | Big Impact from Small Data: Unsupervised Machine Learning Approaches for Chronic Pain Patient Subgrouping
Luke A. Carlson, Jennifer L. St. Sauver, Sunyang Fu, Ahmad P. Tafti, Jungwei Fan 0001, Molly M. Jeffery, Rozalina G. McCoy, Yanshan Wang |
AMIA | 5 |
| 2020 | A Perturbation Approach to Assessing BERT Robustness for Different Linguistic Aspects in Medical Question-Answering
Mohamed Y. Elwazir, Andrew Wen, Sungrim Moon, Jungwei Fan 0001 |
AMIA | 4 |
| 2020 | A Deep Profiling and Visualization Framework to Audit Clinical Assessment VariationabstractClinical assessment variation (CAV) has a profound impact on patient outcomes, and appropriate tooling is critically needed to help understand and guide necessary interventions. In this study, we propose an intuitive approach to visualizing CAV and summarizing the contexts pertinent to decision-making. By superimposing the response variable and clusters learned according to the explanatory variables, a color-coded 2D scatter plot can be rendered to show the spatial proximity and semantic composition of the clusters. Without loss of generality, an example application on preoperative patient assessment demonstrated the approach can assist in auditing inconsistent human decisions and informing the reconciliation process. The methods will also benefit refining of clinical assessment guidelines by systematically eliciting practice-based knowledge. Andrew Wen, Feichen Shen, Sungrim Moon, Jungwei Fan 0001 |
CBMS | 5 |
| 2020 | A review of auditing techniques for the Unified Medical Language SystemabstractOBJECTIVE: The study sought to describe the literature related to the development of methods for auditing the Unified Medical Language System (UMLS), with particular attention to identifying errors and inconsistencies of attributes of the concepts in the UMLS Metathesaurus. MATERIALS AND METHODS: We applied the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) approach by searching the MEDLINE database and Google Scholar for studies referencing the UMLS and any of several terms related to auditing, error detection, and quality assurance. A qualitative analysis and summarization of articles that met inclusion criteria were performed. RESULTS: Eighty-three studies were reviewed in detail. We first categorized techniques based on various aspects including concepts, concept names, and synonymy (n = 37), semantic type assignments (n = 36), hierarchical relationships (n = 24), lateral relationships (n = 12), ontology enrichment (n = 8), and ontology alignment (n = 18). We also categorized the methods according to their level of automation (ie, automated systematic, automated heuristic, or manual) and the type of knowledge used (ie, intrinsic or extrinsic knowledge). CONCLUSIONS: This study is a comprehensive review of the published methods for auditing the various conceptual aspects of the UMLS. Categorizing the auditing techniques according to the various aspects will enable the curators of the UMLS as well as researchers comprehensive easy access to this wealth of knowledge (eg, for auditing lateral relationships in the UMLS). We also reviewed ontology enrichment and alignment techniques due to their critical use of and impact on the UMLS. Zhe He 0001, Duo Helen Wei, Vipina Kuttichi Keloth, Jungwei Fan 0001, Luke Lindemann, James J. Cimino, Yehoshua Perl |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Enhancing Clinical Information Retrieval through Context-Aware Queries and IndicesabstractThe big data revolution has created a hefty demand for searching large-scale electronic health records (EHRs) to support clinical practice, research, and administration. Despite the volume of data involved, fast and accurate identification of clinical narratives pertinent to a clinical case being seen by any given provider is crucial for decision-making at the point of care. In the general domain, this capability is accomplished through a combination of the inverted index data structure, horizontal scaling, and information retrieval (IR) scoring algorithms. These technologies are also being used in the clinical domain, but have met limited success, particularly as clinical cases become more complex. One barrier affecting clinical performance is that contextual information, such as negation, temporality, and the subject of clinical mentions, impact clinical relevance but is not considered in general IR methodologies. In this study, we implemented a solution by identifying and incorporating the aforementioned semantic contexts as part of IR indexing/scoring with Elasticsearch. Experiments were conducted in comparison to baseline approaches with respect to: 1) evaluation of the impact on the quality (relevance) of the returned results, and 2) evaluation of the impact on execution time and storage requirements. The results showed a 5.1-23.1% improvement in retrieval quality, along with achieving 35% faster query execution time. Cost-wise, the solution required 1.5-2 times larger space and about 3 times increase in indexing time. The higher relevance demonstrated the merit of incorporating contextual information into clinical IR, and the near-constant increase in time and space suggested promising scalability. Andrew Wen, Yanshan Wang, Vinod Kaggal, Sijia Liu 0002, Jungwei Fan 0001 |
IEEE BigData | 6 |
| 2015 | Risk factor detection for heart disease by applying text analytics in electronic medical recordsabstractIn the United States, about 600,000 people die of heart disease every year. The annual cost of care services, medications, and lost productivity reportedly exceeds 108.9 billion dollars. Effective disease risk assessment is critical to prevention, care, and treatment planning. Recent advancements in text analytics have opened up new possibilities of using the rich information in electronic medical records (EMRs) to identify relevant risk factors. The 2014 i2b2/UTHealth Challenge brought together researchers and practitioners of clinical natural language processing (NLP) to tackle the identification of heart disease risk factors reported in EMRs. We participated in this track and developed an NLP system by leveraging existing tools and resources, both public and proprietary. Our system was a hybrid of several machine-learning and rule-based components. The system achieved an overall F1 score of 0.9185, with a recall of 0.9409 and a precision of 0.8972. Manabu Torii, Jungwei Fan 0001, Weili Yang, Theodore Lee, Matthew T. Wiley, Daniel Zisook, Yang Huang 0008 |
J. Biomed. Informatics | 2 |
| 2014 | Building a Treebank of hospital discharge summaries
Yang Huang 0008, Jungwei Fan 0001, Elly W. Yang, Hua Xu 0001 |
AMIA | 2 |
| 2013 | Research and applications: Syntactic parsing of clinical text: guideline and corpus development with handling ill-formed sentencesabstractOBJECTIVE: To develop, evaluate, and share: (1) syntactic parsing guidelines for clinical text, with a new approach to handling ill-formed sentences; and (2) a clinical Treebank annotated according to the guidelines. To document the process and findings for readers with similar interest. METHODS: Using random samples from a shared natural language processing challenge dataset, we developed a handbook of domain-customized syntactic parsing guidelines based on iterative annotation and adjudication between two institutions. Special considerations were incorporated into the guidelines for handling ill-formed sentences, which are common in clinical text. Intra- and inter-annotator agreement rates were used to evaluate consistency in following the guidelines. Quantitative and qualitative properties of the annotated Treebank, as well as its use to retrain a statistical parser, were reported. RESULTS: A supplement to the Penn Treebank II guidelines was developed for annotating clinical sentences. After three iterations of annotation and adjudication on 450 sentences, the annotators reached an F-measure agreement rate of 0.930 (while intra-annotator rate was 0.948) on a final independent set. A total of 1100 sentences from progress notes were annotated that demonstrated domain-specific linguistic features. A statistical parser retrained with combined general English (mainly news text) annotations and our annotations achieved an accuracy of 0.811 (higher than models trained purely with either general or clinical sentences alone). Both the guidelines and syntactic annotations are made available at https://sourceforge.net/projects/medicaltreebank. CONCLUSIONS: We developed guidelines for parsing clinical text and annotated a corpus accordingly. The high intra- and inter-annotator agreement rates showed decent consistency in following the guidelines. The corpus was shown to be useful in retraining a statistical parser that achieved moderate accuracy. Jungwei Fan 0001, Elly W. Yang, Min Jiang 0007, Rashmi Prasad, Richard M. Loomis, Daniel Zisook, Joshua C. Denny, Hua Xu 0001, Yang Huang 0008 |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | A review of auditing methods applied to the content of controlled biomedical terminologies
Jungwei Fan 0001, David M. Baorto, Chunhua Weng, James J. Cimino |
J. Biomed. Informatics | 2 |
| 2007 | Gene symbol disambiguation using knowledge-based profilesabstractMOTIVATION: The ambiguity of biomedical entities, particularly of gene symbols, is a big challenge for text-mining systems in the biomedical domain. Existing knowledge sources, such as Entrez Gene and the MEDLINE database, contain information concerning the characteristics of a particular gene that could be used to disambiguate gene symbols. RESULTS: For each gene, we create a profile with different types of information automatically extracted from related MEDLINE abstracts and readily available annotated knowledge sources. We apply the gene profiles to the disambiguation task via an information retrieval method, which ranks the similarity scores between the context where the ambiguous gene is mentioned, and candidate gene profiles. The gene profile with the highest similarity score is then chosen as the correct sense. We evaluated the method on three automatically generated testing sets of mouse, fly and yeast organisms, respectively. The method achieved the highest precision of 93.9% for the mouse, 77.8% for the fly and 89.5% for the yeast. AVAILABILITY: The testing data sets and disambiguation programs are available at http://www.dbmi.columbia.edu/~hux7002/gsd2006 Hua Xu 0001, Jungwei Fan 0001, George Hripcsak, Eneida A. Mendonça, Marianthi Markatou, Carol Friedman |
Bioinform. | 2 |
| 2007 | Using contextual and lexical features to restructure and validate the classification of biomedical conceptsabstractBACKGROUND: Biomedical ontologies are critical for integration of data from diverse sources and for use by knowledge-based biomedical applications, especially natural language processing as well as associated mining and reasoning systems. The effectiveness of these systems is heavily dependent on the quality of the ontological terms and their classifications. To assist in developing and maintaining the ontologies objectively, we propose automatic approaches to classify and/or validate their semantic categories. In previous work, we developed an approach using contextual syntactic features obtained from a large domain corpus to reclassify and validate concepts of the Unified Medical Language System (UMLS), a comprehensive resource of biomedical terminology. In this paper, we introduce another classification approach based on words of the concept strings and compare it to the contextual syntactic approach. RESULTS: The string-based approach achieved an error rate of 0.143, with a mean reciprocal rank of 0.907. The context-based and string-based approaches were found to be complementary, and the error rate was reduced further by applying a linear combination of the two classifiers. The advantage of combining the two approaches was especially manifested on test data with sufficient contextual features, achieving the lowest error rate of 0.055 and a mean reciprocal rank of 0.969. CONCLUSION: The lexical features provide another semantic dimension in addition to syntactic contextual features that support the classification of ontological concepts. The classification errors of each dimension can be further reduced through appropriate combination of the complementary classifiers. Jungwei Fan 0001, Hua Xu 0001, Carol Friedman |
BMC Bioinform. | 1 |