Lixia Yao

dblp:16/7119 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-5187-6120ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2023 Workshop on Applied Data Science for Healthcare: Applications and New Frontiers of Generative Models for Healthcare
abstract
Built on the success of the past five years, KDD DSHealth 2023 will further catalyze the development of links between academic and industrial data science groups. The workshop aims to stimulate discussion on strategic areas for development and to facilitate future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community via timely topics, this year the workshop will focus on the applications and new development of generative models in healthcare, including the new development and application of LLMs. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature two invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, we will invite community members to submit their research works and bring them for discussion. The summary gives a brief description of the half-day workshop to be held on August 7th, 2023.
Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001
KDD7
2022 Workshop on Applied Data Science for Healthcare (DSHealth): Transparent and Human-centered AI
abstract
KDD DSHealth 2022, aims to build on the success of the past four years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the transparency and human-centered AI in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature four invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Journal of Healthcare Informatics Research. The summary gives a brief description of the full-day workshop to be held on August 14th, 2022.
Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001
KDD7
2022 Comparing PSO-based clustering over contextual vector embeddings to modern topic modeling
abstract
Efficient topic modeling is needed to support applications that aim at identifying main themes from a collection of documents. In the present paper, a reduced vector embedding representation and particle swarm optimization (PSO) are combined to develop a topic modeling strategy that is able to identify representative themes from a large collection of documents. Documents are encoded using a reduced, contextual vector embedding from a general-purpose pre-trained language model (sBERT). A modified PSO algorithm (pPSO) that tracks particle fitness on a dimension-by-dimension basis is then applied to these embeddings to create clusters of related documents. The proposed methodology is demonstrated on two datasets. The first dataset consists of posts from the online health forum r/Cancer and the second dataset is a standard benchmark for topic modeling which consists of a collection of messages posted to 20 different news groups. When compared to the state-of-the-art generative document models (i.e., ETM and NVDM), pPSO is able to produce interpretable clusters. The results indicate that pPSO is able to capture both common topics as well as emergent topics. Moreover, the topic coherence of pPSO is comparable to that of ETM and its topic diversity is comparable to NVDM. The assignment parity of pPSO on a document completion task exceeded 90% for the 20NewsGroups dataset. This rate drops to approximately 30% when pPSO is applied to the same Skip-Gram embedding derived from a limited, corpus-specific vocabulary which is used by ETM and NVDM.
Samuel Miles, Lixia Yao, Weilin Meng, Christopher M. Black, Zina Ben-Miled
Inf. Process. Manag.2
2022 Towards quality improvement of vaccine concept mappings in the OMOP vocabulary with a semi-automated method
abstract
The Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) provides a unified model to integrate disparate real-world data (RWD) sources. An integral part of the OMOP CDM is the Standardized Vocabularies (henceforth referred to as the OMOP vocabulary), which enables organization and standardization of medical concepts across various clinical domains of the OMOP CDM. For concepts with the same meaning from different source vocabularies, one is designated as the standard concept, while the others are specified as non-standard or source concepts and mapped to the standard one. However, due to the heterogeneity of source vocabularies, there may exist mapping issues such as erroneous mappings and missing mappings in the OMOP vocabulary, which could affect the results of downstream analyses with RWD. In this paper, we focus on quality assurance of vaccine concept mappings in the OMOP vocabulary, which is necessary to accurately harness the power of RWD on vaccines. We introduce a semi-automated lexical approach to audit vaccine mappings in the OMOP vocabulary. We generated two types of vaccine-pairs: mapped and unmapped, where mapped vaccine-pairs are pairs of vaccine concepts with a "Maps to" relationship, while unmapped vaccine-pairs are those without a "Maps to" relationship. We represented each vaccine concept name as a set of words, and derived term-difference pairs (i.e., name differences) for mapped and unmapped vaccine-pairs. If the same term-difference pair can be obtained by both mapped and unmapped vaccine-pairs, then this is considered as a potential mapping inconsistency. Applying this approach to the vaccine mappings in OMOP, a total of 2087 potentially mapping inconsistencies were obtained. A randomly selected 200 samples were evaluated by domain experts to identify, validate, and categorize the inconsistencies. Experts identified 95 cases revealing valid mapping issues. The remaining 105 cases were found to be invalid due to the external and/or contextual information used in the mappings that were not reflected in the concept names of vaccines. This indicates that our semi-automated approach shows promise in identifying mapping inconsistencies among vaccine concepts in the OMOP vocabulary.
Rashmie Abeysinghe, Adam Black, Denys Kaduk, Christian Reich, Lixia Yao, Licong Cui
J. Biomed. Informatics7
2022 Inferring the patient's age from implicit age clues in health forum posts
abstract
Broader patient-reported experiences in oncology are largely unknown due to the lack of available information from traditional data sources. Online health community data provide an exploratory way to uncover these experiences at a large scale. Analyzing these data can guide further studies towards understanding patients' needs and experiences. However, analysis of online health data is inherently difficult due to the unstructured nature of these data and the variety of ways information can be expressed over text. Specifically, subscribers may not disclose critical information such as the age of the patient in their posts. In fact, the number of health forum posts that explicitly mention the age of the patient is significantly lower than the number of posts that do not include this information in the Reddit r/Cancer health forum under consideration in the present paper. Health-focused studies often need to consider or control for age as a confounder, hence the importance of having sufficient age data. This paper presents a methodology that can help classify health forum posts according to four age groups (0-17, 18-39, 40-64 and 65 + years) even when the posts do not contain explicit mention of the age of the patient. First, the subset of the posts that include explicit mention of the age of the patient is identified. Second, the explicit age clues are removed from these posts and used to train the proposed age classifier. The resulting classifier is able to infer the age of the patient using only implicit age clues with an average true positive rate (TPR) of 71%. This TPR is comparable to the average TPR of 69% obtained from human annotations for the same set of posts.
Christopher M. Black, Weilin Meng, Lixia Yao, Zina Ben-Miled
J. Biomed. Informatics3
2021 KDD Health Day/DSHealth 2021: Joint KDD 2021 Health Day and 2021 KDD Workshop on Applied Data Science for Healthcare: State of XAI and Trustworthiness in Health
abstract
KDD Health Day/DSHealth 2021, aims to build on the success of the past 3 years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the state of explainability and trustworthiness in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature 8 invited talks from eminent speakers across academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Artificial Intelligence in Medicine journal. The summary gives a brief description of the full-day workshop to be held on August, 2021 virtually.
Fei Wang 0001, Prithwish Chakraborty, Tao Xu 0020, Pei-Yun Sabrina Hsueh, Xudong Sun 0014, Gregor Stiglic, Gracy Crane, Jiang Bian 0001, Laleh Haghverdi, Lixia Yao, Florian Buettner 0001
KDD10
2020 Developing a Data Model for Patient Secure Messages Leveraging FHIR
Amrita De, Tinghao Feng, Xiaomeng Yue, Lixia Yao
AMIA5
2018 Probing Technology Innovation on Diseases via Patent Mining
Ming Huang 0006, Maryam Zolnoori, Lixia Yao
AMIA3
2018 Temporal sequence alignment in electronic health records for computable patient representation
Ming Huang 0006, Maryam Zolnoori, Nilay D. Shah, Lixia Yao
BIBM4
2018 OpenHI - An open source framework for annotating histopathological image
Pargorn Puttapirat, Haichuan Zhang 0001, Yuchen Lian, Chunbao Wang 0002, Xiangrong Zhang, Lixia Yao, Chen Li 0011
BIBM6
2018 Detecting Serendipitous Drug Usage in Social Media with Deep Neural Network Models
Boshu Ru, Dingcheng Li, Lixia Yao
BIBM3
2017 Identification of Clinically Meaningful Clusters of Multi-morbidity in a National Cohort of Adults Using Unsupervised Learning
Che Ngufor, Rozalina G. McCoy, Lixia Yao, Lindsey R. Sangaralingham, Shannon M. Dunlay, Nilay D. Shah
AMIA3
2017 DIR - A semantic information resource for healthcare datasets
abstract
It is important for data scientists to have a good understanding of the availability of relevant datasets as well as the content, structure, and existing analyses of these datasets. While a number of efforts are underway to integrate the large amount and variety of datasets, there is a lack of information resources that focus on specific learning needs of some targeted audiences. To address this gap, we have been developing a semantic Dataset Information Resource (DIR) framework to specifically address the challenges of entry-level data scientists in learning to identify, understand, and analyze major datasets with an initial focus on healthcare. The DIR does not contain actual data from the datasets but aims to provide comprehensive knowledge about the datasets and their analyses. The framework leverages Semantic Web technologies and the W3C Dataset Description Standard for knowledge integration and representation and includes natural language processing (NLP)-based methods to enable knowledge extraction and question answering. The prototype DIR implementation includes four major components-dataset metadata and related knowledge, search modules, question answering for frequently-asked questions, and blogs. And the DIR currently includes information on three commonly-used large and complex healthcare datasets: HCUP, MarketScan, and MIMIC. Initial usage evaluation based on health informatics students is encouraging. Further development is underway.
Mingna Zheng, Lixia Yao, Yaorong Ge
BIBM3
2017 Detecting Drinking-Related Contents on Social Media by Classifying Heterogeneous Data Types
Omar ElTayeby, Todd Eaglin, Malak Abdullah, David Burlinson, Wenwen Dou, Lixia Yao
IEA/AIE (2)6
2017 Estimating Disease Burden Using Google Trends and Wikipedia Data
Riyi Qiu, Mirsad Hadzikadic, Lixia Yao
IEA/AIE (2)3
2015 A survey of social media for understanding patient-reported medication outcomes
Kimberly Harris, Boshu Ru, Lixia Yao
AMIA3
2015 30 Day hospital readmission analysis
abstract
Readmissions to a hospital after procedures are costly and considered to be an indication of poor quality. As Per the Affordable Care Act of 2010, hospitals may be reimbursed at a reduced rate for patients readmitted to a hospital within 30 days of discharge. In this project, we used statistical and machine-learning methods to analyze the Nationwide Inpatient Sample dataset provided by HCUP (Healthcare Cost and Utilization Project) to identify various clinical, demographic and socio-economic factors that play crucial roles in predicting the revenue loss due to readmissions. Three medical conditions, namely chronic obstructive pulmonary disorder (COPD), total hip arthroplasty (THA), and total knee arthroplasty (TKA) have been primarily used for this purpose. Our analysis builds on both non-parametric and parametric statistical models and machine learning techniques such as Decision Tree, Gradient Boosting, Logistic Regression and Neural Networks. We evaluated and compared these models based on Area under ROC (AUC) and misclassification rate. By including visual analytics, this analysis not only enables the hospitals to compute the loss of revenue but also monitors their quality of service in a real-time fashion.
Ratna Madhuri Maddipatla, Mirsad Hadzikadic, Dipti Patel Misra, Lixia Yao
IEEE BigData4
2013 Systematic evaluation of unmet medical needs from multiple dimensionalities - A feasibility study
Lixia Yao
AMIA2
2012 Mining Electronic Health Records to Identify Drug Combinations
Lixia Yao
AMIA1
2011 Benchmarking Ontologies: Bigger or Better?
abstract
A scientific ontology is a formal representation of knowledge within a domain, typically including central concepts, their properties, and relations. With the rise of computers and high-throughput data collection, ontologies have become essential to data mining and sharing across communities in the biomedical sciences. Powerful approaches exist for testing the internal consistency of an ontology, but not for assessing the fidelity of its domain representation. We introduce a family of metrics that describe the breadth and depth with which an ontology represents its knowledge domain. We then test these metrics using (1) four of the most common medical ontologies with respect to a corpus of medical documents and (2) seven of the most popular English thesauri with respect to three corpora that sample language from medicine, news, and novels. Here we show that our approach captures the quality of ontological representation and guides efforts to narrow the breach between ontology and collective discourse within a domain. Our results also demonstrate key features of medical ontologies, English thesauri, and discourse from different domains. Medical ontologies have a small intersection, as do English thesauri. Moreover, dialects characteristic of distinct domains vary strikingly as many of the same words are used quite differently in medicine, news, and novels. As ontologies are intended to mirror the state of knowledge, our methods to tighten the fit between ontology and domain will increase their relevance for new areas of biomedical science and improve the accuracy and power of inferences computed across them.
Lixia Yao, Anna Divoli, Ilya Mayzus, James A. Evans, Andrey Rzhetsky
PLoS Comput. Biol.1
2008 Quantitative systems-level determinants of drug targets
Lixia Yao, Andrey Rzhetsky
BMC Bioinform.1