EDBT 2026 Demo / reviewers in the wild / expert
Lixia Yao
dblp:16/7119
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
4since 2021 · last 2023
0000-0002-5187-6120ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 3Information Retrieval & Web Search · 1Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Workshop on Applied Data Science for Healthcare: Applications and New Frontiers of Generative Models for HealthcareabstractBuilt on the success of the past five years, KDD DSHealth 2023 will further catalyze the development of links between academic and industrial data science groups. The workshop aims to stimulate discussion on strategic areas for development and to facilitate future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community via timely topics, this year the workshop will focus on the applications and new development of generative models in healthcare, including the new development and application of LLMs. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature two invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, we will invite community members to submit their research works and bring them for discussion. The summary gives a brief description of the half-day workshop to be held on August 7th, 2023. Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001 |
KDD | 7 |
| 2022 | Workshop on Applied Data Science for Healthcare (DSHealth): Transparent and Human-centered AIabstractKDD DSHealth 2022, aims to build on the success of the past four years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the transparency and human-centered AI in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature four invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Journal of Healthcare Informatics Research. The summary gives a brief description of the full-day workshop to be held on August 14th, 2022. Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001 |
KDD | 7 |
| 2022 | Comparing PSO-based clustering over contextual vector embeddings to modern topic modelingabstractEfficient topic modeling is needed to support applications that aim at identifying main themes from a collection of documents. In the present paper, a reduced vector embedding representation and particle swarm optimization (PSO) are combined to develop a topic modeling strategy that is able to identify representative themes from a large collection of documents. Documents are encoded using a reduced, contextual vector embedding from a general-purpose pre-trained language model (sBERT). A modified PSO algorithm (pPSO) that tracks particle fitness on a dimension-by-dimension basis is then applied to these embeddings to create clusters of related documents. The proposed methodology is demonstrated on two datasets. The first dataset consists of posts from the online health forum r/Cancer and the second dataset is a standard benchmark for topic modeling which consists of a collection of messages posted to 20 different news groups. When compared to the state-of-the-art generative document models (i.e., ETM and NVDM), pPSO is able to produce interpretable clusters. The results indicate that pPSO is able to capture both common topics as well as emergent topics. Moreover, the topic coherence of pPSO is comparable to that of ETM and its topic diversity is comparable to NVDM. The assignment parity of pPSO on a document completion task exceeded 90% for the 20NewsGroups dataset. This rate drops to approximately 30% when pPSO is applied to the same Skip-Gram embedding derived from a limited, corpus-specific vocabulary which is used by ETM and NVDM. Samuel Miles, Lixia Yao, Weilin Meng, Christopher M. Black, Zina Ben-Miled |
Inf. Process. Manag. | 2 |
| 2021 | KDD Health Day/DSHealth 2021: Joint KDD 2021 Health Day and 2021 KDD Workshop on Applied Data Science for Healthcare: State of XAI and Trustworthiness in HealthabstractKDD Health Day/DSHealth 2021, aims to build on the success of the past 3 years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the state of explainability and trustworthiness in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature 8 invited talks from eminent speakers across academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Artificial Intelligence in Medicine journal. The summary gives a brief description of the full-day workshop to be held on August, 2021 virtually. Fei Wang 0001, Prithwish Chakraborty, Tao Xu 0020, Pei-Yun Sabrina Hsueh, Xudong Sun 0014, Gregor Stiglic, Gracy Crane, Jiang Bian 0001, Laleh Haghverdi, Lixia Yao, Florian Buettner 0001 |
KDD | 10 |
| 2015 | 30 Day hospital readmission analysisabstractReadmissions to a hospital after procedures are costly and considered to be an indication of poor quality. As Per the Affordable Care Act of 2010, hospitals may be reimbursed at a reduced rate for patients readmitted to a hospital within 30 days of discharge. In this project, we used statistical and machine-learning methods to analyze the Nationwide Inpatient Sample dataset provided by HCUP (Healthcare Cost and Utilization Project) to identify various clinical, demographic and socio-economic factors that play crucial roles in predicting the revenue loss due to readmissions. Three medical conditions, namely chronic obstructive pulmonary disorder (COPD), total hip arthroplasty (THA), and total knee arthroplasty (TKA) have been primarily used for this purpose. Our analysis builds on both non-parametric and parametric statistical models and machine learning techniques such as Decision Tree, Gradient Boosting, Logistic Regression and Neural Networks. We evaluated and compared these models based on Area under ROC (AUC) and misclassification rate. By including visual analytics, this analysis not only enables the hospitals to compute the loss of revenue but also monitors their quality of service in a real-time fashion. Ratna Madhuri Maddipatla, Mirsad Hadzikadic, Dipti Patel Misra, Lixia Yao |
IEEE BigData | 4 |