David L. Schwartz

dblp:311/0150 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2024
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2024 AI-Ready Multimodal Data Pipeline to Enrich Cancer Care
abstract
This study reports on the progress in designing and developing a framework for integrating heterogeneous datasets—structured, semi-structured, and unstructured—into an AI-ready multimodal data pipeline aimed at predicting radiation therapy interruptions (RTI) and enhancing patient care navigation. The AI-Ready dataset incorporates a broad set of information, including patient demographics, health data, clinical notes, medical imaging, and data on social determinants of health. Preliminary results indicate that this pipeline effectively integrates diverse distributed data sources, providing a foundation for training AI models capable of generating reliable and actionable predictions.
Rezaur Rashid, Soheil Hashtarkhani, Fekede Asefa Kumsa, Lokesh K. Chinthala, Brianna White, Janet A Zink, Christopher L. Brett, Robert L. Davis, David L. Schwartz, Arash Shaban-Nejad
IEEE Big Data9
2021 PRESIDE: A Judge Entity Recognition and Disambiguation Model for US District Court Records
abstract
The docket sheet of a court case contains a wealth of information about the progression of a case, the parties’ and judge’s decision-making along the way, and the case’s ultimate outcome that can be used in analytical applications. However, the unstructured text of the docket sheet and the terse and variable phrasing of docket entries require the development of new models to identify key entities to enable analysis at a systematic level. We developed a judge entity recognition language model and disambiguation pipeline for US District Court records. Our model can robustly identify mentions of judicial entities in free text (~99% F-1 Score) and outperforms general state-of-the-art language models by 13%. Our disambiguation pipeline is able to robustly identify both appointed and non-appointed judicial actors and correctly infer the type of appointment (~99% precision). Lastly, we show with a case study on in forma pauperis decision-making that there is substantial error (~30%) attributing decision outcomes to judicial actors if the free text of the docket is not used to make the identification and attribution.
Adam R. Pah, Christian J. Rozolis, David L. Schwartz, Charlotte Alexander
IEEE BigData3