EDBT 2026 Demo / reviewers in the wild / expert
Saptarshi Purkayastha
dblp:53/10677
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-3625-534XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Logical Observation Identifiers Names and Codes mapping with biomedical natural language processing models: enabling scalable health information exchange via the Open Concept LababstractOBJECTIVES: Efficient exchange of health information requires consistent representation of clinical concepts across laboratories, hospitals, and public health systems. LOINC supports this interoperability by standardizing laboratory test codes, but mapping remains difficult when datasets are incomplete, inconsistently formatted, or structurally diverse. These challenges often create a mismatch between algorithmic performance in controlled settings and real-world deployment. This study aimed to develop a biomedical natural language processing (NLP) approach for mapping heterogeneous laboratory test strings to LOINC v2.81 and to compare its performance with established algorithms in the Open Concept Lab (OCL) Mapper. MATERIALS AND METHODS: We implemented a ScispaCy-based pipeline (ScispaCy-LOINC) that identifies clinical entities, links them to UMLS Concept Unique Identifiers, assembles LOINC codes from LOINC parts, and ranks candidates using a weighted scoring system. Overall and ranked performance was evaluated against 2 OCL algorithms, Elasticsearch Keyword Retrieval (OCL-Keyword) and MiniLM Semantic Search (OCL-Semantic), on 2 datasets: MIMIC-IV lab_d_items and a LOINC-mapped subset of the CIEL interface terminology v2025-07-15. RESULTS: In MIMIC-IV, the ScispaCy-LOINC achieved the highest coverage, correctly identifying the LOINC code in 42.3% of cases, outperforming OCL-Keyword (19.5%) and OCL-Semantic (21.4%). In the CIEL dataset, OCL-Semantic achieved the highest coverage (54.4%), followed by OCL-Keyword (46.9%) and ScispaCy-LOINC (28.4%). DISCUSSION: These results indicate that ScispaCy-LOINC is particularly effective for noisier or structurally sparse inputs, whereas OCL-based approaches perform better for more standardized terminologies, highlighting complementary algorithmic strengths. CONCLUSION: ScispaCy-LOINC offers a flexible approach to LOINC mapping and demonstrates complementary strengths relative to existing OCL algorithms. These findings support the development of an integrated framework that combines algorithmic strategies to improve robustness across diverse clinical datasets. Parvati Naliyatthaliyazchayil, Venkat Ramana Sangam, Joseph Amlung, Andrew S. Kanter, Saptarshi Purkayastha, Jonathan Payne |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | Utilizing ETL Processes to Enhance Healthcare Education: Migrating MIMIC-IV to OpenEMRabstractThe integration of real-world clinical data from the Medical Information Mart for Intensive Care IV (MIMIC-IV) database into the Open Electronic Medical Records (OpenEMR) system presents a unique opportunity to develop and implement innovative Extract, Transform, and Load (ETL) methodologies. This study focuses on the complex process of migrating 40,000 patient records from 33 tables in MIMIC-IV to the intricate schema of OpenEMR, which comprises 264 tables. The primary objective is to outline the novel approaches employed in the ETL process to ensure data fidelity, integrity, and usability within OpenEMR’s frontend interface. The methodology involves a three-stage process: extraction, transformation, and loading. During the extraction phase, relevant data is carefully selected from MIMIC-IV, taking into account data privacy considerations. The transformation phase involves intricate data mapping and manipulation to align the extracted data with OpenEMR’s schema, addressing challenges such as schema mismatches and data format inconsistencies. Finally, the loading phase ensures the accurate and efficient population of OpenEMR with the transformed data. The ETL process is designed to maintain data quality and integrity throughout the migration, employing robust validation and error-handling mechanisms. This study contributes to the advancement of ETL methodologies in the context of integrating real-world clinical data into electronic medical record systems. The detailed description of the innovative approaches employed in this project might serve as a valuable resource for researchers and practitioners working on similar data migration tasks. Mohith Surya Kiran Kasula, Sri Harsha Sudalagunta, Keerthika Sunchu, Saptarshi Purkayastha |
CBMS | 4 |
| 2024 | Interpolating and Forecasting Wound Trajectory using Machine Learning ApproachesabstractChronic wound treatment and management heavily rely on accurate patient outcome predictions. As chronic wounds become increasingly prevalent in many US states, it’s crucial to accurately predict and classify wound progress for timely and effective clinical intervention. The trajectory and outcomes of wounds inform multiple aspects of care, from estimating healing times to understanding disparities in treatment quality and outcomes. For accurate predictions, robust time-series forecast models are required, necessitating comprehensive and consistent wound data. Yet, much of this data is frequently marred by gaps and inaccuracies. In this work, we combined two approaches to overcome this challenge: filling in missing clinical data and forecasting wound healing trajectories. We employed various interpolation techniques, time series, and machine-learning models, including Holt-Winters, ARIMA, Prophet, LSTM, and BiLSTM. Our study assessed 14,571 wounds from 6,171 patients over a five-year span, post-data imputation and interpolation. We introduced a novel approach to firstly impute missing data using XGBoost and then choose the best interpolation method for absent wound data, ensuring precise forecasting without distorting the dataset’s demographic makeup. Data for this research was sourced from four wound centers in Indiana and eight research studies, focusing on a diverse range of wound-related demographics, clinical, and omics data. This comprehensive approach enhances wound trajectory predictions, providing clinicians with valuable insights to formulate personalized treatment plans for both individual patients and demographic groups. This will aid clinicians in better understanding wound healing trends and crafting targeted treatment strategies. Atika Rahman Paddo, Shikhar Shukla, Shomita S. Steiner, Saptarshi Purkayastha, Chandan K. Sen |
CBMS | 4 |
| 2023 | A General-Purpose AI Assistant Embedded in an Open-Source Radiology Information System
Saptarshi Purkayastha, Rohan Satya Isaac, Sharon Anthony, Shikhar Shukla, Elizabeth A. Krupinski, Joshua A. Danish, Judy Gichoya |
AIME | 1 |
| 2023 | Foundational domains and competencies for baccalaureate health informatics educationabstractBACKGROUND: Foundational domains are the building blocks of educational programs. The lack of foundational domains in undergraduate health informatics (HI) education can adversely affect the development of rigorous curricula and may impede the attainment of CAHIIM accreditation of academic programs. OBJECTIVE: This White Paper presents foundational domains developed by AMIA's Academic Forum Baccalaureate Education Committee (BEC) which include corresponding competencies (knowledge, skills, and attitudes) that are intended for curriculum development and CAHIIM accreditation quality assessment for undergraduate education in applied health informatics. METHODS: The AMIA BEC used the previously published master's foundational domains as a guide to creating a set of competencies for health informatics at the undergraduate level to assess graduates from undergraduate health informatics programs for competence at graduation. A consensus method was used to adapt the domains for undergraduate level course work and harmonize the foundational domains with the currently adapted domains for HI master's education. RESULTS: Ten foundational domains were developed to support the development and evaluation of baccalaureate health informatics education. DISCUSSION: This article will inform future work towards building CAHIIM accreditation standards to ensure that higher education institutions meet acceptable levels of quality for undergraduate health informatics education. Saif S. Khairat, Sue S. Feldman, Arif Rana, Mohammad Faysel, Saptarshi Purkayastha, Matthew Scotch, Christina Eldredge |
J. Am. Medical Informatics Assoc. | 5 |
| 2023 | MedShift: Automated Identification of Shift Data for Medical Image Dataset CurationabstractAutomated curation of noisy external data in the medical domain has long been in high demand, as AI technologies need to be validated using various sources with clean, annotated data. Identifying the variance between internal and external sources is a fundamental step in curating a high-quality dataset, as the data distributions from different sources can vary significantly and subsequently affect the performance of AI models. The primary challenges for detecting data shifts are - (1) accessing private data across healthcare institutions for manual detection and (2) the lack of automated approaches to learn efficient shift-data representation without training samples. To overcome these problems, we propose an automated pipeline called MedShift to detect top-level shift samples and evaluate the significance of shift data without sharing data between internal and external organizations. MedShift employs unsupervised anomaly detectors to learn the internal distribution and identify samples showing significant shiftness for external datasets, and then compares their performance. To quantify the effects of detected shift data, we train a multi-class classifier that learns internal domain knowledge and evaluates the classification performance for each class in external domains after dropping the shift data. We also propose a data quality metric to quantify the dissimilarity between internal and external datasets. We verify the efficacy of MedShift using musculoskeletal radiographs (MURA) and chest X-ray datasets from multiple external sources. Our experiments show that our proposed shift data detection pipeline can be beneficial for medical centers to curate high-quality datasets more efficiently. Xiaoyuan Guo, Judy Gichoya, Hari Trivedi, Saptarshi Purkayastha, Imon Banerjee |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | OSCARS: An Outlier-Sensitive Content-Based Radiography Retrieval SystemabstractImproving the retrieval relevance on noisy datasets is an emerging need for the curation of a large-scale clean dataset in the medical domain. While existing methods can be applied for class-wise retrieval (aka. inter-class), they cannot distinguish the granularity of likeness within the same class (aka. intra-class). The problem is exacerbated on medical external datasets, where noisy samples of the same class are treated equally during training. Our goal is to identify both intra/inter-class similarities for fine-grained retrieval. To achieve this, we propose an Outlier-Sensitive Content-based rAdiologhy Retrieval System (OSCARS), consisting of two steps. First, we train an outlier detector on a clean internal dataset in an unsupervised manner. Then we use the trained detector to generate the anomaly scores on the external dataset, whose distribution will be used to bin intra-class variations. Second, we propose a quadruplet (a, p, nintra, ninter) sampling strategy, where intra-class negatives nintra are sampled from bins of the same class other than the bin anchor a belongs to, while n_inter are randomly sampled from inter-classes. We suggest a weighted metric learning objective to balance the intra and inter-class feature learning. We experimented on two representative public radiography datasets. Experiments show the effectiveness of our approach. The training and evaluation code can be found in https://github.com/XiaoyuanGuo/oscars. Xiaoyuan Guo, Jiali Duan, Saptarshi Purkayastha, Hari Trivedi, Judy Gichoya, Imon Banerjee |
ICMR | 3 |
| 2021 | Predicting Opioid Prescriptions based on Patient Demographics in MIMIC-IVabstractOpioids are widely used analgesics because of their efficacy, mild sedative and anxiolytic properties, and flexibility to administer through multiple routes. Understanding the demographics of the patients receiving these medications helps provide customized care for the susceptible group of people. We conducted a demographic evaluation of the frequently prescribed opioid drug prescriptions from the MIMIC IV database. We analyzed prescribing patterns of six commonly used opioids with demographics such as age, gender, ethnicity, marital status, and year predominantly. After conducting exploratory data analysis, we built models using Logistic Regression, Random Forest, and XGBoost to predict opioid prescriptions and demographics for those. We also analyzed the association between demographics and the frequency of prescribed medications for pain management. We found statistically significant differences in opioid prescriptions among the male and female population, married and unmarried, various ages, ethnic groups, and an association with in-hospital deaths. Snigdha Kodela, Jahnavi Pinnamraju, Judy Gichoya, Saptarshi Purkayastha |
CBMS | 4 |
| 2021 | Using Machine Learning Approaches to Identify Exercise Activities from a Triple-Synchronous Biomedical Sensor
Yohan Mahajan, Jahnavi Pinnamraju, John L. Burns, Judy Gichoya, Saptarshi Purkayastha |
ISDA | 5 |
| 2018 | Finding the Best EHR Training Dataset: A Cross-over Trial
Josette F. Jones, Rukhaiya Fatima, Saptarshi Purkayastha, Anand Kulanthaivel, Enming Zhang |
AMIA | 3 |
| 2017 | Overcoming the Maternal Care Crisis: How Can Lessons Learnt in Global Health Informatics Address US Maternal Health Outcomes?
Suranga Nath Kasthurirathne, Burke W. Mamlin, Saptarshi Purkayastha, Theresa A. Cullen |
AMIA | 3 |
| 2015 | Designing a drawing-based tool to manage EBRT process in an open-source oncology EMR system
Manika Maheshwari, Saptarshi Purkayastha |
AMIA | 2 |