Judy Gichoya

dblp:134/2353 · also Judy W. Gichoya, Judy Wawira Gichoya · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0002-1097-316XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Principles and implementation strategies for equitable and representative academic partnerships in global health informatics research
abstract
OBJECTIVE: Developing equitable, sustainable informatics solutions is key to scalability and long-term success for projects in the global health informatics (GHI) domain. This paper presents key strategies for incorporating principles of health equity in the GHI project lifecycle. MATERIALS AND METHODS: The American Medical Informatics Association (AMIA) GHI Working Group organized a collaborative workshop at the 2023 AMIA Annual Symposium that included the presentation of five case studies of how principles of health equity have been incorporated into projects situated in low-and-middle-income countries and with Indigenous communities in the U.S. and best practices for operationalizing these principles into other informatics projects. RESULTS: We present five principles: (1) Inclusion and Participation in Ethical, Sustainable Collaborations; (2) Engaging Community-Based Participatory Research Approaches; (3) Stakeholder Engagement; (4) Scalability and Sustainability; (5) Representation in Knowledge Creation, along with strategies that informatics researchers may use to incorporate these principles into their work. DISCUSSION: Presented case studies and subsequent focus groups yielded key concepts and strategies to promote health equity that may be operationalized across GHI projects. CONCLUSION: Equitable, sustainable, and scalable GHI projects require intentional integration of community and stakeholder perspectives in project development, implementation, and knowledge creation processes.
Elizabeth A. Campbell, Oliver J. Bear Don't Walk IV, Hamish S. F. Fraser, Judy Gichoya, Kavishwar B. Wagholikar, Andrew S. Kanter, Felix Holl, Sansanee Craig
J. Am. Medical Informatics Assoc.4
2025 Adaptable graph neural networks design to support generalizability for clinical event prediction
Amara Tariq, Gurkiran Kaur, Leon Su, Judy Gichoya, Bhavik N. Patel, Imon Banerjee
J. Biomed. Informatics4
2024 Efficient adversarial debiasing with concept activation vector - Medical image case-studies
Ramon Correa, Khushbu Pahwa, Bhavik N. Patel, Celine M. Vachon, Judy Gichoya, Imon Banerjee
J. Biomed. Informatics5
2023 A General-Purpose AI Assistant Embedded in an Open-Source Radiology Information System
Saptarshi Purkayastha, Rohan Satya Isaac, Sharon Anthony, Shikhar Shukla, Elizabeth A. Krupinski, Joshua A. Danish, Judy Gichoya
AIME7
2023 MedShift: Automated Identification of Shift Data for Medical Image Dataset Curation
abstract
Automated curation of noisy external data in the medical domain has long been in high demand, as AI technologies need to be validated using various sources with clean, annotated data. Identifying the variance between internal and external sources is a fundamental step in curating a high-quality dataset, as the data distributions from different sources can vary significantly and subsequently affect the performance of AI models. The primary challenges for detecting data shifts are - (1) accessing private data across healthcare institutions for manual detection and (2) the lack of automated approaches to learn efficient shift-data representation without training samples. To overcome these problems, we propose an automated pipeline called MedShift to detect top-level shift samples and evaluate the significance of shift data without sharing data between internal and external organizations. MedShift employs unsupervised anomaly detectors to learn the internal distribution and identify samples showing significant shiftness for external datasets, and then compares their performance. To quantify the effects of detected shift data, we train a multi-class classifier that learns internal domain knowledge and evaluates the classification performance for each class in external domains after dropping the shift data. We also propose a data quality metric to quantify the dissimilarity between internal and external datasets. We verify the efficacy of MedShift using musculoskeletal radiographs (MURA) and chest X-ray datasets from multiple external sources. Our experiments show that our proposed shift data detection pipeline can be beneficial for medical centers to curate high-quality datasets more efficiently.
Xiaoyuan Guo, Judy Gichoya, Hari Trivedi, Saptarshi Purkayastha, Imon Banerjee
IEEE J. Biomed. Health Informatics2
2022 Graph-based Fusion Modeling and Explanation for Disease Trajectory Prediction
Amara Tariq, Siyi Tang, Hifza Sakhi, Leo A. Celi, Janice M. Newsome, Daniel L. Rubin, Hari Trivedi, Judy Gichoya, Bhavik N. Patel, Imon Banerjee
AMIA8
2022 Augmenting Vision Language Pretraining by Learning Codebook with Visual Semantics
abstract
Language modality within the vision language pre-training framework is innately discretized, endowing each word in the language vocabulary a semantic meaning. In contrast, visual modality is inherently continuous and high-dimensional, which potentially prohibits the alignment as well as fusion between vision and language modalities. We therefore propose to "discretize" the visual representation by joint learning a codebook that imbues each visual token a semantic. We then utilize these discretized visual semantics as self-supervised ground-truths for building our Masked Image Modeling objective, a counterpart of Masked Language Modeling which proves successful for language models. To optimize the codebook, we extend the formulation of VQ-VAE which gives a theoretic guarantee. Experiments validate the effectiveness of our approach across common vision-language benchmarks.
Xiaoyuan Guo, Jiali Duan, C.-C. Jay Kuo, Judy Gichoya, Imon Banerjee
ICPR4
2022 OSCARS: An Outlier-Sensitive Content-Based Radiography Retrieval System
abstract
Improving the retrieval relevance on noisy datasets is an emerging need for the curation of a large-scale clean dataset in the medical domain. While existing methods can be applied for class-wise retrieval (aka. inter-class), they cannot distinguish the granularity of likeness within the same class (aka. intra-class). The problem is exacerbated on medical external datasets, where noisy samples of the same class are treated equally during training. Our goal is to identify both intra/inter-class similarities for fine-grained retrieval. To achieve this, we propose an Outlier-Sensitive Content-based rAdiologhy Retrieval System (OSCARS), consisting of two steps. First, we train an outlier detector on a clean internal dataset in an unsupervised manner. Then we use the trained detector to generate the anomaly scores on the external dataset, whose distribution will be used to bin intra-class variations. Second, we propose a quadruplet (a, p, nintra, ninter) sampling strategy, where intra-class negatives nintra are sampled from bins of the same class other than the bin anchor a belongs to, while n_inter are randomly sampled from inter-classes. We suggest a weighted metric learning objective to balance the intra and inter-class feature learning. We experimented on two representative public radiography datasets. Experiments show the effectiveness of our approach. The training and evaluation code can be found in https://github.com/XiaoyuanGuo/oscars.
Xiaoyuan Guo, Jiali Duan, Saptarshi Purkayastha, Hari Trivedi, Judy Gichoya, Imon Banerjee
ICMR5
2022 Towards an internet-scale overlay network for latency-aware decentralized workflows at the edge
abstract
Small-scale data centers at the edge are becoming prominent in offering various services to the end-users following the cloud model while avoiding the high latency inherent to the classic cloud environments when accessed from remote Internet regions. However, we should address several challenges to facilitate the end-users finding and consuming the relevant services from the edge at the Internet scale. First, the scale and diversity of the edge hinder seamless access. Second, a framework where researchers openly share their services and data in a secured manner among themselves and with external consumers over the Internet does not exist. Third, the lack of a unified interface and trust across the service providers hinder their interchangeability in composing workflows by chaining the services. Thus, creating a workflow from the services deployed on the various edge nodes is presently impractical. This paper designs Viseu, a latency-aware blockchain framework to provide Virtual Internet Services at the Edge. Viseu aims to solve the puzzle of network service discovery at the edge, considering the peers' reputation and latency when choosing the service instances. Viseu enables peers to share their computational resources, services, and data among each other in an untrusted environment, rather than relying on a set of trusted service providers. By composing workflows from the peers' services, rather than confining them to the pre-established service provider and consumer roles, Viseu aims to facilitate scientific collaboration across the peers natively. Furthermore, by offering services from multiple peers close to the end-users, Viseu also minimizes end-to-end latency and data loss in the service execution at the Internet scale.
Pradeeban Kathiravelu, Zach Zaiman, Judy Gichoya, Luís Veiga, Imon Banerjee
Comput. Networks3
2022 Evaluation of federated learning variations for COVID-19 diagnosis using chest radiographs from 42 US and European hospitals
abstract
OBJECTIVE: Federated learning (FL) allows multiple distributed data holders to collaboratively learn a shared model without data sharing. However, individual health system data are heterogeneous. "Personalized" FL variations have been developed to counter data heterogeneity, but few have been evaluated using real-world healthcare data. The purpose of this study is to investigate the performance of a single-site versus a 3-client federated model using a previously described Coronavirus Disease 19 (COVID-19) diagnostic model. Additionally, to investigate the effect of system heterogeneity, we evaluate the performance of 4 FL variations. MATERIALS AND METHODS: We leverage a FL healthcare collaborative including data from 5 international healthcare systems (US and Europe) encompassing 42 hospitals. We implemented a COVID-19 computer vision diagnosis system using the Federated Averaging (FedAvg) algorithm implemented on Clara Train SDK 4.0. To study the effect of data heterogeneity, training data was pooled from 3 systems locally and federation was simulated. We compared a centralized/pooled model, versus FedAvg, and 3 personalized FL variations (FedProx, FedBN, and FedAMP). RESULTS: We observed comparable model performance with respect to internal validation (local model: AUROC 0.94 vs FedAvg: 0.95, P = .5) and improved model generalizability with the FedAvg model (P < .05). When investigating the effects of model heterogeneity, we observed poor performance with FedAvg on internal validation as compared to personalized FL algorithms. FedAvg did have improved generalizability compared to personalized FL algorithms. On average, FedBN had the best rank performance on internal and external validation. CONCLUSION: FedAvg can significantly improve the generalization of the model compared to other personalization FL algorithms; however, at the cost of poor internal validity. Personalized FL may offer an opportunity to develop both internal and externally validated algorithms.
Le Peng, Gaoxiang Luo, Andrew Walker, Zach Zaiman, Emma K. Jones, Hemant Gupta, Kristopher Kersten, John L. Burns, Christopher A. Harle, Tanja Magoc, Benjamin Shickel, Scott D. Steenburg, Tyler J. Loftus, Genevieve B. Melton, Judy Gichoya, Ju Sun, Christopher J. Tignanelli
J. Am. Medical Informatics Assoc.15
2021 A Fusion NLP Model for the Inference of Standardized Thyroid Nodule Malignancy Scores from Radiology Report Text
Thiago Santos, Omar Kallas, Janice M. Newsome, Daniel L. Rubin, Judy Gichoya, Imon Banerjee
AMIA5
2021 Predicting Opioid Prescriptions based on Patient Demographics in MIMIC-IV
abstract
Opioids are widely used analgesics because of their efficacy, mild sedative and anxiolytic properties, and flexibility to administer through multiple routes. Understanding the demographics of the patients receiving these medications helps provide customized care for the susceptible group of people. We conducted a demographic evaluation of the frequently prescribed opioid drug prescriptions from the MIMIC IV database. We analyzed prescribing patterns of six commonly used opioids with demographics such as age, gender, ethnicity, marital status, and year predominantly. After conducting exploratory data analysis, we built models using Logistic Regression, Random Forest, and XGBoost to predict opioid prescriptions and demographics for those. We also analyzed the association between demographics and the frequency of prescribed medications for pain management. We found statistically significant differences in opioid prescriptions among the male and female population, married and unmarried, various ages, ethnic groups, and an association with in-hospital deaths.
Snigdha Kodela, Jahnavi Pinnamraju, Judy Gichoya, Saptarshi Purkayastha
CBMS3
2021 Using Machine Learning Approaches to Identify Exercise Activities from a Triple-Synchronous Biomedical Sensor
Yohan Mahajan, Jahnavi Pinnamraju, John L. Burns, Judy Gichoya, Saptarshi Purkayastha
ISDA4
2021 Query bot for retrieving patients' clinical history: A COVID-19 use-case
Yibo Wang 0001, Amara Tariq, Fiza Khan, Judy Gichoya, Hari Trivedi, Imon Banerjee
J. Biomed. Informatics4
2017 Toward better public health reporting using existing off the shelf approaches: The value of medical dictionaries in automated cancer detection using plaintext medical data
Suranga Nath Kasthurirathne, Brian E. Dixon, Judy Gichoya, Huiping Xu, Yuni Xia, Burke W. Mamlin, Shaun J. Grannis
J. Biomed. Informatics3
2016 The OpenMRS Community's Experience: A Decade of Developing and Implementing Medical Record Systems Within Constraint
Theresa A. Cullen, Paul G. Biondich, Burke W. Mamlin, Judy Gichoya, Hamish S. F. Fraser
AMIA4
2016 Toward better public health reporting using existing off the shelf approaches: A comparison of alternative cancer detection approaches using plaintext medical data and non-dictionary based feature selection
Suranga Nath Kasthurirathne, Brian E. Dixon, Judy Gichoya, Huiping Xu, Yuni Xia, Burke W. Mamlin, Shaun J. Grannis
J. Biomed. Informatics3
2012 An Evaluation of the Rates of Repeat Notifiable Disease Reporting and Patient Crossover Using a Health Information Exchange-based Automated Electronic Laboratory Reporting System
Judy Gichoya, Brian E. Dixon, John T. Finnell, Daniel J. Vreeman, Roland E. Gamache, Shaun J. Grannis
AMIA1