EDBT 2026 Demo / reviewers in the wild / expert
Yindalon Aphinyanagphongs
dblp:96/7338 · also Yin Aphinyanaphongs, Yindalon Aphinyanaphongs
· DBLP profile ↗
25ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0001-8605-5392ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Health system-wide access to generative artificial intelligence: the New York University Langone Health experienceabstractOBJECTIVES: The study aimed to assess the usage and impact of a private and secure instance of a generative artificial intelligence (GenAI) application in a large academic health center. The goal was to understand how employees interact with this technology and the influence on their perception of skill and work performance. MATERIALS AND METHODS: New York University Langone Health (NYULH) established a secure, private, and managed Azure OpenAI service (GenAI Studio) and granted widespread access to employees. Usage was monitored and users were surveyed about their experiences. RESULTS: Over 6 months, over 1007 individuals applied for access, with high usage among research and clinical departments. Users felt prepared to use the GenAI studio, found it easy to use, and would recommend it to a colleague. Users employed the GenAI studio for diverse tasks such as writing, editing, summarizing, data analysis, and idea generation. Challenges included difficulties in educating the workforce in constructing effective prompts and token and API limitations. DISCUSSION: The study demonstrated high interest in and extensive use of GenAI in a healthcare setting, with users employing the technology for diverse tasks. While users identified several challenges, they also recognized the potential of GenAI and indicated a need for more instruction and guidance on effective usage. CONCLUSION: The private GenAI studio provided a useful tool for employees to augment their skills and apply GenAI to their daily tasks. The study underscored the importance of workforce education when implementing system-wide GenAI and provided insights into its strengths and weaknesses. Kiran Malhotra, Batia Mishan Wiesenfeld, Vincent J. Major, Himanshu Grover, Yindalon Aphinyanagphongs, Paul A. Testa, Jonathan S. Austrian |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | Evaluation of GPT-4 ability to identify and generate patient instructions for actionable incidental radiology findingsabstractOBJECTIVES: To evaluate the proficiency of a HIPAA-compliant version of GPT-4 in identifying actionable, incidental findings from unstructured radiology reports of Emergency Department patients. To assess appropriateness of artificial intelligence (AI)-generated, patient-facing summaries of these findings. MATERIALS AND METHODS: Radiology reports extracted from the electronic health record of a large academic medical center were manually reviewed to identify non-emergent, incidental findings with high likelihood of requiring follow-up, further sub-stratified as "definitely actionable" (DA) or "possibly actionable-clinical correlation" (PA-CC). Instruction prompts to GPT-4 were developed and iteratively optimized using a validation set of 50 reports. The optimized prompt was then applied to a test set of 430 unseen reports. GPT-4 performance was primarily graded on accuracy identifying either DA or PA-CC findings, then secondarily for DA findings alone. Outputs were reviewed for hallucinations. AI-generated patient-facing summaries were assessed for appropriateness via Likert scale. RESULTS: For the primary outcome (DA or PA-CC), GPT-4 achieved 99.3% recall, 73.6% precision, and 84.5% F-1. For the secondary outcome (DA only), GPT-4 demonstrated 95.2% recall, 77.3% precision, and 85.3% F-1. No findings were "hallucinated" outright. However, 2.8% of cases included generated text about recommendations that were inferred without specific reference. The majority of True Positive AI-generated summaries required no or minor revision. CONCLUSION: GPT-4 demonstrates proficiency in detecting actionable, incidental findings after refined instruction prompting. AI-generated patient instructions were most often appropriate, but rarely included inferred recommendations. While this technology shows promise to augment diagnostics, active clinician oversight via "human-in-the-loop" workflows remains critical for clinical implementation. Kar-mun C. Woo, Gregory W. Simon, Olumide Akindutire, Yindalon Aphinyanagphongs, Jonathan S. Austrian, Jung G. Kim, Nicholas Genes, Jacob A Goldenring, Vincent J. Major, Chloé S. Pariente, Edwin G. Pineda, Stella K. Kang |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Quantitative and Qualitative Evaluation of Provider Use of a Novel Machine Learning Model for Favorable Outcome Prediction
Elisabeth Yang, Yindalon Aphinyanagphongs, Paawan Punjabi, Jonathan S. Austrian, Batia Mishan Wiesenfeld |
AMIA | 2 |
| 2022 | Automated interpretable discovery of heterogeneous treatment effectiveness: A COVID-19 case study
Benjamin J. Lengerich, Mark E. Nunnally, Yindalon Aphinyanagphongs, Caleb Ellington, Rich Caruana |
J. Biomed. Informatics | 3 |
| 2021 | Have We Learned to Explain?: How Interpretability Methods Can Learn to Encode Predictions in their InterpretationsabstractWhile the need for interpretable machine learning has been established, many common approaches are slow, lack fidelity, or hard to evaluate. Amortized explanation methods reduce the cost of providing interpretations by learning a global selector model that returns feature importances for a single instance of data. The selector model is trained to optimize the fidelity of the interpretations, as evaluated by a predictor model for the target. Popular methods learn the selector and predictor model in concert, which we show allows predictions to be encoded within interpretations. We introduce EVAL-X as a method to quantitatively evaluate interpretations and REAL-X as an amortized explanation method, which learn a predictor model that approximates the true data generating distribution given any subset of the input. We show EVAL-X can detect when predictions are encoded in interpretations and show the advantages of REAL-X through quantitative and radiologist evaluation. Neil Jethani, Mukund Sudarshan, Yindalon Aphinyanagphongs, Rajesh Ranganath |
AISTATS | 3 |
| 2021 | Data-Driven Patterns in Protective Effects of Ibuprofen and Ketorolac on Hospitalized Covid-19 Patients
Rich Caruana, Benjamin J. Lengerich, Yindalon Aphinyanagphongs |
AMIA | 3 |
| 2021 | Prediction of Resuscitation for Pediatric Sepsis from Data Available at Triage
Peter A. Stella, Elizabeth Haines, Yindalon Aphinyanagphongs |
AMIA | 3 |
| 2019 | Translating, Implementing, Deploying, and Evaluating Clinical Interventions Using Machine Learning Based Predictive Models: Illustrative Case Studies
Yindalon Aphinyanagphongs, Jonathan K. Wilt, Corey Chivers, Mark P. Sendak |
AMIA | 1 |
| 2019 | The Process of Developing, Validating and Operationalizing a Personalized Machine Learning Algorithm for Clinical Decision Support: A Case Study
Sara Kuppin Chokshi, Roshini Hegde, Eduardo Iturrate, Yindalon Aphinyanagphongs, Devin M. Mann |
AMIA | 5 |
| 2018 | Utility of General and Specific Word Embeddings for Classifying Translational Stages of Research
Vincent J. Major, Alisa Surkis, Yindalon Aphinyanagphongs |
AMIA | 3 |
| 2017 | The Good, The Bad, and The Ugly of Deploying and Adopting Machine Learning Based Models in Clinical Practice
Yindalon Aphinyanagphongs, David Holmes, Berkman Sahiner, Parsa Mirhaji, Michael Draugelis |
AMIA | 1 |
| 2017 | Optimizing Parameters of word2vec for a Text Classification Task
Vincent J. Major, Alisa Surkis, Yindalon Aphinyanagphongs |
AMIA | 3 |
| 2017 | Neural Network Word Embeddings for Text Classification of MeSH terms
Walter Wang, Yindalon Aphinyanagphongs |
AMIA | 2 |
| 2016 | Simplifying PheWAS Analysis: An R Package and Teaching Module
Emily Kawaler, Yindalon Aphinyanagphongs |
AMIA | 2 |
| 2016 | Reusable Filtering Functions for Application in ICU data: a case study
Vincent J. Major, Monique S. Tanna, Simon Jones 0002, Yindalon Aphinyanagphongs |
AMIA | 4 |
| 2015 | Building and evaluating predictive models for postoperative ileus prior to colorectal surgery
Xuya Wang, Raul Caso Caso, Yindalon Aphinyanagphongs |
AMIA | 3 |
| 2014 | A comprehensive empirical comparison of modern supervised classification and feature selection methods for text categorizationabstractAn important aspect to performing text categorization is selecting appropriate supervised classification and feature selection methods. A comprehensive benchmark is needed to inform best practices in this broad application field. Previous benchmarks have evaluated performance for a few supervised classification and feature selection methods and limited ways to optimize them. The present work updates prior benchmarks by increasing the number of classifiers and feature selection methods order of magnitude, including adding recently developed, state‐of‐the‐art methods. Specifically, this study used 229 text categorization data sets/tasks, and evaluated 28 classification methods (both well‐established and proprietary/commercial) and 19 feature selection methods according to 4 classification performance metrics. We report several key findings that will be helpful in establishing best methodological practices for text categorization. Yindalon Aphinyanagphongs, Lawrence D. Fu, Eric R. Peskin, Efstratios Efstathiadis, Constantin F. Aliferis, Alexander R. Statnikov |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2013 | Feasibility of Machine Learning Based Automatic Classification of Medical School Curricula
Bisakha Ray, Lawrence D. Fu, William Holloway, Yindalon Aphinyanagphongs |
AMIA | 4 |
| 2012 | Using Knowledge from APIs to Disambiguate Affiliation Names in MEDLINE
Karen Hanson, Yindalon Aphinyanagphongs, Lawrence D. Fu |
AMIA | 2 |
| 2011 | A comparison of evaluation metrics for biomedical journals, articles, and websites in terms of sensitivity to topic
Lawrence D. Fu, Yindalon Aphinyanagphongs, Lily Wang 0001, Constantin F. Aliferis |
J. Biomed. Informatics | 2 |
| 2006 | Prospective Validation of Text Categorization Filters for Identifying High-Quality, Content-Specific Articles in MEDLINE
Yindalon Aphinyanagphongs, Constantin F. Aliferis |
AMIA | 1 |
| 2006 | Research Paper: A Comparison of Citation Metrics to Machine Learning Filters for the Identification of High Quality MEDLINE DocumentsabstractOBJECTIVE: The present study explores the discriminatory performance of existing and novel gold-standard-specific machine learning (GSS-ML) focused filter models (i.e., models built specifically for a retrieval task and a gold standard against which they are evaluated) and compares their performance to citation count and impact factors, and non-specific machine learning (NS-ML) models (i.e., models built for a different task and/or different gold standard). DESIGN: Three gold standard corpora were constructed using the SSOAB bibliography, the ACPJ-cited treatment articles, and the ACPJ-cited etiology articles. Citation counts and impact factors were obtained for each article. Support vector machine models were used to classify the articles using combinations of content, impact factors, and citation counts as predictors. MEASUREMENTS: Discriminatory performance was estimated using the area under the receiver operating characteristic curve and n-fold cross-validation. RESULTS: For all three gold standards and tasks, GSS-ML filters outperformed citation count, impact factors, and NS-ML filters. Combinations of content with impact factor or citation count produced no or negligible improvements to the GSS machine learning filters. CONCLUSIONS: These experiments provide evidence that when building information retrieval filters focused on a retrieval task and corresponding gold standard, the filter models have to be built specifically for this task and gold standard. Under those conditions, machine learning filters outperform standard citation metrics. Furthermore, citation counts and impact factors add marginal value to discriminatory performance. Previous research that claimed better performance of citation metrics than machine learning in one of the corpora examined here is attributed to using machine learning filters built for a different gold standard and task. Yindalon Aphinyanagphongs, Alexander R. Statnikov, Constantin F. Aliferis |
J. Am. Medical Informatics Assoc. | 1 |
| 2006 | Research Paper: Using Citation Data to Improve Retrieval from MEDLINEabstractOBJECTIVE: To determine whether algorithms developed for the World Wide Web can be applied to the biomedical literature in order to identify articles that are important as well as relevant. DESIGN AND MEASUREMENTS A direct comparison of eight algorithms: simple PubMed queries, clinical queries (sensitive and specific versions), vector cosine comparison, citation count, journal impact factor, PageRank, and machine learning based on polynomial support vector machines. The objective was to prioritize important articles, defined as being included in a pre-existing bibliography of important literature in surgical oncology. RESULTS Citation-based algorithms were more effective than noncitation-based algorithms at identifying important articles. The most effective strategies were simple citation count and PageRank, which on average identified over six important articles in the first 100 results compared to 0.85 for the best noncitation-based algorithm (p < 0.001). The authors saw similar differences between citation-based and noncitation-based algorithms at 10, 20, 50, 200, 500, and 1,000 results (p < 0.001). Citation lag affects performance of PageRank more than simple citation count. However, in spite of citation lag, citation-based algorithms remain more effective than noncitation-based algorithms. CONCLUSION Algorithms that have proved successful on the World Wide Web can be applied to biomedical information retrieval. Citation-based algorithms can help identify important articles within large sets of relevant results. Further studies are needed to determine whether citation-based algorithms can effectively meet actual user information needs. Elmer V. Bernstam, Jorge R. Herskovic, Yindalon Aphinyanagphongs, Constantin F. Aliferis, Madurai G. Sriram, William R. Hersh |
J. Am. Medical Informatics Assoc. | 3 |
| 2005 | Research Paper: Text Categorization Models for High-Quality Article Retrieval in Internal MedicineabstractOBJECTIVE Finding the best scientific evidence that applies to a patient problem is becoming exceedingly difficult due to the exponential growth of medical publications. The objective of this study was to apply machine learning techniques to automatically identify high-quality, content-specific articles for one time period in internal medicine and compare their performance with previous Boolean-based PubMed clinical query filters of Haynes et al. DESIGN The selection criteria of the ACP Journal Club for articles in internal medicine were the basis for identifying high-quality articles in the areas of etiology, prognosis, diagnosis, and treatment. Naive Bayes, a specialized AdaBoost algorithm, and linear and polynomial support vector machines were applied to identify these articles. MEASUREMENTS The machine learning models were compared in each category with each other and with the clinical query filters using area under the receiver operating characteristic curves, 11-point average recall precision, and a sensitivity/specificity match method. RESULTS In most categories, the data-induced models have better or comparable sensitivity, specificity, and precision than the clinical query filters. The polynomial support vector machine models perform the best among all learning methods in ranking the articles as evaluated by area under the receiver operating curve and 11-point average recall precision. CONCLUSION This research shows that, using machine learning methods, it is possible to automatically build models for retrieving high-quality, content-specific articles using inclusion or citation by the ACP Journal Club as a gold standard in a given time period in internal medicine that perform better than the 1994 PubMed clinical query filters. Yindalon Aphinyanagphongs, Ioannis Tsamardinos, Alexander R. Statnikov, Douglas P. Hardin, Constantin F. Aliferis |
J. Am. Medical Informatics Assoc. | 1 |
| 2003 | Text Categorization Models for Retrieval of High Quality Articles in Internal Medicine
Yindalon Aphinyanagphongs, Constantin F. Aliferis |
AMIA | 1 |