EDBT 2026 Demo / reviewers in the wild / expert
Ameen Abu-Hanna
dblp:31/59
· DBLP profile ↗
41ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-4324-7954ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainability of decoder-only clinical large language models: A scoping reviewabstractClinical large language models (LLMs) are increasingly used for documentation, diagnosis, and decision support, but their opaque reasoning can limit clinician trust, regulatory assessment, and safe deployment. Explainability research has expanded rapidly, yet existing reviews largely address traditional machine learning or general-domain LLMs. We conducted a PRISMA-ScR scoping review to map explainability approaches for decoder-only clinical LLMs with over one billion parameters, searching PubMed, Scopus, Web of Science, ACM Digital Library, and arXiv through early 2026. Among 69 included studies, LLM-native generative and interactive methods dominated (58.0%, n = 40), spanning chain-of-thought rationales, retrieval-augmented evidence citation, and agentic decomposition. Intrinsic by-design methods accounted for 24.6% (n = 17); post-hoc XAI methods accounted for 17.4% (n = 12). General medicine was the most represented clinical domain, and diagnosis was the dominant task. Proprietary models were used in 75.4% of studies, yet every mechanistic analysis relied on open-source models, revealing a transparency asymmetry: most deployed models are the least transparent. Although 59.4% of studies quantitatively evaluated explanations, metrics remain non-standardized and rarely assess faithfulness. Local explanations predominated, and no study prospectively evaluated explanations in live clinical workflows. These findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited. To support trustworthy deployment, we highlight three regulatory priorities: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling. Nishant Mishra 0001, Ameen Abu-Hanna, Iacer Calixto |
Artif. Intell. Medicine | 2 |
| 2024 | Mitigating Overconfidence in Out-of-Distribution Detection by Capturing Extreme ActivationsabstractDetecting out-of-distribution (OOD) instances is crucial for the reliable deployment of machine learning models in real-world scenarios. OOD inputs are commonly expected to cause a more uncertain prediction in the primary task; however, there are OOD cases for which the model returns a highly confident prediction. This phenomenon, denoted as "overconfidence", presents a challenge to OOD detection. Specifically, theoretical evidence indicates that overconfidence is an intrinsic property of certain neural network architectures, leading to poor OOD detection. In this work, we address this issue by measuring extreme activation values in the penultimate layer of neural networks and then leverage this proxy of overconfidence to improve on several OOD detection baselines. We test our method on a wide array of experiments spanning synthetic data and real-world data, tabular and image datasets, multiple architectures such as ResNet and Transformer, different training loss functions, and include the scenarios examined in previous theoretical work. Compared to the baselines, our method often grants substantial improvements, with double-digit increases in OOD detection AUC, and it does not damage performance in any scenario. Mohammad Azizmalayeri, Ameen Abu-Hanna, Giovanni Cinà |
UAI | 2 |
| 2023 | Soft-Prompt Tuning to Predict Lung Cancer Using Primary Care Free-Text Dutch Medical Notes
Auke Elfrink, Iacopo Vagliano, Ameen Abu-Hanna, Iacer Calixto |
AIME | 3 |
| 2023 | Autoencoder-Based Prediction of ICU Clinical Codes
Tsvetan R. Yordanov, Ameen Abu-Hanna, Anita C. J. Ravelli, Iacopo Vagliano |
AIME | 2 |
| 2023 | Electronic health record-based prediction models for in-hospital adverse drug event diagnosis or prognosis: a systematic reviewabstractOBJECTIVE: We conducted a systematic review to characterize and critically appraise developed prediction models based on structured electronic health record (EHR) data for adverse drug event (ADE) diagnosis and prognosis in adult hospitalized patients. MATERIALS AND METHODS: We searched the Embase and Medline databases (from January 1, 1999, to July 4, 2022) for articles utilizing structured EHR data to develop ADE prediction models for adult inpatients. For our systematic evidence synthesis and critical appraisal, we applied the Checklist for Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHARMS). RESULTS: Twenty-five articles were included. Studies often did not report crucial information such as patient characteristics or the method for handling missing data. In addition, studies frequently applied inappropriate methods, such as univariable screening for predictor selection. Furthermore, the majority of the studies utilized ADE labels that only described an adverse symptom while not assessing causality or utilizing a causal model. None of the models were externally validated. CONCLUSIONS: Several challenges should be addressed before the models can be widely implemented, including the adherence to reporting standards and the adoption of best practice methods for model development and validation. In addition, we propose a reorientation of the ADE prediction modeling domain to include causality as a fundamental challenge that needs to be addressed in future studies, either through acquiring ADE labels via formal causality assessments or the usage of adverse event labels in combination with causal prediction modeling. Izak A. R. Yasrebi-de Kom, Dave Dongelmans, Nicolette de Keizer, Kitty J. Jager, Martijn C. Schut, Ameen Abu-Hanna, Joanna E. Klopotowska |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Prognostic models of in-hospital mortality of intensive care patients using neural representation of unstructured text: A systematic review and critical appraisalabstractOBJECTIVE: To review and critically appraise published and preprint reports of prognostic models of in-hospital mortality of patients in the intensive-care unit (ICU) based on neural representations (embeddings) of clinical notes. METHODS: PubMed and arXiv were searched up to August 1, 2022. At least two reviewers independently selected the studies that developed a prognostic model of in-hospital mortality of intensive-care patients using free-text represented as embeddings and extracted data using the CHARMS checklist. Risk of bias was assessed using PROBAST. Reporting on the model was assessed with the TRIPOD guideline. To assess the machine learning components that were used in the models, we present a new descriptive framework based on different techniques to represent text and provide predictions from text. The study protocol was registered in the PROSPERO database (CRD42022354602). RESULTS: Eighteen studies out of 2,825 were included. All studies used the publicly-available MIMIC dataset. Context-independent word embeddings are widely used. Model discrimination was provided by all studies (AUROC 0.75-0.96), but measures of calibration were scarce. Seven studies used both structural clinical variables and notes. Model discrimination improved when adding clinical notes to variables. None of the models was externally validated and often a simple train/test split was used for internal validation. Our critical appraisal demonstrated a high risk of bias in all studies and concerns regarding their applicability in clinical practice. CONCLUSION: All studies used a neural architecture for prediction and were based on one publicly available dataset. Clinical notes were reported to improve predictive performance when used in addition to only clinical variables. Most studies had methodological, reporting, and applicability issues. We recommend reporting both model discrimination and calibration, using additional data sources, and using more robust evaluation strategies, including prospective and external validation. Finally, sharing data and code is encouraged to improve study reproducibility. Iacopo Vagliano, Noman Dormosh, Miguel Ángel Ríos-Gaona, Torec T. Luik, Tommaso Mario Buonocore, Paul W. G. Elbers, Dave Dongelmans, Martijn C. Schut, Ameen Abu-Hanna |
J. Biomed. Informatics | 9 |
| 2022 | Evaluating pointwise reliability of machine learning predictionabstractInterest in Machine Learning applications to tackle clinical and biological problems is increasing. This is driven by promising results reported in many research papers, the increasing number of AI-based software products, and by the general interest in Artificial Intelligence to solve complex problems. It is therefore of importance to improve the quality of machine learning output and add safeguards to support their adoption. In addition to regulatory and logistical strategies, a crucial aspect is to detect when a Machine Learning model is not able to generalize to new unseen instances, which may originate from a population distant to that of the training population or from an under-represented subpopulation. As a result, the prediction of the machine learning model for these instances may be often wrong, given that the model is applied outside its "reliable" space of work, leading to a decreasing trust of the final users, such as clinicians. For this reason, when a model is deployed in practice, it would be important to advise users when the model's predictions may be unreliable, especially in high-stakes applications, including those in healthcare. Yet, reliability assessment of each machine learning prediction is still poorly addressed. Here, we review approaches that can support the identification of unreliable predictions, we harmonize the notation and terminology of relevant concepts, and we highlight and extend possible interrelationships and overlap among concepts. We then demonstrate, on simulated and real data for ICU in-hospital death prediction, a possible integrative framework for the identification of reliable and unreliable predictions. To do so, our proposed approach implements two complementary principles, namely the density principle and the local fit principle. The density principle verifies that the instance we want to evaluate is similar to the training set. The local fit principle verifies that the trained model performs well on training subsets that are more similar to the instance under evaluation. Our work can contribute to consolidating work in machine learning especially in medicine. Giovanna Nicora, Miguel Ángel Ríos-Gaona, Ameen Abu-Hanna, Riccardo Bellazzi |
J. Biomed. Informatics | 3 |
| 2021 | The Effectiveness of Phrase Skip-Gram in Primary Care NLP for the Prediction of Lung Cancer
Torec T. Luik, Miguel Ángel Ríos-Gaona, Ameen Abu-Hanna, Henk C. P. M. van Weert, Martijn C. Schut |
AIME | 3 |
| 2021 | Deep Kernel Learning for Mortality Prediction in the Face of Temporal Shift
Miguel Ángel Ríos-Gaona, Ameen Abu-Hanna |
AIME | 2 |
| 2021 | Explaining heterogeneity of individual treatment causal effects by subgroup discovery: An observational case study in antibiotics treatment of acute rhino-sinusitisabstractOBJECTIVES: Individuals may respond differently to the same treatment, and there is a need to understand such heterogeneity of causal individual treatment effects. We propose and evaluate a modelling approach to better understand this heterogeneity from observational studies by identifying patient subgroups with a markedly deviating response to treatment. We illustrate this approach in a primary care case-study of antibiotic (AB) prescription on recovery from acute rhino-sinusitis (ARS). METHODS: Our approach consists of four stages and is applied to a large dataset in primary care dataset of 24,392 patients suspected of suffering from ARS. We first identify pre-treatment variables that either confound the relationship between treatment and outcome or are risk factors of the outcome. Second, based on the pre-treatment variables we create Synthetic Random Forest (SRF) models to compute the potential outcomes and subsequently the causal individual treatment effect (ITE) estimates. Third, we perform subgroup discovery using the ITE estimates as outcomes to identify positive and negative responders. Fourth, we evaluate the predictive performance of the identified subgroups for predicting the outcome in two ways: the likelihood ratio test, and whether the subgroups are selected via the Akaike Information Criterion (AIC) using backward stepwise variable selection. We validate the whole modelling strategy by means of 10-fold-cross-validation. RESULTS: Based on 20 pre-treatment variables, four subgroups (three for positive responders and one for negative responders) were identified. The log likelihood ratio tests showed that the subgroups were significant. Variable selection using the AIC kept two of the four subgroups, one for positive responders and one for negative responders. As for the validation of the whole modelling strategy, all reported measures (the number of pre-treatment variables associated with the outcome, number of subgroups, number of subgroups surviving variable selection and coverage) showed little variation. CONCLUSIONS: With the proposed approach, we identified subgroups of positive and negative responders to treatment that markedly deviate from the mean response. The subgroups showed additive predictive value of the outcome. The modelling approach strategy was shown to be robust on this dataset. Our approach was thus able to discover understandable subgroups from observational data that have predictive value and which may be considered by the clinical users to get insight into who responds positively or negatively to a proposed treatment. W. Qi, Ameen Abu-Hanna, Thamar Eva Maria van Esch, Derek de Beurs, Linda E. Flinterman, Martijn C. Schut |
Artif. Intell. Medicine | 2 |
| 2021 | De-novo FAIRification via an Electronic Data Capture system by automated transformation of filled electronic Case Report Forms into machine-readable dataabstractINTRODUCTION: Existing methods to make data Findable, Accessible, Interoperable, and Reusable (FAIR) are usually carried out in a post hoc manner: after the research project is conducted and data are collected. De-novo FAIRification, on the other hand, incorporates the FAIRification steps in the process of a research project. In medical research, data is often collected and stored via electronic Case Report Forms (eCRFs) in Electronic Data Capture (EDC) systems. By implementing a de novo FAIRification process in such a system, the reusability and, thus, scalability of FAIRification across research projects can be greatly improved. In this study, we developed and implemented a novel method for de novo FAIRification via an EDC system. We evaluated our method by applying it to the Registry of Vascular Anomalies (VASCA). METHODS: Our EDC and research project independent method ensures that eCRF data entered into an EDC system can be transformed into machine-readable, FAIR data using a semantic data model (a canonical representation of the data, based on ontology concepts and semantic web standards) and mappings from the model to questions on the eCRF. The FAIRified data are stored in a triple store and can, together with associated metadata, be accessed and queried through a FAIR Data Point. The method was implemented in Castor EDC, an EDC system, through a data transformation application. The FAIRness of the output of the method, the FAIRified data and metadata, was evaluated using the FAIR Evaluation Services. RESULTS: We successfully applied our FAIRification method to the VASCA registry. Data entered on eCRFs is automatically transformed into machine-readable data and can be accessed and queried using SPARQL queries in the FAIR Data Point. Twenty-one FAIR Evaluator tests pass and one test regarding the metadata persistence policy fails, since this policy is not in place yet. CONCLUSION: In this study, we developed a novel method for de novo FAIRification via an EDC system. Its application in the VASCA registry and the automated FAIR evaluation show that the method can be used to make clinical research data FAIR when they are entered in an eCRF without any intervention from data management and data entry personnel. Due to the generic approach and developed tooling, we believe that our method can be used in other registries and clinical trials as well. Martijn G. Kersloot, Annika Jacobsen, Karlijn H. J. Groenen, Bruna dos Santos Vieira, Rajaram Kaliyaperumal, Ameen Abu-Hanna, Ronald Cornet, Peter A. C. 't Hoen, Marco Roos, Leo J. Schultze Kool, Derk L. Arts |
J. Biomed. Informatics | 6 |
| 2016 | Modeling information flows in clinical decision support: key insights for enhancing system effectivenessabstractA fundamental challenge in the field of clinical decision support is to determine what characteristics of systems make them effective in supporting particular types of clinical decisions. However, we lack such a theory of decision support itself and a model to describe clinical decisions and the systems to support them. This article outlines such a framework. We present a two-stream model of information flow within clinical decision-support systems (CDSSs): reasoning about the patient (the clinical stream), and reasoning about the user (the cognitive-behavioral stream). We propose that CDSS "effectiveness" be measured not only in terms of a system's impact on clinical care, but also in terms of how (and by whom) the system is used, its effect on work processes, and whether it facilitates appropriate decisions by clinicians and patients. Future research into which factors improve the effectiveness of decision support should not regard CDSSs as a single entity, but should instead differentiate systems based on their attributes, users, and the decision being supported. Stephanie Medlock, Jeremy C. Wyatt, Vimla L. Patel, Edward H. Shortliffe, Ameen Abu-Hanna |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Comparison of Probabilistic versus Non-probabilistic Electronic Nose Classification Methods in an Animal Model
Camilla Colombo, Jan Hendrik Leopold, Lieuwe D. J. Bos, Riccardo Bellazzi, Ameen Abu-Hanna |
AIME | 5 |
| 2015 | Evaluation of a Self-Triage Decision Aid System in Pregnancy-Induced Hypertension and Diabetes Mellitus: Preliminary Results of a Randomized Control Trial
Azam Aslani, Fatemeh Tara, Leila Galichi, Sina Madani, Ameen Abu-Hanna, Saeid Eslami |
AMIA | 5 |
| 2014 | Electronic medical record systems are associated with appropriate placement of HIV patients on antiretroviral therapy in rural health facilities in Kenya: a retrospective pre-post studyabstractBACKGROUND AND OBJECTIVE: There is little evidence that electronic medical record (EMR) use is associated with better compliance with clinical guidelines on initiation of antiretroviral therapy (ART) among ART-eligible HIV patients. We assessed the effect of transitioning from paper-based to an EMR-based system on appropriate placement on ART among eligible patients. METHODS: We conducted a retrospective, pre-post EMR study among patients enrolled in HIV care and eligible for ART at 17 rural Kenyan clinics and compared the: (1) proportion of patients eligible for ART based on CD4 count or WHO staging who initiate therapy; (2) time from eligibility for ART to ART initiation; (3) time from ART initiation to first CD4 test. RESULTS: 7298 patients were eligible for ART; 54.8% (n=3998) were enrolled in HIV care using a paper-based system while 45.2% (n=3300) were enrolled after the implementation of the EMR. EMR was independently associated with a 22% increase in the odds of initiating ART among eligible patients (adjusted OR (aOR) 1.22, 95% CI 1.12 to 1.33). The proportion of ART-eligible patients not receiving ART was 20.3% and 15.1% for paper and EMR, respectively (χ(2)=33.5, p<0.01). Median time from ART eligibility to ART initiation was 29.1 days (IQR: 14.1-62.1) for paper compared to 27 days (IQR: 12.9-50.1) for EMR. CONCLUSIONS: EMRs can improve quality of HIV care through appropriate placement of ART-eligible patients on treatment in resource limited settings. However, other non-EMR factors influence timely initiation of ART. Tom Oluoch, Abraham Katana, Victor Ssempijja, Daniel Kwaro, Patrick Langat, Davies Kimanga, Nicky Okeyo, Ameen Abu-Hanna, Nicolette de Keizer |
J. Am. Medical Informatics Assoc. | 8 |
| 2013 | A Modified Real AdaBoost algorithm to discover Intensive Care Unit subgroups with a poor outcome
Antonie Koetsier, Nicolette de Keizer, Ameen Abu-Hanna, Niels Peek |
AMIA | 3 |
| 2013 | Assessing and combining repeated prognosis of physicians and temporal models in the intensive care
Lilian Minne, Tudor Toma, Evert de Jonge, Ameen Abu-Hanna |
Artif. Intell. Medicine | 4 |
| 2012 | Statistical process control for validating a classification tree model for predicting mortality - A novel approach towards temporal validation
Lilian Minne, Saeid Eslami, Nicolette de Keizer, Evert de Jonge, Sophia E. de Rooij, Ameen Abu-Hanna |
J. Biomed. Informatics | 6 |
| 2011 | Lessons Learned from Implementing and Evaluating Computerized Decision Support Systems
Saeid Eslami, Nicolette de Keizer, Evert de Jonge, Dave Dongelmans, Marcus J. Schultz, Ameen Abu-Hanna |
AIME | 6 |
| 2011 | Repeated Prognosis in the Intensive Care: How Well Do Physicians and Temporal Models Perform?
Lilian Minne, Evert de Jonge, Ameen Abu-Hanna |
AIME | 3 |
| 2010 | PRIM versus CART in subgroup discovery: When patience is harmful
Ameen Abu-Hanna, Barry Nannings, Dave Dongelmans, Arie Hasman |
J. Biomed. Informatics | 1 |
| 2010 | Using hierarchical dynamic Bayesian networks to investigate dynamics of organ failure in patients in the Intensive Care Unit
Linda Peelen, Nicolette de Keizer, Evert de Jonge, Robert-Jan Bosman, Ameen Abu-Hanna, Niels Peek |
J. Biomed. Informatics | 5 |
| 2010 | Learning predictive models that use pattern discovery - A bootstrap evaluative approach applied in organ functioning sequences
Tudor Toma, Robert-Jan Bosman, Arno Siebes, Niels Peek, Ameen Abu-Hanna |
J. Biomed. Informatics | 5 |
| 2009 | Artificial intelligence in medicine AIME'07
Riccardo Bellazzi, Ameen Abu-Hanna |
Artif. Intell. Medicine | 2 |
| 2009 | The coming of age of artificial intelligence in medicine
Vimla L. Patel, Edward H. Shortliffe, Mario Stefanelli, Peter Szolovits, Michael R. Berthold, Riccardo Bellazzi, Ameen Abu-Hanna |
Artif. Intell. Medicine | 7 |
| 2008 | Application of Statistical Process Control Methods to Monitor Guideline Adherence: A Case Study
Niels Peek, Rick Goud, Ameen Abu-Hanna |
AMIA | 3 |
| 2008 | Discovery and integration of univariate patterns from daily individual organ-failure scores for intensive care mortality prediction
Tudor Toma, Ameen Abu-Hanna, Robert-Jan Bosman |
Artif. Intell. Medicine | 2 |
| 2007 | Discovery and Integration of Organ-Failure Episodes in Mortality Prediction
Tudor Toma, Ameen Abu-Hanna, Robert-Jan Bosman |
AIME | 2 |
| 2007 | Review Paper: Evaluation of Outpatient Computerized Physician Medication Order Entry Systems: A Systematic ReviewabstractThis paper provides a systematic literature review of CPOE evaluation studies in the outpatient setting on: safety; cost and efficiency; adherence to guideline; alerts; time; and satisfaction, usage, and usability. Thirty articles with original data (randomized clinical trial, non-randomized clinical trial, or observational study designs) met the inclusion criteria. Only four studies assessed the effect of CPOE on safety. The effect was not significant on the number of adverse drug events. Only one study showed a significant reduction of the number of medication errors. Three studies showed significant reductions in medication costs; five other studies could not support this. Most studies on adherence to guidelines showed a significant positive effect. The relatively small number of evaluation studies published to date do not provide adequate evidence that CPOE systems enhance safety and reduce cost in the outpatient settings. There is however evidence for (a) increasing adherence to guidelines, (b) increasing total prescribing time, and (c) high frequency of ignored alerts. Saeid Eslami, Ameen Abu-Hanna, Nicolette de Keizer |
J. Am. Medical Informatics Assoc. | 2 |
| 2007 | Discovery and inclusion of SOFA score episodes in mortality prediction
Tudor Toma, Ameen Abu-Hanna, Robert-Jan Bosman |
J. Biomed. Informatics | 2 |
| 2005 | Data-Driven Analysis of Blood Glucose Management Effectiveness
Barry Nannings, Ameen Abu-Hanna, Robert-Jan Bosman |
AIME | 2 |
| 2005 | Two DL-based Methods for Auditing Medical Terminological Systems
Ronald Cornet, Ameen Abu-Hanna |
AMIA | 2 |
| 2005 | Description logic-based methods for auditing frame-based medical terminological systems
Ronald Cornet, Ameen Abu-Hanna |
Artif. Intell. Medicine | 2 |
| 2005 | protégé as a vehicle for developing medical terminological systems
Ameen Abu-Hanna, Ronald Cornet, Nicolette de Keizer, Monica Crubézy, Samson W. Tu |
Int. J. Hum. Comput. Stud. | 1 |
| 2004 | Bayesian networks in biomedicine and health-care
Peter J. F. Lucas, Linda C. van der Gaag, Ameen Abu-Hanna |
Artif. Intell. Medicine | 3 |
| 2003 | Using Description Logics for Managing Medical Terminologies
Ronald Cornet, Ameen Abu-Hanna |
AIME | 2 |
| 2003 | Integrating classification trees with local logistic regression in Intensive Care prognosis
Ameen Abu-Hanna, Nicolette de Keizer |
Artif. Intell. Medicine | 1 |
| 2002 | Usability of expressive description logics-a case study in UMLS
Ronald Cornet, Ameen Abu-Hanna |
AMIA | 2 |
| 2001 | A Classification-Tree Hybrid Method for Studying Prognostic Models in Intensive Care
Ameen Abu-Hanna, Nicolette de Keizer |
AIME | 1 |
| 2001 | An Architecture for Reasoning with Terminological Systems
Ronald Cornet, Ameen Abu-Hanna, A. E. Blohm |
AMIA | 2 |
| 1995 | A case study in ontology library construction
Gertjan van Heijst, Sabina Falasconi, Ameen Abu-Hanna, Guus Schreiber, Mario Stefanelli |
Artif. Intell. Medicine | 3 |