EDBT 2026 Demo / reviewers in the wild / expert
Anton van der Vegt
dblp:215/9176 · also Anton H. van der Vegt
· DBLP profile ↗
9ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0001-5642-5188ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A novel, standardised approach to balancing effectiveness, efficiency and utility of surveillance AI prediction models for hospitalised patients using sepsis prediction as an exemplarabstractOBJECTIVE: To introduce a novel, standardised approach to evaluating AI prediction models in balancing effectiveness, efficiency and utility, using a sepsis prediction model case study. MATERIALS AND METHODS: Retrospective patient data from electronic medical records of 7 public hospitals was used to retrain and evaluate a machine learning sepsis prediction model. Four conventional metrics-area under the receiver operating curve (AUROC), sensitivity, positive predictive value, and specificity-were compared with a novel graphical display integrating metrics of predictive accuracy (effectiveness), alert burden (efficiency) and lead time of alerts relative to clinical events (utility) for different alert thresholds. RESULTS: The dataset comprised 977,506 inpatient admissions. The novel methodology produced a plot of four vertically aligned graphs that enables decision-makers to identify an alert threshold that optimally balances effectiveness, efficiency and utility (EEU) at the level of an entire admission, and which differs from that derived using conventional metrics. DISCUSSION: Conventional evaluation metrics do not consider alert timing relative to clinical events and are often applied to different evaluation datasets (sample and admission level), introducing bias and confusion. In contrast, the EEU methodology (i) generates admission level evaluations at different alert thresholds; (ii) measures alert timing relative to clinical events; and (iii) provides a visual display that enables identification of the alert threshold that optimally balances EEU factors. CONCLUSION: Evaluations of prediction models for adverse events in hospitalised patients should incorporate the EEU approach in assessing model suitability and selecting alert thresholds. Anton van der Vegt, Victoria Campbell, Robert Webb, Balasubramanian Venkatesh, Paul J. Lane, Kathryn Wilks, Steven McPhail, Michael Rice, Tara Isaacs, Ahmad Abdel-Hafez, Stephen Whebell, Adam Irwin, Rudolf J. Schnetler, Amith Shetty, Ian A. Scott |
J. Am. Medical Informatics Assoc. | 1 |
| 2025 | RARR Unraveled: Component-Level Insights into Hallucination Detection and MitigationabstractLarge Language Models (LLMs) often exhibit hallucinations, which makes detecting and mitigating these errors a critical challenge. The Retrofit Attribution using Research and Revision (RARR) framework addresses this challenge by extracting key aspects of an LLM response, verifying them against retrieved evidence, and resolving errors through re-prompting. In this work, we critically examine RARR and adapt its framework to incorporate publicly available evidence retrieval systems and generative models, thereby operationalizing the approach. We focus on hallucination detection, analyzing how each pipeline component contributes to this task. We also conduct a sentence-level analysis of hallucinations to provide a more granular assessment of RARR's performance. A key finding is that while query generation and retrieval are effective, the agreement module emerges as the weakest link in the RARR pipeline. We offer deeper insights into RARR's strengths, limitations, and potential areas for improvement, thereby broadening our understanding of hallucination detection in LLMs. Jonathan J. Ross, Ekaterina Khramtsova, Anton van der Vegt, Bevan Koopman, Guido Zuccon |
SIGIR | 3 |
| 2025 | Factors underpinning the performance of implemented artificial intelligence-based patient deterioration prediction systems: reasons for selection and implications for hospitals and researchersabstractOBJECTIVE: The degree to which deployed artificial intelligence-based deterioration prediction algorithms (AI-DPA) differ in their development, the reasons for these differences, and how this may impact their performance remains unclear. Our primary objective was to identify design factors and associated decisions related to the development of AI-DPA and highlight deficits that require further research. MATERIALS AND METHODS: Based on a systematic review of 14 deployed AI-DPA and an updated systematic search, we identified studies of 12 eligible AI-DPA from which data were extracted independently by 2 investigators on all design factors, decisions, and justifications pertaining to 6 machine learning development stages: (1) model requirements, (2) data collection, (3) data cleaning, (4) data labeling, (5) feature engineering, and (6) model training. RESULTS: We found 13 design factors and 315 decision alternatives likely to impact AI-DPA performance, all of which varied, together with their rationales, between all included AI-DPA. Variable selection, data imputation methods, training data exclusions, training sample definitions, length of lookback periods, and definition of outcome labels were key design factors accounting for most variation. In justifying decisions, most studies made no reference to prior research or compared with other state-of-the-art algorithms. DISCUSSION: Algorithm design decisions regarding factors impacting AI-DPA performance have little supporting evidence, are inconsistent, do not learn from prior work, and lack reference standards. CONCLUSION: Several deficits in AI-DPA development that prevent implementers selecting the most accurate algorithm have been identified, and future research needs to address these deficits as a priority. Anton van der Vegt, Victoria Campbell, James Malycha, Ian A. Scott |
J. Am. Medical Informatics Assoc. | 1 |
| 2025 | Presenting Artificial Intelligence Predictions Based on Electronic Medical Records to Clinicians in Hospitals: A Systematic ReviewabstractOur objective was to investigate how artificial intelligence (AI) predictions calculated on structured hospital data are presented to clinicians. We performed a systematic review of 5 databases and 9 other reviews, identifying 31 studies on 21 implemented clinical AI systems. We report current approaches to presenting AI predictions to clinicians, whether and how interaction on the user interface (UI) is used, how UIs have been evaluated and the extent to which clinicians have been involved in UI design and testing. The results indicate variation across systems in presentation content and styles, evaluation methods, and interaction approaches. Half of the systems implemented a co-design approach to UI development. Our findings provide valuable insights for future AI-based clinical decision support system designers, clinical AI researchers and healthcare organisations seeking to implement clinical AI solutions. Monica Noselli, Anton van der Vegt, Maxime Cordeil, Victoria Campbell, Audrey P. Wang, Amith Shetty, Ian A. Scott |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Systematic review and longitudinal analysis of implementing Artificial Intelligence to predict clinical deterioration in adult hospitals: what is known and what remains uncertainabstractOBJECTIVE: To identify factors influencing implementation of machine learning algorithms (MLAs) that predict clinical deterioration in hospitalized adult patients and relate these to a validated implementation framework. MATERIALS AND METHODS: A systematic review of studies of implemented or trialed real-time clinical deterioration prediction MLAs was undertaken, which identified: how MLA implementation was measured; impact of MLAs on clinical processes and patient outcomes; and barriers, enablers and uncertainties within the implementation process. Review findings were then mapped to the SALIENT end-to-end implementation framework to identify the implementation stages at which these factors applied. RESULTS: Thirty-seven articles relating to 14 groups of MLAs were identified, each trialing or implementing a bespoke algorithm. One hundred and seven distinct implementation evaluation metrics were identified. Four groups reported decreased hospital mortality, 1 significantly. We identified 24 barriers, 40 enablers, and 14 uncertainties and mapped these to the 5 stages of the SALIENT implementation framework. DISCUSSION: Algorithm performance across implementation stages decreased between in silico and trial stages. Silent plus pilot trial inclusion was associated with decreased mortality, as was the use of logistic regression algorithms that used less than 39 variables. Mitigation of alert fatigue via alert suppression and threshold configuration was commonly employed across groups. CONCLUSIONS: : There is evidence that real-world implementation of clinical deterioration prediction MLAs may improve clinical outcomes. Various factors identified as influencing success or failure of implementation can be mapped to different stages of implementation, thereby providing useful and practical guidance for implementers. Anton van der Vegt, Victoria Campbell, Imogen Mitchell, James Malycha, Joanna Simpson, Tracy Flenady, Arthas Flabouris, Paul J. Lane, Naitik Mehta, Vikrant R. Kalke, Jovie A Decoyna, Nicholas Es'haghi, Chun-Huei Liu, Ian A. Scott |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Deployment of machine learning algorithms to predict sepsis: systematic review and application of the SALIENT clinical AI implementation frameworkabstractOBJECTIVE: To retrieve and appraise studies of deployed artificial intelligence (AI)-based sepsis prediction algorithms using systematic methods, identify implementation barriers, enablers, and key decisions and then map these to a novel end-to-end clinical AI implementation framework. MATERIALS AND METHODS: Systematically review studies of clinically applied AI-based sepsis prediction algorithms in regard to methodological quality, deployment and evaluation methods, and outcomes. Identify contextual factors that influence implementation and map these factors to the SALIENT implementation framework. RESULTS: The review identified 30 articles of algorithms applied in adult hospital settings, with 5 studies reporting significantly decreased mortality post-implementation. Eight groups of algorithms were identified, each sharing a common algorithm. We identified 14 barriers, 26 enablers, and 22 decision points which were able to be mapped to the 5 stages of the SALIENT implementation framework. DISCUSSION: Empirical studies of deployed sepsis prediction algorithms demonstrate their potential for improving care and reducing mortality but reveal persisting gaps in existing implementation guidance. In the examined publications, key decision points reflecting real-word implementation experience could be mapped to the SALIENT framework and, as these decision points appear to be AI-task agnostic, this framework may also be applicable to non-sepsis algorithms. The mapping clarified where and when barriers, enablers, and key decisions arise within the end-to-end AI implementation process. CONCLUSIONS: A systematic review of real-world implementation studies of sepsis prediction algorithms was used to validate an end-to-end staged implementation framework that has the ability to account for key factors that warrant attention in ensuring successful deployment, and which extends on previous AI implementation frameworks. Anton van der Vegt, Ian A. Scott, Krishna Dermawan, Rudolf J. Schnetler, Vikrant R. Kalke, Paul J. Lane |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Implementation frameworks for end-to-end clinical AI: derivation of the SALIENT frameworkabstractOBJECTIVE: To derive a comprehensive implementation framework for clinical AI models within hospitals informed by existing AI frameworks and integrated with reporting standards for clinical AI research. MATERIALS AND METHODS: (1) Derive a provisional implementation framework based on the taxonomy of Stead et al and integrated with current reporting standards for AI research: TRIPOD, DECIDE-AI, CONSORT-AI. (2) Undertake a scoping review of published clinical AI implementation frameworks and identify key themes and stages. (3) Perform a gap analysis and refine the framework by incorporating missing items. RESULTS: The provisional AI implementation framework, called SALIENT, was mapped to 5 stages common to both the taxonomy and the reporting standards. A scoping review retrieved 20 studies and 247 themes, stages, and subelements were identified. A gap analysis identified 5 new cross-stage themes and 16 new tasks. The final framework comprised 5 stages, 7 elements, and 4 components, including the AI system, data pipeline, human-computer interface, and clinical workflow. DISCUSSION: This pragmatic framework resolves gaps in existing stage- and theme-based clinical AI implementation guidance by comprehensively addressing the what (components), when (stages), and how (tasks) of AI implementation, as well as the who (organization) and why (policy domains). By integrating research reporting standards into SALIENT, the framework is grounded in rigorous evaluation methodologies. The framework requires validation as being applicable to real-world studies of deployed AI models. CONCLUSIONS: A novel end-to-end framework has been developed for implementing AI within hospital clinical practice that builds on previous AI implementation frameworks and research reporting standards. Anton van der Vegt, Ian A. Scott, Krishna Dermawan, Rudolf J. Schnetler, Vikrant R. Kalke, Paul J. Lane |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | Do better search engines really equate to better clinical decisions? If not, why not?abstractAbstract Previous research has found that improved search engine effectiveness—evaluated using a batch‐style approach—does not always translate to significant improvements in user task performance; however, these prior studies focused on simple recall and precision‐based search tasks. We investigated the same relationship, but for realistic, complex search tasks required in clinical decision making. One hundred and nine clinicians and final year medical students answered 16 clinical questions. Although the search engine did improve answer accuracy by 20 percentage points, there was no significant difference when participants used a more effective, state‐of‐the‐art search engine. We also found that the search engine effectiveness difference, identified in the lab, was diminished by around 70% when the search engines were used with real users. Despite the aid of the search engine, half of the clinical questions were answered incorrectly. We further identified the relative contribution of search engine effectiveness to the overall end task success. We found that the ability to interpret documents correctly was a much more important factor impacting task success. If these findings are representative, information retrieval research may need to reorient its emphasis towards helping users to better understand information, rather than just finding it for them. Anton van der Vegt, Guido Zuccon, Bevan Koopman |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2019 | Learning Inter-Sentence, Disorder-Centric, Biomedical Relationships from Medical Literature
Anton van der Vegt, Guido Zuccon, Bevan Koopman |
AMIA | 1 |