VLDB 2026 Research / reviewers in the wild / expert
Peter R. Rijnbeek
dblp:31/8268
· DBLP profile ↗
16ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-0621-1979ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 12 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Loss function influence on hyperparameter optimization for observational healthcare prediction modelsabstractOBJECTIVES: Prediction models are increasingly used in healthcare for risk stratification and personalized care. Many models are developed using machine learning, which requires tuning hyperparameters to maximize performance based on a chosen loss function metric. In healthcare, the area under the receiver operating characteristic curve (AUROC) is commonly used for this purpose, but it may not always be the most appropriate choice for every clinical application. We empirically characterize whether the choice of loss function metric in hyperparameter optimization leads to systematic differences in model behavior across several clinical prediction tasks using real-world healthcare data. METHODS: We utilized fifteen different loss function metrics to guide hyperparameter selection across three clinical prediction tasks and four machine learning algorithms. We then compared how loss function metric choice affected selected hyperparameters, overall performance, and individual predicted probabilities. RESULTS: We observed that certain hyperparameters tended to have similar optimal values across different loss function metrics, although this pattern differed by algorithm. The best-performing models, evaluated using AUROC, were often not the models with hyperparameters optimized using AUROC. While models performed similarly at a population level, based on discrimination and calibration. The choice of the loss function metric had significant impact on the individual predicted risk for a patient. DISCUSSION: The predictive multiplicity observed can have significant impact on the patient level, while not observed in the population level model evaluation. CONCLUSION: Predictive multiplicity can have a serious impact on patient treatment decisions but is not yet well understood. Fleur Vereijken, Jenna Reps, Peter R. Rijnbeek, Ross D. Williams |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Evaluation of the impact of defining observable time in real-world data on outcome incidenceabstractOBJECTIVE: In real-world data (RWD), defining the observation period-the time during which a patient is considered observable-is critical for estimating incidence rates (IRs) and other outcomes. Yet, in the absence of explicit enrollment information, this period must often be inferred, introducing potential bias. MATERIALS AND METHODS: This study evaluates methods for defining observation periods and their impact on IR estimates across multiple database types. We applied 3 methods for defining observation periods: (1) a persistence + surveillance window approach, (2) an age- and gender-adjusted method based on time between healthcare events, and (3) the min/max method. These were tested across 11 RWD databases, including both enrollment-based and encounter-based sources. Enrollment time was used as the reference standard in eligible databases. To assess the impact on epidemiologic results, we replicated a prior study of adverse event incidence, comparing IRs and calculating mean squared error between methods. RESULTS: Incidence rates decreased as observation periods lengthened, driven by increases in the person-time denominator. The persistence + surveillance method produced estimates closest to enrollment-based rates when appropriately balanced. The min/max approach yielded inconsistent results, particularly in encounter-based databases, with greater error observed in databases with longer time spans. DISCUSSION: These findings suggest that assumptions about data completeness and population observability significantly affect incidence estimates. Observation period definitions substantially influence outcome measurement in RWD studies. CONCLUSION: Standardized, transparent approaches are necessary to ensure valid, reproducible results-especially in databases lacking defined enrollment. Clair Blacketer, Frank J. DeFalco, Mitchell Conover, Patrick B. Ryan, Martijn J. Schuemie, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Assessing the feasibility of observational data sources for multicenter clinical studiesabstractThe availability of large Electronic Health Records (EHR) databases has created new opportunities for clinical research. To improve interoperability of these databases a Common Data Model (CDM) can be used which enables standardized analytics. However, identifying the appropriate data sources for a specific study remains a challenge. The current strategy used for database discovery are based on catalogues that contain metadata which is often not rich enough for the task at hand. Additionally, sometimes this information is incorrect since it is inserted in such platforms manually by the data owners. In response to this challenge, we proposed the Concept Browser tool. The tool aims to streamline the process of efficiently exploring and selecting suitable OMOP CDM databases aligned with the study purpose. It was developed and validated in the EHDEN project, which currently contains information from more than 180 health databases across Europe and is now being used also in the DARWIN EU®initiative of the European Medicines Agency. João Rafael Almeida, Peter R. Rijnbeek, Maxim Moinat, José Luís Oliveira |
CBMS | 2 |
| 2024 | Comparing penalization methods for linear models on large observational health dataabstractOBJECTIVE: This study evaluates regularization variants in logistic regression (L1, L2, ElasticNet, Adaptive L1, Adaptive ElasticNet, Broken adaptive ridge [BAR], and Iterative hard thresholding [IHT]) for discrimination and calibration performance, focusing on both internal and external validation. MATERIALS AND METHODS: We use data from 5 US claims and electronic health record databases and develop models for various outcomes in a major depressive disorder patient population. We externally validate all models in the other databases. We use a train-test split of 75%/25% and evaluate performance with discrimination and calibration. Statistical analysis for difference in performance uses Friedman's test and critical difference diagrams. RESULTS: Of the 840 models we develop, L1 and ElasticNet emerge as superior in both internal and external discrimination, with a notable AUC difference. BAR and IHT show the best internal calibration, without a clear external calibration leader. ElasticNet typically has larger model sizes than L1. Methods like IHT and BAR, while slightly less discriminative, significantly reduce model complexity. CONCLUSION: L1 and ElasticNet offer the best discriminative performance in logistic regression for healthcare predictions, maintaining robustness across validations. For simpler, more interpretable models, L0-based methods (IHT and BAR) are advantageous, providing greater parsimony and calibration with fewer features. This study aids in selecting suitable regularization techniques for healthcare prediction models, balancing performance, complexity, and interpretability. Egill A. Fridgeirsson, Ross D. Williams, Peter R. Rijnbeek, Marc A. Suchard, Jenna Reps |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | OHDSI Standardized Vocabularies - a large-scale centralized reference ontology for international data harmonizationabstractIMPORTANCE: The Observational Health Data Sciences and Informatics (OHDSI) is the largest distributed data network in the world encompassing more than 331 data sources with 2.1 billion patient records across 34 countries. It enables large-scale observational research through standardizing the data into a common data model (CDM) (Observational Medical Outcomes Partnership [OMOP] CDM) and requires a comprehensive, efficient, and reliable ontology system to support data harmonization. MATERIALS AND METHODS: We created the OHDSI Standardized Vocabularies-a common reference ontology mandatory to all data sites in the network. It comprises imported and de novo-generated ontologies containing concepts and relationships between them, and the praxis of converting the source data to the OMOP CDM based on these. It enables harmonization through assigned domains according to clinical categories, comprehensive coverage of entities within each domain, support for commonly used international coding schemes, and standardization of semantically equivalent concepts. RESULTS: The OHDSI Standardized Vocabularies comprise over 10 million concepts from 136 vocabularies. They are used by hundreds of groups and several large data networks. More than 8600 users have performed 50 000 downloads of the system. This open-source resource has proven to address an impediment of large-scale observational research-the dependence on the context of source data representation. With that, it has enabled efficient phenotyping, covariate construction, patient-level prediction, population-level estimation, and standard reporting. DISCUSSION AND CONCLUSION: OHDSI has made available a comprehensive, open vocabulary system that is unmatched in its ability to support global observational research. We encourage researchers to exploit it and contribute their use cases to this dynamic resource. Christian G. Reich, Anna Ostropolets, Patrick B. Ryan, Peter R. Rijnbeek, Martijn J. Schuemie, Dmitry Dymshyts, George Hripcsak |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction toolsabstractOBJECTIVE: To explore the feasibility of validating Dutch concept extraction tools using annotated corpora translated from English, focusing on preserving annotations during translation and addressing the scarcity of non-English annotated clinical corpora. MATERIALS AND METHODS: Three annotated corpora were standardized and translated from English to Dutch using 2 machine translation services, Google Translate and OpenAI GPT-4, with annotations preserved through a proposed method of embedding annotations in the text before translation. The performance of 2 concept extraction tools, MedSpaCy and MedCAT, was assessed across the corpora in both Dutch and English. RESULTS: The translation process effectively generated Dutch annotated corpora and the concept extraction tools performed similarly in both English and Dutch. Although there were some differences in how annotations were preserved across translations, these did not affect extraction accuracy. Supervised MedCAT models consistently outperformed unsupervised models, whereas MedSpaCy demonstrated high recall but lower precision. DISCUSSION: Our validation of Dutch concept extraction tools on corpora translated from English was successful, highlighting the efficacy of our annotation preservation method and the potential for efficiently creating multilingual corpora. Further improvements and comparisons of annotation preservation techniques and strategies for corpus synthesis could lead to more efficient development of multilingual corpora and accurate non-English concept extraction tools. CONCLUSION: This study has demonstrated that translated English corpora can be used to validate non-English concept extraction tools. The annotation preservation method used during translation proved effective, and future research can apply this corpus translation method to additional languages and clinical settings. Tom M. Seinen, Jan A. Kors, Erik M. van Mulligen, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | The added value of text from Dutch general practitioner notes in predictive modelingabstractOBJECTIVE: This work aims to explore the value of Dutch unstructured data, in combination with structured data, for the development of prognostic prediction models in a general practitioner (GP) setting. MATERIALS AND METHODS: We trained and validated prediction models for 4 common clinical prediction problems using various sparse text representations, common prediction algorithms, and observational GP electronic health record (EHR) data. We trained and validated 84 models internally and externally on data from different EHR systems. RESULTS: On average, over all the different text representations and prediction algorithms, models only using text data performed better or similar to models using structured data alone in 2 prediction tasks. Additionally, in these 2 tasks, the combination of structured and text data outperformed models using structured or text data alone. No large performance differences were found between the different text representations and prediction algorithms. DISCUSSION: Our findings indicate that the use of unstructured data alone can result in well-performing prediction models for some clinical prediction problems. Furthermore, the performance improvement achieved by combining structured and text data highlights the added value. Additionally, we demonstrate the significance of clinical natural language processing research in languages other than English and the possibility of validating text-based prediction models across various EHR systems. CONCLUSION: Our study highlights the potential benefits of incorporating unstructured data in clinical prediction models in a GP setting. Although the added value of unstructured data may vary depending on the specific prediction task, our findings suggest that it has the potential to enhance patient care. Tom M. Seinen, Jan A. Kors, Erik M. van Mulligen, Egill A. Fridgeirsson, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 5 |
| 2023 | European Health Data & Evidence Network - learnings from building out a standardized international health data networkabstractOBJECTIVE: Health data standardized to a common data model (CDM) simplifies and facilitates research. This study examines the factors that make standardizing observational health data to the Observational Medical Outcomes Partnership (OMOP) CDM successful. MATERIALS AND METHODS: Twenty-five data partners (DPs) from 11 countries received funding from the European Health Data Evidence Network (EHDEN) to standardize their data. Three surveys, DataQualityDashboard results, and statistics from the conversion process were analyzed qualitatively and quantitatively. Our measures of success were the total number of days to transform source data into the OMOP CDM and participation in network research. RESULTS: The health data converted to CDM represented more than 133 million patients. 100%, 88%, and 84% of DPs took Surveys 1, 2, and 3. The median duration of the 6 key extract, transform, and load (ETL) processes ranged from 4 to 115 days. Of the 25 DPs, 21 DPs were considered applicable for analysis of which 52% standardized their data on time, and 48% participated in an international collaborative study. DISCUSSION: This study shows that the consistent workflow used by EHDEN proves appropriate to support the successful standardization of observational data across Europe. Over the 25 successful transformations, we confirmed that getting the right people for the ETL is critical and vocabulary mapping requires specific expertise and support of tools. Additionally, we learned that teams that proactively prepared for data governance issues were able to avoid considerable delays improving their ability to finish on time. CONCLUSION: This study provides guidance for future DPs to standardize to the OMOP CDM and participate in distributed networks. We demonstrate that the Observational Health Data Sciences and Informatics community must continue to evaluate and provide guidance and support for what ultimately develops the backbone of how community members generate evidence. Erica A. Voss, Clair Blacketer, Sebastiaan van Sandijk, Maxim Moinat, Michael Kallfelz, Michel Van Speybroeck, Daniel Prieto-Alhambra, Martijn J. Schuemie, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 9 |
| 2022 | Use of unstructured text in prognostic clinical prediction models: a systematic reviewabstractOBJECTIVE: This systematic review aims to assess how information from unstructured text is used to develop and validate clinical prognostic prediction models. We summarize the prediction problems and methodological landscape and determine whether using text data in addition to more commonly used structured data improves the prediction performance. MATERIALS AND METHODS: We searched Embase, MEDLINE, Web of Science, and Google Scholar to identify studies that developed prognostic prediction models using information extracted from unstructured text in a data-driven manner, published in the period from January 2005 to March 2021. Data items were extracted, analyzed, and a meta-analysis of the model performance was carried out to assess the added value of text to structured-data models. RESULTS: We identified 126 studies that described 145 clinical prediction problems. Combining text and structured data improved model performance, compared with using only text or only structured data. In these studies, a wide variety of dense and sparse numeric text representations were combined with both deep learning and more traditional machine learning methods. External validation, public availability, and attention for the explainability of the developed models were limited. CONCLUSION: The use of unstructured text in the development of prognostic prediction models has been found beneficial in addition to structured data in most studies. The text data are source of valuable information for prediction model development and should not be neglected. We suggest a future focus on explainability and external validation of the developed models, promoting robust and trustworthy prediction models in clinical practice. Tom M. Seinen, Egill A. Fridgeirsson, Solomon Ioannou, Daniel Jeannetot, Luis H. John, Jan A. Kors, Aniek F. Markus, Victor Pera, Alexandros Rekkas, Ross D. Williams, Cynthia Yang, Erik M. van Mulligen, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 13 |
| 2022 | Trends in the conduct and reporting of clinical prediction model development and validation: a systematic reviewabstractOBJECTIVES: This systematic review aims to provide further insights into the conduct and reporting of clinical prediction model development and validation over time. We focus on assessing the reporting of information necessary to enable external validation by other investigators. MATERIALS AND METHODS: We searched Embase, Medline, Web-of-Science, Cochrane Library, and Google Scholar to identify studies that developed 1 or more multivariable prognostic prediction models using electronic health record (EHR) data published in the period 2009-2019. RESULTS: We identified 422 studies that developed a total of 579 clinical prediction models using EHR data. We observed a steep increase over the years in the number of developed models. The percentage of models externally validated in the same paper remained at around 10%. Throughout 2009-2019, for both the target population and the outcome definitions, code lists were provided for less than 20% of the models. For about half of the models that were developed using regression analysis, the final model was not completely presented. DISCUSSION: Overall, we observed limited improvement over time in the conduct and reporting of clinical prediction model development and validation. In particular, the prediction problem definition was often not clearly reported, and the final model was often not completely presented. CONCLUSION: Improvement in the reporting of information necessary to enable external validation by other investigators is still urgently needed to increase clinical adoption of developed models. Cynthia Yang, Jan A. Kors, Solomon Ioannou, Luis H. John, Aniek F. Markus, Alexandros Rekkas, Maria A. J. de Ridder, Tom M. Seinen, Ross D. Williams, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 10 |
| 2021 | Increasing trust in real-world evidence through evaluation of observational data qualityabstractOBJECTIVE: Advances in standardization of observational healthcare data have enabled methodological breakthroughs, rapid global collaboration, and generation of real-world evidence to improve patient outcomes. Standardizations in data structure, such as use of common data models, need to be coupled with standardized approaches for data quality assessment. To ensure confidence in real-world evidence generated from the analysis of real-world data, one must first have confidence in the data itself. MATERIALS AND METHODS: We describe the implementation of check types across a data quality framework of conformance, completeness, plausibility, with both verification and validation. We illustrate how data quality checks, paired with decision thresholds, can be configured to customize data quality reporting across a range of observational health data sources. We discuss how data quality reporting can become part of the overall real-world evidence generation and dissemination process to promote transparency and build confidence in the resulting output. RESULTS: The Data Quality Dashboard is an open-source R package that reports potential quality issues in an OMOP CDM instance through the systematic execution and summarization of over 3300 configurable data quality checks. DISCUSSION: Transparently communicating how well common data model-standardized databases adhere to a set of quality measures adds a crucial piece that is currently missing from observational research. CONCLUSION: Assessing and improving the quality of our data will inherently improve the quality of the evidence we generate. Clair Blacketer, Frank J. DeFalco, Patrick B. Ryan, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | The role of explainability in creating trustworthy artificial intelligence for health care: A comprehensive survey of the terminology, design choices, and evaluation strategiesabstractArtificial intelligence (AI) has huge potential to improve the health and well-being of people, but adoption in clinical practice is still limited. Lack of transparency is identified as one of the main barriers to implementation, as clinicians should be confident the AI system can be trusted. Explainable AI has the potential to overcome this issue and can be a step towards trustworthy AI. In this paper we review the recent literature to provide guidance to researchers and practitioners on the design of explainable AI systems for the health-care domain and contribute to formalization of the field of explainable AI. We argue the reason to demand explainability determines what should be explained as this determines the relative importance of the properties of explainability (i.e. interpretability and fidelity). Based on this, we propose a framework to guide the choice between classes of explainable AI methods (explainable modelling versus post-hoc explanation; model-based, attribution-based, or example-based explanations; global and local explanations). Furthermore, we find that quantitative evaluation metrics, which are important for objective standardized evaluation, are still lacking for some properties (e.g. clarity) and types of explanations (e.g. example-based methods). We conclude that explainable modelling can contribute to trustworthy AI, but the benefits of explainability still need to be proven in practice and complementary measures might be needed to create trustworthy AI in health care (e.g. reporting data quality, performing extensive (external) validation, and regulation). Aniek F. Markus, Jan A. Kors, Peter R. Rijnbeek |
J. Biomed. Informatics | 3 |
| 2019 | Supplementing claims data analysis using self-reported data to develop a probabilistic phenotype model for current smoking status
Jenna Reps, Peter R. Rijnbeek, Patrick B. Ryan |
J. Biomed. Informatics | 2 |
| 2018 | Design and implementation of a standardized framework to generate and evaluate patient-level prediction models using observational healthcare dataabstractObjective: To develop a conceptual prediction model framework containing standardized steps and describe the corresponding open-source software developed to consistently implement the framework across computational environments and observational healthcare databases to enable model sharing and reproducibility. Methods: Based on existing best practices we propose a 5 step standardized framework for: (1) transparently defining the problem; (2) selecting suitable datasets; (3) constructing variables from the observational data; (4) learning the predictive model; and (5) validating the model performance. We implemented this framework as open-source software utilizing the Observational Medical Outcomes Partnership Common Data Model to enable convenient sharing of models and reproduction of model evaluation across multiple observational datasets. The software implementation contains default covariates and classifiers but the framework enables customization and extension. Results: As a proof-of-concept, demonstrating the transparency and ease of model dissemination using the software, we developed prediction models for 21 different outcomes within a target population of people suffering from depression across 4 observational databases. All 84 models are available in an accessible online repository to be implemented by anyone with access to an observational database in the Common Data Model format. Conclusions: The proof-of-concept study illustrates the framework's ability to develop reproducible models that can be readily shared and offers the potential to perform extensive external validation of models, and improve their likelihood of clinical uptake. In future work the framework will be applied to perform an "all-by-all" prediction analysis to assess the observational data prediction domain across numerous target populations, outcomes and time, and risk settings. Jenna Reps, Martijn J. Schuemie, Marc A. Suchard, Patrick B. Ryan, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | Accuracy of an automated knowledge base for identifying drug adverse reactions
Erica A. Voss, Richard D. Boyce, Patrick B. Ryan, Johan van der Lei, Peter R. Rijnbeek, Martijn J. Schuemie |
J. Biomed. Informatics | 5 |
| 2010 | Finding a short and accurate decision rule in disjunctive normal form by exhaustive searchabstractGreedy approaches suffer from a restricted search space which could lead to suboptimal classifiers in terms of performance and classifier size. This study discusses exhaustive search as an alternative to greedy search for learning short and accurate decision rules. The Exhaustive Procedure for LOgic-Rule Extraction (EXPLORE) algorithm is presented, to induce decision rules in disjunctive normal form (DNF) in a systematic and efficient manner. We propose a method based on subsumption to reduce the number of values considered for instantiation in the literals, by taking into account the relational operator without loss of performance. Furthermore, we describe a branch-and-bound approach that makes optimal use of user-defined performance constraints. To improve the generalizability we use a validation set to determine the optimal length of the DNF rule. The performance and size of the DNF rules induced by EXPLORE are compared to those of eight well-known rule learners. Our results show that an exhaustive approach to rule learning in DNF results in significantly smaller classifiers than those of the other rule learners, while securing comparable or even better performance. Clearly, exhaustive search is computer-intensive and may not always be feasible. Nevertheless, based on this study, we believe that exhaustive search should be considered an alternative for greedy search in many problems. Peter R. Rijnbeek, Jan A. Kors |
Mach. Learn. | 1 |