Folkert W. Asselbergs

dblp:53/7969 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-1692-8669ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 58% Medical and health informatics · 42%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Medical and health informatics
clinical decision support
0.412020
Model selection for metabolomics: predicting diagnosis of coronary artery disease using automated machine learning · Bioinform. 2020
Medical and health informatics › clinical diagnosis
coronary artery disease diagnosis
0.412020
Model selection for metabolomics: predicting diagnosis of coronary artery disease using automated machine learning · Bioinform. 2020
Bioinformatics and computational biology › metabolomics
metabolic profiling
0.412020
Model selection for metabolomics: predicting diagnosis of coronary artery disease using automated machine learning · Bioinform. 2020
Bioinformatics and computational biology
metabolomics
0.412020
Model selection for metabolomics: predicting diagnosis of coronary artery disease using automated machine learning · Bioinform. 2020
Bioinformatics and computational biology › statistical genetics
gene-environment interaction
0.112010
Bioinformatics challenges for genome-wide association studies · Bioinform. 2010
Bioinformatics and computational biology › statistical genetics
gene-gene interaction
0.112010
Bioinformatics challenges for genome-wide association studies · Bioinform. 2010
Bioinformatics and computational biology › genomics
genome-wide association study
0.112010
Bioinformatics challenges for genome-wide association studies · Bioinform. 2010

Methods — techniques the papers use, named apart from their topics

tree-based pipeline optimization · 0.4genetic programming · 0.4automated machine learning · 0.4biostatistical methods · 0.1
YearPublicationVenuePosition
2025 High-Quality Analytic Isosurface Rendering for Web-Based Cardiac Imaging
abstract
Advancements in GPU-accelerated web-based 3D rendering have enabled high-quality, interactive medical visualization directly within standard browsers, eliminating the need for native software installations. Building on this progress, we present a fully browser-native isosurface renderer designed for volumetric medical datasets. The system supports both trilinear and tricubic interpolation, enabling sub-voxel accurate surface reconstruction through analytic root-finding techniques. To improve rendering efficiency, we extend empty-space skipping strategies to accommodate tricubic interpolation. Additionally, we introduce a novel early rejection test based on Bernstein polynomials, which efficiently identifies and excludes regions devoid of isosurface intersections-yielding up to$1.5 \times$speedups, especially in the computationally intensive tricubic case. While each component exists in native graphics pipelines, their integration into a browser-based framework bridges the gap between desktop-level performance and installation-free access. We validate our approach using contrast-enhanced cardiac CT volumes, achieving interactive frame rates ranging from 25 to 45 frames per second (FPS). on consumer-grade hardware. In a targeted usability study, 7 out of 8 clinicians reported that the system improved their ability to visualize patient data and communicate medical findings.
Vasileios Lazaros Charalampidis, M. Louis Handoko, Eduard Ródenas-Alesinao, Folkert W. Asselbergs, Konstantinos Votis, Paschalis Bizopoulos, Andreas Triantafyllidis
BIBE4
2023 Integrated rapid-cycle comparative effectiveness trials using flexible point of care randomisation in electronic health record systems
Matthew G. Wilson, Edward Palmer, Folkert W. Asselbergs, Steve K. Harris
J. Biomed. Informatics3
2022 Transforming and evaluating the UK Biobank to the OMOP Common Data Model for COVID-19 research and beyond
abstract
OBJECTIVE: The coronavirus disease 2019 (COVID-19) pandemic has demonstrated the value of real-world data for public health research. International federated analyses are crucial for informing policy makers. Common data models (CDMs) are critical for enabling these studies to be performed efficiently. Our objective was to convert the UK Biobank, a study of 500 000 participants with rich genetic and phenotypic data to the Observational Medical Outcomes Partnership (OMOP) CDM. MATERIALS AND METHODS: We converted UK Biobank data to OMOP CDM v. 5.3. We transformedparticipant research data on diseases collected at recruitment and electronic health records (EHRs) from primary care, hospitalizations, cancer registrations, and mortality from providers in England, Scotland, and Wales. We performed syntactic and semantic validations and compared comorbidities and risk factors between source and transformed data. RESULTS: We identified 502 505 participants (3086 with COVID-19) and transformed 690 fields (1 373 239 555 rows) to the OMOP CDM using 8 different controlled clinical terminologies and bespoke mappings. Specifically, we transformed self-reported noncancer illnesses 946 053 (83.91% of all source entries), cancers 37 802 (70.81%), medications 1 218 935 (88.25%), and prescriptions 864 788 (86.96%). In EHR, we transformed 13 028 182 (99.95%) hospital diagnoses, 6 465 399 (89.2%) procedures, 337 896 333 primary care diagnoses (CTV3, SNOMED-CT), 139 966 587 (98.74%) prescriptions (dm+d) and 77 127 (99.95%) deaths (ICD-10). We observed good concordance across demographic, risk factor, and comorbidity factors between source and transformed data. DISCUSSION AND CONCLUSION: Our study demonstrated that the OMOP CDM can be successfully leveraged to harmonize complex large-scale biobanked studies combining rich multimodal phenotypic data. Our study uncovered several challenges when transforming data from questionnaires to the OMOP CDM which require further research. The transformed UK Biobank resource is a valuable tool that can enable federated research, like COVID-19 studies.
Václav Papez, Maxim Moinat, Erica A. Voss, Sofia Bazakou, Anne Van Winzum, Alessia Peviani, Stefan Payralbe, Elena Garcia Lara, Michael Kallfelz, Folkert W. Asselbergs, Daniel Prieto-Alhambra, Richard J. B. Dobson, Spiros C. Denaxas
J. Am. Medical Informatics Assoc.10
2020 Model selection for metabolomics: predicting diagnosis of coronary artery disease using automated machine learning
abstract
MOTIVATION: Selecting the optimal machine learning (ML) model for a given dataset is often challenging. Automated ML (AutoML) has emerged as a powerful tool for enabling the automatic selection of ML methods and parameter settings for the prediction of biomedical endpoints. Here, we apply the tree-based pipeline optimization tool (TPOT) to predict angiographic diagnoses of coronary artery disease (CAD). With TPOT, ML models are represented as expression trees and optimal pipelines discovered using a stochastic search method called genetic programing. We provide some guidelines for TPOT-based ML pipeline selection and optimization-based on various clinical phenotypes and high-throughput metabolic profiles in the Angiography and Genes Study (ANGES). RESULTS: We analyzed nuclear magnetic resonance-derived lipoprotein and metabolite profiles in the ANGES cohort with a goal to identify the role of non-obstructive CAD patients in CAD diagnostics. We performed a comparative analysis of TPOT-generated ML pipelines with selected ML classifiers, optimized with a grid search approach, applied to two phenotypic CAD profiles. As a result, TPOT-generated ML pipelines that outperformed grid search optimized models across multiple performance metrics including balanced accuracy and area under the precision-recall curve. With the selected models, we demonstrated that the phenotypic profile that distinguishes non-obstructive CAD patients from no CAD patients is associated with higher precision, suggesting a discrepancy in the underlying processes between these phenotypes. AVAILABILITY AND IMPLEMENTATION: TPOT is freely available via http://epistasislab.github.io/tpot/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alena Orlenko, Daniel Kofink, Leo-Pekka Lyytikäinen, Kjell Nikus, Pashupati P. Mishra, Pekka Kuukasjärvi, Pekka J. Karhunen, Mika Kähönen, Jari O. Laurikka, Terho Lehtimäki, Folkert W. Asselbergs, Jason H. Moore
Bioinform.11
2020 ETM: Enrichment by topic modeling for automated clinical sentence classification to detect patients' disease history
abstract
Abstract Given the rapid rate at which text data are being digitally gathered in the medical domain, there is growing need for automated tools that can analyze clinical notes and classify their sentences in electronic health records (EHRs). This study uses EHR texts to detect patients’ disease history from clinical sentences. However, in EHRs, sentences are less topic-focused and shorter than that in general domain, which leads to the sparsity of co-occurrence patterns and the lack of semantic features. To tackle this challenge, current approaches for clinical sentence classification are dependent on external information to improve classification performance. However, this is implausible owing to a lack of universal medical dictionaries. This study proposes the ETM (enrichment by topic modeling) algorithm, based on latent Dirichlet allocation, to smoothen the semantic representations of short sentences. The ETM enriches text representation by incorporating probability distributions generated by an unsupervised algorithm into it. It considers the length of the original texts to enhance representation by using an internal knowledge acquisition procedure. When it comes to clinical predictive modeling, interpretability improves the acceptance of the model. Thus, for clinical sentence classification, the ETM approach employs an initial TFiDF (term frequency inverse document frequency) representation, where we use the support vector machine and neural network algorithms for the classification task. We conducted three sets of experiments on a data set consisting of clinical cardiovascular notes from the Netherlands to test the sentence classification performance of the proposed method in comparison with prevalent approaches. The results show that the proposed ETM approach outperformed state-of-the-art baselines.
Ayoub Bagheri, Arjan Sammani, Peter G. M. van der Heijden, Folkert W. Asselbergs, Daniel L. Oberski
J. Intell. Inf. Syst.4
2020 Natural Language Processing for Mimicking Clinical Trial Recruitment in Critical Care: A Semi-Automated Simulation Based on the LeoPARDS Trial
abstract
Clinical trials often fail to recruit an adequate number of appropriate patients. Identifying eligible trial participants is resource-intensive when relying on manual review of clinical notes, particularly in critical care settings where the time window is short. Automated review of electronic health records (EHR) may help, but much of the information is in free text rather than a computable form. We applied natural language processing (NLP) to free text EHR data using the CogStack platform to simulate recruitment into the LeoPARDS study, a clinical trial aiming to reduce organ dysfunction in septic shock. We applied an algorithm to identify eligible patients using a moving 1-hour time window, and compared patients identified by our approach with those actually screened and recruited for the trial, for the time period that data were available. We manually reviewed records of a random sample of patients identified by the algorithm but not screened in the original trial. Our method identified 376 patients, including 34 patients with EHR data available who were actually recruited to LeoPARDS in our centre. The sensitivity of CogStack for identifying patients screened was 90% (95% CI 85%, 93%). Of the 203 patients identified by both manual screening and CogStack, the index date matched in 95 (47%) and CogStack was earlier in 94 (47%). In conclusion, analysis of EHR data using NLP could effectively replicate recruitment in a critical care trial, and identify some eligible patients at an earlier stage, potentially improving trial recruitment if implemented in real time.
Hegler Tissot, Anoop D. Shah, David Brealey, Steve K. Harris, Ruth Agbakoba, Amos Folarin, Luis Romao, Lukasz Roguski, Richard J. B. Dobson, Folkert W. Asselbergs
IEEE J. Biomed. Health Informatics10
2015 An Independent Filter for Gene Set Testing Based on Spectral Enrichment
abstract
Gene set testing has become an indispensable tool for the analysis of high-dimensional genomic data. An important motivation for testing gene sets, rather than individual genomic variables, is to improve statistical power by reducing the number of tested hypotheses. Given the dramatic growth in common gene set collections, however, testing is often performed with nearly as many gene sets as underlying genomic variables. To address the challenge to statistical power posed by large gene set collections, we have developed spectral gene set filtering (SGSF), a novel technique for independent filtering of gene set collections prior to gene set testing. The SGSF method uses as a filter statistic the p-value measuring the statistical significance of the association between each gene set and the sample principal components (PCs), taking into account the significance of the associated eigenvalues. Because this filter statistic is independent of standard gene set test statistics under the null hypothesis but dependent under the alternative, the proportion of enriched gene sets is increased without impacting the type I error rate. As shown using simulated and real gene expression data, the SGSF algorithm accurately filters gene sets unrelated to the experimental outcome resulting in significantly increased gene set testing power.
H. Robert Frost, Folkert W. Asselbergs, Jason H. Moore
IEEE ACM Trans. Comput. Biol. Bioinform.3
2010 Bioinformatics challenges for genome-wide association studies
abstract
MOTIVATION: The sequencing of the human genome has made it possible to identify an informative set of >1 million single nucleotide polymorphisms (SNPs) across the genome that can be used to carry out genome-wide association studies (GWASs). The availability of massive amounts of GWAS data has necessitated the development of new biostatistical methods for quality control, imputation and analysis issues including multiple testing. This work has been successful and has enabled the discovery of new associations that have been replicated in multiple studies. However, it is now recognized that most SNPs discovered via GWAS have small effects on disease susceptibility and thus may not be suitable for improving health care through genetic testing. One likely explanation for the mixed results of GWAS is that the current biostatistical analysis paradigm is by design agnostic or unbiased in that it ignores all prior knowledge about disease pathobiology. Further, the linear modeling framework that is employed in GWAS often considers only one SNP at a time thus ignoring their genomic and environmental context. There is now a shift away from the biostatistical approach toward a more holistic approach that recognizes the complexity of the genotype-phenotype relationship that is characterized by significant heterogeneity and gene-gene and gene-environment interaction. We argue here that bioinformatics has an important role to play in addressing the complexity of the underlying genetic basis of common human diseases. The goal of this review is to identify and discuss those GWAS challenges that will require computational methods.
Jason H. Moore, Folkert W. Asselbergs, Scott M. Williams
Bioinform.2