Marleen Balvert

dblp:222/1814 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-2376-9301ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2024 Computing linkage disequilibrium aware genome embeddings using autoencoders
abstract
MOTIVATION: The completion of the genome has paved the way for genome-wide association studies (GWAS), which explained certain proportions of heritability. GWAS are not optimally suited to detect non-linear effects in disease risk, possibly hidden in non-additive interactions (epistasis). Alternative methods for epistasis detection using, e.g. deep neural networks (DNNs) are currently under active development. However, DNNs are constrained by finite computational resources, which can be rapidly depleted due to increasing complexity with the sheer size of the genome. Besides, the curse of dimensionality complicates the task of capturing meaningful genetic patterns for DNNs; therefore necessitates dimensionality reduction. RESULTS: We propose a method to compress single nucleotide polymorphism (SNP) data, while leveraging the linkage disequilibrium (LD) structure and preserving potential epistasis. This method involves clustering correlated SNPs into haplotype blocks and training per-block autoencoders to learn a compressed representation of the block's genetic content. We provide an adjustable autoencoder design to accommodate diverse blocks and bypass extensive hyperparameter tuning. We applied this method to genotyping data from Project MinE, and achieved 99% average test reconstruction accuracy-i.e. minimal information loss-while compressing the input to nearly 10% of the original size. We demonstrate that haplotype-block based autoencoders outperform linear Principal Component Analysis (PCA) by approximately 3% chromosome-wide accuracy of reconstructed variants. To the extent of our knowledge, our approach is the first to simultaneously leverage haplotype structure and DNNs for dimensionality reduction of genetic data. AVAILABILITY AND IMPLEMENTATION: Data are available for academic use through Project MinE at https://www.projectmine.com/research/data-sharing/, contingent upon terms and requirements specified by the source studies. Code is available at https://github.com/gizem-tas/haploblock-autoencoders.
Gizem Tas, Timo Westerdijk, Eric O. Postma, Jan Veldink, Alexander Schönhuth, Marleen Balvert
Bioinform.7
2024 Iterative Rule Extension for Logic Analysis of Data: An MILP-Based Heuristic to Derive Interpretable Binary Classifiers from Large Data Sets
abstract
Data-driven decision making is rapidly gaining popularity, fueled by the ever-increasing amounts of available data and encouraged by the development of models that can identify nonlinear input–output relationships. Simultaneously, the need for interpretable prediction and classification methods is increasing as this improves both our trust in these models and the amount of information we can abstract from data. An important aspect of this interpretability is to obtain insight in the sensitivity–specificity trade-off constituted by multiple plausible input–output relationships. These are often shown in a receiver operating characteristic curve. These developments combined lead to the need for a method that can identify complex yet interpretable input–output relationships from large data, that is, data containing large numbers of samples and features. Boolean phrases in disjunctive normal form (DNF) are highly suitable for explaining nonlinear input–output relationships in a comprehensible way. Mixed integer linear programming can be used to obtain these Boolean phrases from binary data though its computational complexity prohibits the analysis of large data sets. This work presents IRELAND, an algorithm that allows for abstracting Boolean phrases in DNF from data with up to 10,000 samples and features. The results show that, for large data sets, IRELAND outperforms the current state of the art in terms of prediction accuracy. Additionally, by construction, IRELAND allows for an efficient computation of the sensitivity–specificity trade-off curve, allowing for further understanding of the underlying input–output relationship. History: Accepted by Andrea Lodi, Area Editor for Design & Analysis of Algorithms–Discrete. Funding: This work was supported by the Netherlands Organization for Scientific Research (Nederlandse Organisatie voor Wetenschappelijk Onderzoek) [Grant VI.VENI.192.043]. Supplemental Material: The online supplement is available at https://doi.org/10.1287/ijoc.2021.0284 .
Marleen Balvert
INFORMS J. Comput.1
2021 OGRE: Overlap Graph-based metagenomic Read clustEring
abstract
MOTIVATION: The microbes that live in an environment can be identified from the combined genomic material, also referred to as the metagenome. Sequencing a metagenome can result in large volumes of sequencing reads. A promising approach to reduce the size of metagenomic datasets is by clustering reads into groups based on their overlaps. Clustering reads are valuable to facilitate downstream analyses, including computationally intensive strain-aware assembly. As current read clustering approaches cannot handle the large datasets arising from high-throughput metagenome sequencing, a novel read clustering approach is needed. In this article, we propose OGRE, an Overlap Graph-based Read clustEring procedure for high-throughput sequencing data, with a focus on shotgun metagenomes. RESULTS: We show that for small datasets OGRE outperforms other read binners in terms of the number of species included in a cluster, also referred to as cluster purity, and the fraction of all reads that is placed in one of the clusters. Furthermore, OGRE is able to process metagenomic datasets that are too large for other read binners into clusters with high cluster purity. CONCLUSION: OGRE is the only method that can successfully cluster reads in species-specific clusters for large metagenomic datasets without running into computation time- or memory issues. AVAILABILITYAND IMPLEMENTATION: Code is made available on Github (https://github.com/Marleen1/OGRE). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marleen Balvert, Ernestina Hauptfeld, Alexander Schönhuth, Bas E. Dutilh
Bioinform.1
2019 Using the structure of genome data in the design of deep neural networks for predicting amyotrophic lateral sclerosis from genotype
abstract
MOTIVATION: Amyotrophic lateral sclerosis (ALS) is a neurodegenerative disease caused by aberrations in the genome. While several disease-causing variants have been identified, a major part of heritability remains unexplained. ALS is believed to have a complex genetic basis where non-additive combinations of variants constitute disease, which cannot be picked up using the linear models employed in classical genotype-phenotype association studies. Deep learning on the other hand is highly promising for identifying such complex relations. We therefore developed a deep-learning based approach for the classification of ALS patients versus healthy individuals from the Dutch cohort of the Project MinE dataset. Based on recent insight that regulatory regions harbor the majority of disease-associated variants, we employ a two-step approach: first promoter regions that are likely associated to ALS are identified, and second individuals are classified based on their genotype in the selected genomic regions. Both steps employ a deep convolutional neural network. The network architecture accounts for the structure of genome data by applying convolution only to parts of the data where this makes sense from a genomics perspective. RESULTS: Our approach identifies potentially ALS-associated promoter regions, and generally outperforms other classification methods. Test results support the hypothesis that non-additive combinations of variants contribute to ALS. Architectures and protocols developed are tailored toward processing population-scale, whole-genome data. We consider this a relevant first step toward deep learning assisted genotype-phenotype association in whole genome-sized data. AVAILABILITY AND IMPLEMENTATION: Our code will be available on Github, together with a synthetic dataset (https://github.com/byin-cwi/ALS-Deeplearning). The data used in this study is available to bona-fide researchers upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bojian Yin, Marleen Balvert, Rick A. A. van der Spek, Bas E. Dutilh, Sander M. Bohté, Jan Veldink, Alexander Schönhuth
Bioinform.2
2019 Robust Optimization of Dose-Volume Metrics for Prostate HDR-Brachytherapy Incorporating Target and OAR Volume Delineation Uncertainties
abstract
In radiation therapy planning, uncertainties in the definition of the target volume yield a risk of underdosing the tumor. The traditional corrective action in the context of external beam radiotherapy (EBRT) expands the clinical target volume (CTV) with an isotropic margin to obtain the planning target volume (PTV). However, the EBRT-based PTV concept is not directly applicable to brachytherapy (BT) since it can lead to undesirable dose escalation. Here, we present a treatment plan optimization model that uses worst-case robust optimization to account for delineation uncertainties in interstitial high-dose-rate BT of the prostate. A scenario-based method was developed that handles uncertainties in index sets. Heuristics were included to reduce the calculation times to acceptable proportions. The approach was extended to account for delineation uncertainties of an organ at risk (OAR) as well. The method was applied on data from prostate cancer patients and evaluated in terms of commonly used dosimetric performance criteria for the CTV and relevant OARs. The robust optimization approach was compared against the classical PTV margin concept and against a scenario-based CTV margin approach. The results show that the scenario-based margin and the robust optimization method are capable of reducing the risk of underdosage to the tumor. As expected, the scenario-based CTV margin approach leads to dose escalation within the target, whereas this can be prevented with the robust model. For cases where rectum sparing was a binding restriction, including uncertainties in rectum delineation in the planning model led to a reduced risk of a rectum overdose, and in some cases, to reduced targetcoverage. The online supplement is available at https://doi.org/10.1287/ijoc.2018.0815 .
Marleen Balvert, Dick den Hertog, Aswin L. Hoffmann
INFORMS J. Comput.1
2018 An image representation based convolutional network for DNA classification
Bojian Yin, Marleen Balvert, Davide Zambrano, Alexander Schönhuth, Sander M. Bohté
ICLR (Poster)2