Marieke L. Kuijjer

dblp:188/6346 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0001-6280-3130ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2025 Gene regulatory network integration with multi-omics data enhances survival predictions in cancer
abstract
The emergence of high-throughput omics technologies has resulted in their wide application to cancer studies, greatly increasing our understanding of the disruptions occurring at different molecular levels. To fully harness these data, integrative approaches have emerged as essential tools, enabling the combination of multiple omics modalities to uncover disease mechanisms. However, many such approaches overlook gene regulatory mechanisms, which play a central role in the development and progression of cancer. Patient-specific gene regulatory networks (GRNs), representing interactions between regulators (such as transcription factors) and their target genes in each individual tumour, offer a powerful framework to bridge this gap and investigate the regulatory landscape of cancer. In this study, we introduce a novel approach for integrating patient-specific GRNs with multi-omic data and assess whether their inclusion in joint dimensionality reduction models improves survival prediction across multiple cancer types. By applying our method on ten cancer datasets from The Cancer Genome Atlas, we demonstrate that incorporating GRNs enhances associations with patient survival in several cancer types. Focusing on liver cancer, with validation in independent data, our methodology identifies potential mechanisms of gene regulatory dysregulation associated with cancer progression. These were linked to dysregulated fatty acid metabolism, and identified JUND as a potential novel transcriptional regulator driving these processes. Our findings highlight the value of network-based multi-omics integration for uncovering clinically relevant regulatory mechanisms and improving our understanding of cancer biology at the patient-specific level.
Romana T. Pop, Ping-Han Hsieh, Tatiana Belova, Anthony Mathelier, Marieke L. Kuijjer
Briefings Bioinform.5
2025 SPONGE: simple prior omics network GEnerator
abstract
SUMMARY: Gene regulatory networks modelled from experimental data can be improved through the use of prior biological knowledge, e.g. transcription factor binding. There are several tools that utilize this information. However, the prior networks used with them are often not updated and may fail to reflect the most up-to-date information. Here we present SPONGE, a Python module designed to access information across biological databases, chiefly JASPAR and STRING, to model two types of networks-a prior gene regulatory network mapping transcription factors to genes based on their predicted binding sites, and a prior protein-protein interaction network mapping potential interactions between transcription factors. SPONGE is mainly designed to work with the PANDA algorithm and the corresponding NetZoo family of tools. However, the networks are provided in an easily adaptable format for other tools. SPONGE was designed with ease of use in mind, and it provides sensible default values for all of its parameters while giving the users the freedom to fine-tune them. AVAILABILITY AND IMPLEMENTATION: The code for the Python module and the documentation can be found in our GitHub repository.
Ladislav Hovan, Marieke L. Kuijjer
Bioinform.2
2024 Improving bioinformatics software quality through teamwork
abstract
SUMMARY: Since high-throughput techniques became a staple in biological science laboratories, computational algorithms, and scientific software have boomed. However, the development of bioinformatics software usually lacks software development quality standards. The resulting software code is hard to test, reuse, and maintain. We believe that the root of inefficiency in implementing the best software development practices in academic settings is the individualistic approach, which has traditionally been the norm for recognizing scientific achievements and, by extension, for developing specialized software. Software development is a collective effort in most software-heavy endeavors. Indeed, the literature suggests teamwork directly impacts code quality through knowledge sharing, collective software development, and established coding standards. In our computational biology research groups, we sustainably involve all group members in learning, sharing, and discussing software development while maintaining the personal ownership of research projects and related software products. We found that group members involved in this endeavor improved their coding skills, became more efficient bioinformaticians, and obtained detailed knowledge about their peers' work, triggering new collaborative projects. We strongly advocate for improving software development culture within bioinformatics through collective effort in computational biology groups or institutes with three or more bioinformaticians. AVAILABILITY AND IMPLEMENTATION: Additional information and guidance on how to get started is available at https://ferenckata.github.io/ImprovingSoftwareTogether.github.io/.
Katalin Ferenc, Ieva Rauluseviciute, Ladislav Hovan, Marieke L. Kuijjer, Anthony Mathelier
Bioinform.5
2023 Adjustment of spurious correlations in co-expression measurements from RNA-Sequencing data
abstract
MOTIVATION: Gene co-expression measurements are widely used in computational biology to identify coordinated expression patterns across a group of samples. Coordinated expression of genes may indicate that they are controlled by the same transcriptional regulatory program, or involved in common biological processes. Gene co-expression is generally estimated from RNA-Sequencing data, which are commonly normalized to remove technical variability. Here, we demonstrate that certain normalization methods, in particular quantile-based methods, can introduce false-positive associations between genes. These false-positive associations can consequently hamper downstream co-expression network analysis. Quantile-based normalization can, however, be extremely powerful. In particular, when preprocessing large-scale heterogeneous data, quantile-based normalization methods such as smooth quantile normalization can be applied to remove technical variability while maintaining global differences in expression for samples with different biological attributes. RESULTS: We developed SNAIL (Smooth-quantile Normalization Adaptation for the Inference of co-expression Links), a normalization method based on smooth quantile normalization specifically designed for modeling of co-expression measurements. We show that SNAIL avoids formation of false-positive associations in co-expression as well as in downstream network analyses. Using SNAIL, one can avoid arbitrary gene filtering and retain associations to genes that only express in small subgroups of samples. This highlights the method's potential future impact on network modeling and other association-based approaches in large-scale heterogeneous data. AVAILABILITY AND IMPLEMENTATION: The implementation of the SNAIL algorithm and code to reproduce the analyses described in this work can be found in the GitHub repository https://github.com/kuijjerlab/PySNAIL.
Ping-Han Hsieh, Camila Miranda Lopes-Ramos, Manuela Zucknick, Geir Kjetil Sandve, Kimberly Glass, Marieke L. Kuijjer
Bioinform.6
2022 rPanglaoDB: an R package to download and merge labeled single-cell RNA-seq data from the PanglaoDB database
abstract
MOTIVATION: Characterizing cells with rare molecular phenotypes is one of the promises of high throughput single-cell RNA sequencing (scRNA-seq) techniques. However, collecting enough cells with the desired molecular phenotype in a single experiment is challenging, requiring several samples preprocessing steps to filter and collect the desired cells experimentally before sequencing. Data integration of multiple public single-cell experiments stands as a solution for this problem, allowing the collection of enough cells exhibiting the desired molecular signatures. By increasing the sample size of the desired cell type, this approach enables a robust cell type transcriptome characterization. RESULTS: Here, we introduce rPanglaoDB, an R package to download and merge the uniformly processed and annotated scRNA-seq data provided by the PanglaoDB database. To show the potential of rPanglaoDB for collecting rare cell types by integrating multiple public datasets, we present a biological application collecting and characterizing a set of 157 fibrocytes. Fibrocytes are a rare monocyte-derived cell type, that exhibits both the inflammatory features of macrophages and the tissue remodeling properties of fibroblasts. This constitutes the first fibrocytes' unbiased transcriptome profile report. We compared the transcriptomic profile of the fibrocytes against the fibroblasts collected from the same tissue samples and confirm their associated relationship with healing processes in tissue damage and infection through the activation of the prostaglandin biosynthesis and regulation pathway. AVAILABILITY AND IMPLEMENTATION: rPanglaoDB is implemented as an R package available through the CRAN repositories https://CRAN.R-project.org/package=rPanglaoDB.
Daniel Osório 0004, Marieke L. Kuijjer, James J. Cai
Bioinform.2
2020 Catalysis Clustering with GAN by Incorporating Domain Knowledge
abstract
Clustering is an important unsupervised learning method with serious challenges when data is sparse and high-dimensional. Generated clusters are often evaluated with general measures, which may not be meaningful or useful for practical applications and domains. Using a distance metric, a clustering algorithm searches through the data space, groups close items into one cluster, and assigns far away samples to different clusters. In many real-world applications, the number of dimensions is high and data space becomes very sparse. Selection of a suitable distance metric is very difficult and becomes even harder when categorical data is involved. Moreover, existing distance metrics are mostly generic, and clusters created based on them will not necessarily make sense to domain-specific applications. One option to address these challenges is to integrate domain-defined rules and guidelines into the clustering process. In this work we propose a GAN-based approach called Catalysis Clustering to incorporate domain knowledge into the clustering process. With GANs we generate catalysts, which are special synthetic points drawn from the original data distribution and verified to improve clustering quality when measured by a domain-specific metric. We then perform clustering analysis using both catalysts and real data. Final clusters are produced after catalyst points are removed. Experiments on two challenging real-world datasets clearly show that our approach is effective and can generate clusters that are meaningful and useful for real-world applications.
Olga Andreeva, Wei Li 0121, Wei Ding 0003, Marieke L. Kuijjer, John Quackenbush, Ping Chen 0001
KDD4
2020 PUMA: PANDA Using MicroRNA Associations
abstract
MOTIVATION: Conventional methods to analyze genomic data do not make use of the interplay between multiple factors, such as between microRNAs (miRNAs) and the messenger RNA (mRNA) transcripts they regulate, and thereby often fail to identify the cellular processes that are unique to specific tissues. We developed PUMA (PANDA Using MicroRNA Associations), a computational tool that uses message passing to integrate a prior network of miRNA target predictions with target gene co-expression information to model genome-wide gene regulation by miRNAs. We applied PUMA to 38 tissues from the Genotype-Tissue Expression project, integrating RNA-Seq data with two different miRNA target predictions priors, built on predictions from TargetScan and miRanda, respectively. We found that while target predictions obtained from these two different resources are considerably different, PUMA captures similar tissue-specific miRNA-target regulatory interactions in the different network models. Furthermore, the tissue-specific functions of miRNAs we identified based on regulatory profiles (available at: https://kuijjer.shinyapps.io/puma_gtex/) are highly similar between networks modeled on the two target prediction resources. This indicates that PUMA consistently captures important tissue-specific miRNA regulatory processes. In addition, using PUMA we identified miRNAs regulating important tissue-specific processes that, when mutated, may result in disease development in the same tissue. AVAILABILITY AND IMPLEMENTATION: PUMA is available in C++, MATLAB and Python on GitHub (https://github.com/kuijjerlab and https://netzoo.github.io/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marieke L. Kuijjer, Maud Fagny, Alessandro Marin, John Quackenbush, Kimberly Glass
Bioinform.1
2019 Machine learning analysis of gene expression data reveals novel diagnostic and prognostic biomarkers and identifies therapeutic targets for soft tissue sarcomas
abstract
Based on morphology it is often challenging to distinguish between the many different soft tissue sarcoma subtypes. Moreover, outcome of disease is highly variable even between patients with the same disease. Machine learning on transcriptome sequencing data could be a valuable new tool to understand differences between and within entities. Here we used machine learning analysis to identify novel diagnostic and prognostic markers and therapeutic targets for soft tissue sarcomas. Gene expression data was used from the Cancer Genome Atlas, the Genotype-Tissue Expression project and the French Sarcoma Group. We identified three groups of tumors that overlap in their molecular profiles as seen with unsupervised t-Distributed Stochastic Neighbor Embedding clustering and a deep neural network. The three groups corresponded to subtypes that are morphologically overlapping. Using a random forest algorithm, we identified novel diagnostic markers for soft tissue sarcoma that distinguished between synovial sarcoma and MPNST, and that we validated using qRT-PCR in an independent series. Next, we identified prognostic genes that are strong predictors of disease outcome when used in a k-nearest neighbor algorithm. The prognostic genes were further validated in expression data from the French Sarcoma Group. One of these, HMMR, was validated in an independent series of leiomyosarcomas using immunohistochemistry on tissue micro array as a prognostic gene for disease-free interval. Furthermore, reconstruction of regulatory networks combined with data from the Connectivity Map showed, amongst others, that HDAC inhibitors could be a potential effective therapy for multiple soft tissue sarcoma subtypes. A viability assay with two HDAC inhibitors confirmed that both leiomyosarcoma and synovial sarcoma are sensitive to HDAC inhibition. In this study we identified novel diagnostic markers, prognostic markers and therapeutic leads from multiple soft tissue sarcoma gene expression datasets. Thus, machine learning algorithms are powerful new tools to improve our understanding of rare tumor entities.
David G. P. van IJzendoorn, Karoly Szuhai, Inge H. Briaire-de Bruijn, Marie Kostine, Marieke L. Kuijjer, Judith V. M. G. Bovée
PLoS Comput. Biol.5
2018 Clustering on Sparse Data in Non-overlapping Feature Space with Applications to Cancer Subtyping
abstract
This paper presents a new algorithm, Reinforced and Informed Network-based Clustering(RINC), for finding unknown groups of similar data objects in sparse and largely non-overlapping feature space where a network structure among features can be observed. Sparse and non-overlapping unlabeled data become increasingly common and available especially in text mining and biomedical data mining. RINC inserts a domain informed model into a modelless neural network. In particular, our approach integrates physically meaningful feature dependencies into the neural network architecture and soft computational constraint. Our learning algorithm efficiently clusters sparse data through integrated smoothing and sparse auto-encoder learning. The informed design requires fewer samples for training and at least part of the model becomes explainable. The architecture of the reinforced network layers smooths sparse data over the network dependency in the feature space. Most importantly, through back-propagation, the weights of the reinforced smoothing layers are simultaneously constrained by the remaining sparse auto-encoder layers that set the target values to be equal to the raw inputs. Empirical results demonstrate that RINC achieves improved accuracy and renders physically meaningful clustering results.
Tianyu Kang, Kourosh Zarringhalam, Marieke L. Kuijjer, Ping Chen 0001, John Quackenbush, Wei Ding 0003
ICDM3
2017 Tissue-aware RNA-Seq processing and normalization for heterogeneous and sparse data
abstract
BACKGROUND: Although ultrahigh-throughput RNA-Sequencing has become the dominant technology for genome-wide transcriptional profiling, the vast majority of RNA-Seq studies typically profile only tens of samples, and most analytical pipelines are optimized for these smaller studies. However, projects are generating ever-larger data sets comprising RNA-Seq data from hundreds or thousands of samples, often collected at multiple centers and from diverse tissues. These complex data sets present significant analytical challenges due to batch and tissue effects, but provide the opportunity to revisit the assumptions and methods that we use to preprocess, normalize, and filter RNA-Seq data - critical first steps for any subsequent analysis. RESULTS: We find that analysis of large RNA-Seq data sets requires both careful quality control and the need to account for sparsity due to the heterogeneity intrinsic in multi-group studies. We developed Yet Another RNA Normalization software pipeline (YARN), that includes quality control and preprocessing, gene filtering, and normalization steps designed to facilitate downstream analysis of large, heterogeneous RNA-Seq data sets and we demonstrate its use with data from the Genotype-Tissue Expression (GTEx) project. CONCLUSIONS: An R package instantiating YARN is available at http://bioconductor.org/packages/yarn .
Joseph N. Paulson, Cho-Yi Chen, Camila Miranda Lopes-Ramos, Marieke L. Kuijjer, John Platig, Abhijeet R. Sonawane, Maud Fagny, Kimberly Glass, John Quackenbush
BMC Bioinform.4
2016 PyPanda: a Python package for gene regulatory network reconstruction
abstract
PANDA (Passing Attributes between Networks for Data Assimilation) is a gene regulatory network inference method that uses message-passing to integrate multiple sources of 'omics data. PANDA was originally coded in C ++. In this application note we describe PyPanda, the Python version of PANDA. PyPanda runs considerably faster than the C ++ version and includes additional features for network analysis. AVAILABILITY AND IMPLEMENTATION: The open source PyPanda Python package is freely available at http://github.com/davidvi/pypanda CONTACT: [email protected] or [email protected].
David G. P. van IJzendoorn, Kimberly Glass, John Quackenbush, Marieke L. Kuijjer
Bioinform.4