Marcel J. T. Reinders

dblp:66/1184 · DBLP profile ↗
← Back
109ranked-venue papers
4as first author
9since 2021 · last 2024
0000-0002-1148-1562ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 53 · 7 since 2021Artificial intelligence and machine learning · 27 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 13 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1Software engineering, systems software and programming languages · 1Theory of computation · 1
YearPublicationVenuePosition
2024 Federated K-Means Clustering
Swier Garst, Marcel J. T. Reinders
ICPR (2)2
2024 PATE: Proximity-Aware Time Series Anomaly Evaluation
abstract
Evaluating anomaly detection algorithms in time series data is critical as inaccuracies can lead to flawed decision-making in various domains where real-time analytics and data-driven strategies are essential. Traditional performance metrics assume iid data and fail to capture the complex temporal dynamics and specific characteristics of time series anomalies, such as early and delayed detections. We introduce Proximity-Aware Time series anomaly Evaluation (PATE), a novel evaluation metric that incorporates the temporal relationship between prediction and anomaly intervals. PATE uses proximity-based weighting considering buffer zones around anomaly intervals, enabling a more detailed and informed assessment of a detection. Using these weights, PATE computes a weighted version of the area under the Precision and Recall curve. Our experiments with synthetic and real-world datasets show the superiority of PATE in providing more sensible and accurate evaluations than other evaluation metrics. We also tested several state-of-the-art anomaly detectors across various benchmark datasets using the PATE evaluation scheme. The results show that a common metric like Point-Adjusted F1 Score fails to characterize the detection performances well, and that PATE is able to provide a more fair model comparison. By introducing PATE, we redefine the understanding of model efficacy that steers future studies toward developing more effective and accurate detection models.
Ramin Ghorbani, Marcel J. T. Reinders, David M. J. Tax
KDD2
2024 An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics
abstract
Multi-omic analyses are necessary to understand the complex biological processes taking place at the tissue and cell level, but also to make reliable predictions about, for example, disease outcome. Several linear methods exist that create a joint embedding using paired information per sample, but recently there has been a rise in the popularity of neural architectures that embed paired -omics into the same non-linear manifold. This work describes a head-to-head comparison of linear and non-linear joint embedding methods using both bulk and single-cell multi-modal datasets. We found that non-linear methods have a clear advantage with respect to linear ones for missing modality imputation. Performance comparisons in the downstream tasks of survival analysis for bulk tumor data and cell type classification for single-cell data lead to the following insights: First, concatenating the principal components of each modality is a competitive baseline and hard to beat if all modalities are available at test time. However, if we only have one modality available at test time, training a predictive model on the joint space of that modality can lead to performance improvements with respect to just using the unimodal principal components. Second, -omic profiles imputed by neural joint embedding methods are realistic enough to be used by a classifier trained on real data with limited performance drops. Taken together, our comparisons give hints to which joint embedding to use for which downstream task. Overall, product-of-experts performed well in most tasks and was reasonably fast, while early integration (concatenation) of modalities did quite poorly.
Stavros Makrodimitris, Bram Pronk, Tamim Abdelaal, Marcel J. T. Reinders
Briefings Bioinform.4
2023 Percolate: An Exponential Family JIVE Model to Design DNA-Based Predictors of Drug Response
abstract
Abstract Motivation: Anti-cancer drugs may elicit resistance or sensitivity through mechanisms which involve several genomic layers. Nevertheless, we have demonstrated that gene expression contains most of the predictive capacity compared to the remaining omic data types. Unfortunately, this comes at a price: gene expression biomarkers are often hard to interpret and show poor robustness. Results: To capture the best of both worlds, i.e. the accuracy of gene expression and the robustness of other genomic levels, such as mutations, copy-number or methylation, we developed Percolate, a computational approach which extracts the joint signal between gene expression and the other omic data types. We developed an out-of-sample extension of Percolate which allows predictions on unseen samples without the necessity to recompute the joint signal on all data. We employed Percolate to extract the joint signal between gene expression and either mutations, copy-number or methylation, and used the out-of sample extension to perform response prediction on unseen samples. We showed that the joint signal recapitulates, and sometimes exceeds, the predictive performance achieved with each data type individually. Importantly, molecular signatures created by Percolate do not require gene expression to be evaluated, rendering them suitable to clinical applications where only one data type is available. Availability: Percolate is available as a Python 3.7 package and the scripts to reproduce the results are available here .
Soufiane Mourragui, Marco Loog, Mirrelijn M. van Nee, Mark A. van de Wiel, Marcel J. T. Reinders, Lodewyk F. A. Wessels
RECOMB5
2023 Cell type matching across species using protein embeddings and transfer learning
abstract
MOTIVATION: Knowing the relation between cell types is crucial for translating experimental results from mice to humans. Establishing cell type matches, however, is hindered by the biological differences between the species. A substantial amount of evolutionary information between genes that could be used to align the species is discarded by most of the current methods since they only use one-to-one orthologous genes. Some methods try to retain the information by explicitly including the relation between genes, however, not without caveats. RESULTS: In this work, we present a model to transfer and align cell types in cross-species analysis (TACTiCS). First, TACTiCS uses a natural language processing model to match genes using their protein sequences. Next, TACTiCS employs a neural network to classify cell types within a species. Afterward, TACTiCS uses transfer learning to propagate cell type labels between species. We applied TACTiCS on scRNA-seq data of the primary motor cortex of human, mouse, and marmoset. Our model can accurately match and align cell types on these datasets. Moreover, our model outperforms Seurat and the state-of-the-art method SAMap. Finally, we show that our gene matching method results in better cell type matches than BLAST in our model. AVAILABILITY AND IMPLEMENTATION: The implementation is available on GitHub (https://github.com/kbiharie/TACTiCS). The preprocessed datasets and trained models can be downloaded from Zenodo (https://doi.org/10.5281/zenodo.7582460).
Kirti Biharie, Lieke Michielsen, Marcel J. T. Reinders, Ahmed Mahfouz
Bioinform.3
2023 Determining epitope specificity of T-cell receptors with transformers
abstract
SUMMARY: T-cell receptors (TCRs) on T cells recognize and bind to epitopes presented by the major histocompatibility complex in case of an infection or cancer. However, the high diversity of TCRs, as well as their unique and complex binding mechanisms underlying epitope recognition, make it difficult to predict the binding between TCRs and epitopes. Here, we present the utility of transformers, a deep learning strategy that incorporates an attention mechanism that learns the informative features, and show that these models pre-trained on a large set of protein sequences outperform current strategies. We compared three pre-trained auto-encoder transformer models (ProtBERT, ProtAlbert, and ProtElectra) and one pre-trained auto-regressive transformer model (ProtXLNet) to predict the binding specificity of TCRs to 25 epitopes from the VDJdb database (human and murine). Two additional modifications were performed to incorporate gene usage of the TCRs in the four transformer models. Of all 12 transformer implementations (four models with three different modifications), a modified version of the ProtXLNet model could predict TCR-epitope pairs with the highest accuracy (weighted F1 score 0.55 simultaneously considering all 25 epitopes). The modification included additional features representing the gene names for the TCRs. We also showed that the basic implementation of transformers outperformed the previously available methods, i.e. TCRGP, TCRdist, and DeepTCR, developed for the same biological problem, especially for the hard-to-classify labels. We show that the proficiency of transformers in attention learning can be made operational in a complex biological setting like TCR binding prediction. Further ingenuity in utilizing the full potential of transformers, either through attention head visualization or introducing additional features, can extend T-cell research avenues. AVAILABILITY AND IMPLEMENTATION: Data and code are available on https://github.com/InduKhatri/tcrformer.
Abdul Rehman Khan, Marcel J. T. Reinders, Indu Khatri
Bioinform.2
2022 MiMIR: R-shiny application to infer risk factors and endpoints from Nightingale Health's 1H-NMR metabolomics data
abstract
MOTIVATION: 1H-NMR metabolomics is rapidly becoming a standard resource in large epidemiological studies to acquire metabolic profiles in large numbers of samples in a relatively low-priced and standardized manner. Concomitantly, metabolomics-based models are increasingly developed that capture disease risk or clinical risk factors. These developments raise the need for user-friendly toolbox to inspect new 1H-NMR metabolomics data and project a wide array of previously established risk models. RESULTS: We present MiMIR (Metabolomics-based Models for Imputing Risk), a graphical user interface that provides an intuitive framework for ad hoc statistical analysis of Nightingale Health's 1H-NMR metabolomics data and allows for the projection and calibration of 24 pre-trained metabolomics-based models, without any pre-required programming knowledge. AVAILABILITY AND IMPLEMENTATION: The R-shiny package is available in CRAN or downloadable at https://github.com/DanieleBizzarri/MiMIR, together with an extensive user manual (also available as Supplementary Documents to the article). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniele Bizzarri, Marcel J. T. Reinders, Marian Beekman, P. Eline Slagboom, Erik van den Akker 0001
Bioinform.2
2022 A framework for employing longitudinally collected multicenter electronic health records to stratify heterogeneous patient populations on disease history
abstract
OBJECTIVE: To facilitate patient disease subset and risk factor identification by constructing a pipeline which is generalizable, provides easily interpretable results, and allows replication by overcoming electronic health records (EHRs) batch effects. MATERIAL AND METHODS: We used 1872 billing codes in EHRs of 102 880 patients from 12 healthcare systems. Using tools borrowed from single-cell omics, we mitigated center-specific batch effects and performed clustering to identify patients with highly similar medical history patterns across the various centers. Our visualization method (PheSpec) depicts the phenotypic profile of clusters, applies a novel filtering of noninformative codes (Ranked Scope Pervasion), and indicates the most distinguishing features. RESULTS: We observed 114 clinically meaningful profiles, for example, linking prostate hyperplasia with cancer and diabetes with cardiovascular problems and grouping pediatric developmental disorders. Our framework identified disease subsets, exemplified by 6 "other headache" clusters, where phenotypic profiles suggested different underlying mechanisms: migraine, convulsion, injury, eye problems, joint pain, and pituitary gland disorders. Phenotypic patterns replicated well, with high correlations of ≥0.75 to an average of 6 (2-8) of the 12 different cohorts, demonstrating the consistency with which our method discovers disease history profiles. DISCUSSION: Costly clinical research ventures should be based on solid hypotheses. We repurpose methods from single-cell omics to build these hypotheses from observational EHR data, distilling useful information from complex data. CONCLUSION: We establish a generalizable pipeline for the identification and replication of clinically meaningful (sub)phenotypes from widely available high-dimensional billing codes. This approach overcomes datatype problems and produces comprehensive visualizations of validation-ready phenotypes.
Marc P. Maurits, Ilya Korsunsky, Soumya Raychaudhuri, Shawn N. Murphy, Jordan W. Smoller, Scott T. Weiss, Thomas W. J. Huizinga, Marcel J. T. Reinders, Elizabeth W. Karlson, Erik van den Akker 0001, Rachel Knevel
J. Am. Medical Informatics Assoc.8
2021 Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function
abstract
MOTIVATION: Protein function prediction is a difficult bioinformatics problem. Many recent methods use deep neural networks to learn complex sequence representations and predict function from these. Deep supervised models require a lot of labeled training data which are not available for this task. However, a very large amount of protein sequences without functional labels is available. RESULTS: We applied an existing deep sequence model that had been pretrained in an unsupervised setting on the supervised task of protein molecular function prediction. We found that this complex feature representation is effective for this task, outperforming hand-crafted features such as one-hot encoding of amino acids, k-mer counts, secondary structure and backbone angles. Also, it partly negates the need for complex prediction models, as a two-layer perceptron was enough to achieve competitive performance in the third Critical Assessment of Functional Annotation benchmark. We also show that combining this sequence representation with protein 3D structure information does not lead to performance improvement, hinting that 3D structure is also potentially learned during the unsupervised pretraining. AVAILABILITY AND IMPLEMENTATION: Implementations of all used models can be found at https://github.com/stamakro/GCN-for-Structure-and-Function. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Amelia Villegas-Morcillo, Stavros Makrodimitris, Roeland C. H. J. van Ham, Ángel M. Gómez, Victoria E. Sánchez, Marcel J. T. Reinders
Bioinform.6
2020 PREDICT: Efficient Private Disease Susceptibility Testing in Direct-to-Consumer Model
abstract
Genome sequencing has rapidly advanced in the last decade, making it easier for anyone to obtain digital genomes at low costs from companies such as Helix, MyHeritage, and 23andMe. Companies now offer their services in a direct-to-consumer (DTC) model without the intervention of a medical institution. Thereby, providing people with direct services for paternity testing, ancestry testing and disease susceptibility testing (DST) to infer diseases' predisposition. Genome analyses are partly motivated by curiosity and people often want to partake without fear of privacy invasion. Existing privacy protection solutions for DST adopt cryptographic techniques to protect the genome of a patient from the party responsible for computing the analysis. Said techniques include homomorphic encryption, which can be computationally expensive and could take minutes for only a few single-nucleotide polymorphisms (SNPs). A predominant approach is a solution that computes DST over encrypted data, but the design depends on a medical unit and exposes test results of patients to the medical unit, making the design uncomfortable for privacy-aware individuals. Hence it is pertinent to have an efficient privacy-preserving DST solution with a DTC service. We propose a novel DTC model that protects the privacy of SNPs and prevents leakage of test results to any other party save for the genome owner. Conversely, we protect the privacy of the algorithms or trade secrets used by the genome analyzing companies. Our work utilizes a secure obfuscation technique in computing DST, eliminating expensive computations over encrypted data. Our approach significantly outperforms existing state-of-the-art solutions in runtime and scales linearly for equivalent levels of security. As an example, computing DST for 10,000 SNPs requires approximately 96 milliseconds on commodity hardware. With this efficient and privacy-preserving solution which is also simulation-based secure, we open possibilities for performing genome analyses on collectively shared data resources.
Chibuike Ugwuoke, Zekeriya Erkin, Marcel J. T. Reinders, Reginald L. Lagendijk
CODASPY3
2020 SCHNEL: scalable clustering of high dimensional single-cell data
abstract
MOTIVATION: Single cell data measures multiple cellular markers at the single-cell level for thousands to millions of cells. Identification of distinct cell populations is a key step for further biological understanding, usually performed by clustering this data. Dimensionality reduction based clustering tools are either not scalable to large datasets containing millions of cells, or not fully automated requiring an initial manual estimation of the number of clusters. Graph clustering tools provide automated and reliable clustering for single cell data, but suffer heavily from scalability to large datasets. RESULTS: We developed SCHNEL, a scalable, reliable and automated clustering tool for high-dimensional single-cell data. SCHNEL transforms large high-dimensional data to a hierarchy of datasets containing subsets of data points following the original data manifold. The novel approach of SCHNEL combines this hierarchical representation of the data with graph clustering, making graph clustering scalable to millions of cells. Using seven different cytometry datasets, SCHNEL outperformed three popular clustering tools for cytometry data, and was able to produce meaningful clustering results for datasets of 3.5 and 17.2 million cells within workable time frames. In addition, we show that SCHNEL is a general clustering tool by applying it to single-cell RNA sequencing data, as well as a popular machine learning benchmark dataset MNIST. AVAILABILITY AND IMPLEMENTATION: Implementation is available on GitHub (https://github.com/biovault/SCHNELpy). All datasets used in this study are publicly available. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tamim Abdelaal, Paul de Raadt, Boudewijn P. F. Lelieveldt, Marcel J. T. Reinders, Ahmed Mahfouz
Bioinform.4
2020 ImSpectR: R package to quantify immune repertoire diversity in spectratype and repertoire sequencing data
abstract
SUMMARY: An effective immune system is characterized by a diverse immune repertoire. There is a strong demand for accurate and quantitative methods to assess the diversity of the immune repertoire for various (pre-)clinical applications, including the diagnosis and prognosis of primary immune deficiencies, or to assess the response to therapy. Current strategies for immune diversity assessment generally comprise the visual inspection of the length distribution of rearranged T- and B-cell receptors. Visual inspections, however, are prone to subjective assessments and thus lead to biases. Here, we introduce ImSpectR, a unified approach to quantify immunodiversity using either spectratype, repertoire sequencing or single cell RNA sequencing data. ImSpectR scores various types of deviations from the expected length distribution and integrates these into one measure, allowing for robust quantitative comparisons of immune diversity across individuals or conditions. AVAILABILITY: R-package is available for download on GitHub at https://github.com/martijn-cordes/ImSpectR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martijn Cordes, Karin Pike-Overzet, Marja van Eggermond, Sandra Vloemans, Miranda R. Baert, Laura Garcia-Perez, Frank J. T. Staal, Marcel J. T. Reinders, Erik van den Akker 0001
Bioinform.8
2020 Metric learning on expression data for gene function prediction
abstract
MOTIVATION: Co-expression of two genes across different conditions is indicative of their involvement in the same biological process. However, when using RNA-Seq datasets with many experimental conditions from diverse sources, only a subset of the experimental conditions is expected to be relevant for finding genes related to a particular Gene Ontology (GO) term. Therefore, we hypothesize that when the purpose is to find similarly functioning genes, the co-expression of genes should not be determined on all samples but only on those samples informative for the GO term of interest. RESULTS: To address this, we developed Metric Learning for Co-expression (MLC), a fast algorithm that assigns a GO-term-specific weight to each expression sample. The goal is to obtain a weighted co-expression measure that is more suitable than the unweighted Pearson correlation for applying Guilt-By-Association-based function predictions. More specifically, if two genes are annotated with a given GO term, MLC tries to maximize their weighted co-expression and, in addition, if one of them is not annotated with that term, the weighted co-expression is minimized. Our experiments on publicly available Arabidopsis thaliana RNA-Seq data demonstrate that MLC outperforms standard Pearson correlation in term-centric performance. Moreover, our method is particularly good at more specific terms, which are the most interesting. Finally, by observing the sample weights for a particular GO term, one can identify which experiments are important for learning that term and potentially identify novel conditions that are relevant, as demonstrated by experiments in both A. thaliana and Pseudomonas Aeruginosa. AVAILABILITY AND IMPLEMENTATION: MLC is available as a Python package at www.github.com/stamakro/MLC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Stavros Makrodimitris, Marcel J. T. Reinders, Roeland C. H. J. van Ham
Bioinform.2
2019 Bioinformatics in the Netherlands: the value of a nationwide community
abstract
This review provides a historical overview of the inception and development of bioinformatics research in the Netherlands. Rooted in theoretical biology by foundational figures such as Paulien Hogeweg (at Utrecht University since the 1970s), the developments leading to organizational structures supporting a relatively large Dutch bioinformatics community will be reviewed. We will show that the most valuable resource that we have built over these years is the close-knit national expert community that is well engaged in basic and translational life science research programmes. The Dutch bioinformatics community is accustomed to facing the ever-changing landscape of data challenges and working towards solutions together. In addition, this community is the stable factor on the road towards sustainability, especially in times where existing funding models are challenged and change rapidly.
Celia W. G. van Gelder, Rob W. W. Hooft, Merlijn N. van Rijswijk, Linda van den Berg, Ruben G. Kok, Marcel J. T. Reinders, Barend Mons, Jaap Heringa
Briefings Bioinform.6
2019 CyTOFmerge: integrating mass cytometry data across multiple panels
abstract
MOTIVATION: High-dimensional mass cytometry (CyTOF) allows the simultaneous measurement of multiple cellular markers at single-cell level, providing a comprehensive view of cell compositions. However, the power of CyTOF to explore the full heterogeneity of a biological sample at the single-cell level is currently limited by the number of markers measured simultaneously on a single panel. RESULTS: To extend the number of markers per cell, we propose an in silico method to integrate CyTOF datasets measured using multiple panels that share a set of markers. Additionally, we present an approach to select the most informative markers from an existing CyTOF dataset to be used as a shared marker set between panels. We demonstrate the feasibility of our methods by evaluating the quality of clustering and neighborhood preservation of the integrated dataset, on two public CyTOF datasets. We illustrate that by computationally extending the number of markers we can further untangle the heterogeneity of mass cytometry data, including rare cell-population detection. AVAILABILITY AND IMPLEMENTATION: Implementation is available on GitHub (https://github.com/tabdelaal/CyTOFmerge). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tamim Abdelaal, Thomas Höllt, Vincent van Unen, Boudewijn P. F. Lelieveldt, Frits Koning, Marcel J. T. Reinders, Ahmed Mahfouz
Bioinform.6
2019 Improving protein function prediction using protein sequence and GO-term similarities
abstract
MOTIVATION: Most automatic functional annotation methods assign Gene Ontology (GO) terms to proteins based on annotations of highly similar proteins. We advocate that proteins that are less similar are still informative. Also, despite their simplicity and structure, GO terms seem to be hard for computers to learn, in particular the Biological Process ontology, which has the most terms (>29 000). We propose to use Label-Space Dimensionality Reduction (LSDR) techniques to exploit the redundancy of GO terms and transform them into a more compact latent representation that is easier to predict. RESULTS: We compare proteins using a sequence similarity profile (SSP) to a set of annotated training proteins. We introduce two new LSDR methods, one based on the structure of the GO, and one based on semantic similarity of terms. We show that these LSDR methods, as well as three existing ones, improve the Critical Assessment of Functional Annotation performance of several function prediction algorithms. Cross-validation experiments on Arabidopsis thaliana proteins pinpoint the superiority of our GO-aware LSDR over generic LSDR. Our experiments on A.thaliana proteins show that the SSP representation in combination with a kNN classifier outperforms state-of-the-art and baseline methods in terms of cross-validated F-measure. AVAILABILITY AND IMPLEMENTATION: Source code for the experiments is available at https://github.com/stamakro/SSP-LSDR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Stavros Makrodimitris, Roeland C. H. J. van Ham, Marcel J. T. Reinders
Bioinform.3
2019 PRECISE: a domain adaptation approach to transfer predictors of drug response from pre-clinical models to tumors
abstract
MOTIVATION: Cell lines and patient-derived xenografts (PDXs) have been used extensively to understand the molecular underpinnings of cancer. While core biological processes are typically conserved, these models also show important differences compared to human tumors, hampering the translation of findings from pre-clinical models to the human setting. In particular, employing drug response predictors generated on data derived from pre-clinical models to predict patient response remains a challenging task. As very large drug response datasets have been collected for pre-clinical models, and patient drug response data are often lacking, there is an urgent need for methods that efficiently transfer drug response predictors from pre-clinical models to the human setting. RESULTS: We show that cell lines and PDXs share common characteristics and processes with human tumors. We quantify this similarity and show that a regression model cannot simply be trained on cell lines or PDXs and then applied on tumors. We developed PRECISE, a novel methodology based on domain adaptation that captures the common information shared amongst pre-clinical models and human tumors in a consensus representation. Employing this representation, we train predictors of drug response on pre-clinical data and apply these predictors to stratify human tumors. We show that the resulting domain-invariant predictors show a small reduction in predictive performance in the pre-clinical domain but, importantly, reliably recover known associations between independent biomarkers and their companion drugs on human tumors. AVAILABILITY AND IMPLEMENTATION: PRECISE and the scripts for running our experiments are available on our GitHub page (https://github.com/NKI-CCB/PRECISE). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Soufiane Mourragui, Marco Loog, Mark A. van de Wiel, Marcel J. T. Reinders, Lodewyk F. A. Wessels
Bioinform.4
2018 Bioinformatics in the Netherlands: the value of a nationwide community
abstract
Briefings in Bioinformatics, 2017. https://doi.org/10.1093/bib/bbx087 The authors have corrected the acknowledgements section to include Antoine van Kampen. The section has been corrected online and now reads as follows: Huge thanks are due to Gert Vriend, Jacob de Vlieg and Bob Hertzberger for securing funding for NBIC, initiating its activities, and for leading NBIC during its initial years. We are indebted in the same vein to Antoine van Kampen, who chaired and directed NBIC from 2006 to 2010. The authors would also like to acknowledge the Dutch bioinformatics community at large for its commitment and energy to drive the field further.
Celia W. G. van Gelder, Rob W. W. Hooft, Merlijn N. van Rijswijk, Linda van den Berg, Ruben G. Kok, Marcel J. T. Reinders, Barend Mons, Jaap Heringa
Briefings Bioinform.6
2016 ECCB 2016: The 15th European Conference on Computational Biology
abstract
This special issue includes the proceeding papers accepted for presentation at the 15th European Conference on Computational Biology (ECCB 2016), to be held from September 3 to 7, 2016 at the World Forum Convention Center in The Hague, The Netherlands. Details of the conference are available on the conference web site (www.eccb2016.org) and will later be archived at eccb.iscb.org/2016/. ECCB is the premier European conference in computational biology and bioinformatics, and together with ISMB (Intelligent Systems in Molecular Biology) and RECOMB (Research in Computational Molecular Biology), it is one of the major international conference series in this domain. The ECCB conferences are gathering about a thousand scientists and industry staff working at the intersection of a broad range of disciplines including computer science, mathematics, biology and medicine. New challenges are now emerging in these fields with the recent advances in low-cost ultra-fast sequencing, bio-imaging and big data. Computational analysis platforms are challenged by the enormous complexity of biological systems and the sheer amount of data resulting from high-throughput measuring techniques, such as single-cell or single-molecule measurements. As a consequence, databases and software are evolving rapidly, and new algorithms are required to improve computational analyses of massive biological or biomedical datasets. Recent advances are presented at the conference in the field of data interoperability and machine learning, in particular ‘deep learning’ and network-based analysis techniques. The impact on the field of public data repositories such as TCGA (Weinstein et al., 2013), ENCODE (ENCODE Project Consortium, 2012) and GDSC (Yang et al., 2013) is palpable with many submissions revolving around these resources. ECCB is held annually in a different country, while it is held jointly with the ISMB conference biennially. Going back in time, the fourteen previous editions of ECCB have been held in: Dublin, Ireland, together with ISMB (Moreau and Beerenwinkel, 2015); Strasbourg, France (Devignes, 2014); Berlin, Germany, together with ISMB (Ben-Tal, 2013); Basel, Switzerland (Schwede and Iber, 2012); Vienna, Austria, together with ISMB (Gaasterland and Vingron, 2011); Ghent, Belgium (Moreau and Heringa, 2010); Stockholm, Sweden, together with ISMB (Gusfield and Tramontano, 2009); Cagliari, Italy (Tramontano, 2008); Vienna, Austria, together with ISMB (Lengauer et al., 2007); Eilat, Israel (Wolfson and Safer, 2007); Madrid, Spain (Guigo et al., 2005); Glasgow, United Kingdom, together with ISMB (Thornton et al., 2004); Paris, France (Lenhof and Sagot, 2003); and Saarbrücken, Germany (Lengauer, 2002). The ECCB 2016 edition features keynote lectures by distinguished speakers. The opening keynote lecture will be delivered by the 2013 Breakthrough Prize-winner Hans Clevers (Hubrecht Institute and Princess Maxima Centre for Pediatric Oncology, Utrecht, The Netherlands). Further keynote presentations will be given by Amos Tanay (Weizmann Institute, Tel Aviv, Israel), John Marioni (EMBL-EBI, Hinxton, UK), Nuria Lopez-Bigas (Universitat Pompeu Fabra, Barcelona, Spain), Benedict Paten (UCSC, Santa Cruz, USA), Pauline Hogeweg (Utrecht University, Utrecht, The Netherlands) and Christina Leslie (Memorial Sloan Kettering Cancer Center, New York, USA). The conference topics span all areas of methodological developments for computational biology and innovative applications of computational methods to molecular biology and biomedicine. To present a more unified view of where the science has gone over recent years, new to ECCB this year is that the conference presentations are divided over five broad themes: (i) Data (organization, management, categorization, integration, analysis of data, knowledge discovery); (ii) Genome (sequence analysis, alignment, evolution, phylogeny, genetics, epigenetics, 3D conformation); (iii) Genes (expression, function, regulation, transcription, translation, geno/phenotype); (iv) Proteins (structure, function, alterations, assemblies, interactions, design, proteomics) and (v) Systems (systems biology, pathways, molecular networks, dynamics, signalling, multi-scale modelling). This year four different tracks were created for ECCB: two scientific and two application-oriented ones. On the scientific side, the Proceedings Track presents novel scientific contributions, while the Highlights Track showcases already published high-impact science in computational biology. These two tracks were coordinated and managed across the five themes by a board of 5 × 4 = 20 co-chairs, both overseen by a dedicated track chair. A novelty of ECCB this year is that we have given a prominent platform to applications in the new Application and ELIXIR tracks. The Application Track is an initiative to promote application of computational biology in industry and other fields beyond academia. Submissions to this track should cross the boundaries of traditional academic science, or show developments that are directly relevant beyond academia or have potential for it. Consequently, submissions relating to (pure) research were deemed out of scope. The ELIXIR Track, running for the first time at ECCB 2016, is managed by ELIXIR and focuses on developments relating to services and infrastructure within the ELIXIR nodes. ELIXIR is the pan-European life science infrastructural network (‘Data for the Life Sciences’). It coordinates, integrates and sustains bioinformatics resources across its member states, enabling users in academia and industry to access vital data, tools, standards, compute and training services for research. ELIXIR has chosen ECCB as their major dissemination platform, acting as co-organising sponsor. Following the call for Proceedings papers, we received 150 submissions. Submission authors were asked to rank the five themes for fit of their paper. These rankings were subsequently used to assign papers to themes and to arrive at an optimally balanced distribution of submissions over the themes. Within each theme, the theme co-chairs assigned papers to expert referees, taking care to avoid any conflict of interest. Together, the Programme Committee (PC) was composed of 217 reviewers and 35 co-reviewers. The reviewing and selection process was carried out using the EasyChair multi-track conference reviewing system (www.easychair.org). The review form explicitly differentiated between impact and suitability for ECCB on the one hand, and scientific quality and reproducibility on the other. This distinction was intended to create more clarity for both the reviewers and authors; the reviewing criteria were therefore explicitly stated in the submission guidelines. The added focus on reproducibility resulted in many authors opting to make their methods open source. After the reviewers reached a consensus, a final ranking and selection for each theme was carried out by the theme (co-)chairs. A total of 48 papers were conditionally accepted (acceptance ratio of 32%) based on a predefined number of acceptances per theme based on the distribution of initial assignments over the themes. The authors had two weeks to modify their papers according to the suggestions made by the reviewers and to respond to the reviewers’ comments, which was checked by the theme (co-)chairs. We thank the authors for incorporating these suggestions as they were given little time to carry out (minor) revisions. Essentially, these efforts contribute to the success and reputation of the ECCB conference! It is worth noting that many authors expressed their gratitude to the reviewers for their comments and suggestions. We also gratefully acknowledge the hard and diligent work performed by the reviewers over a short period of time. We believe that for all of the rejected submissions, the reviewers provided high-quality reports. We hope that authors of rejected manuscripts will benefit from these remarks in their future research. We thank all theme co-chairs for their availability throughout the reviewing process, for their very positive attitude in the final selection and for their valuable help in re-examining the modified submissions. The 48 accepted papers are included in this special issue. The Proceedings Track papers with their supplementary files are available free-for-view in electronic format from Oxford’s press journal Bioinformatics from September 1st, 2016. Highlight presentations were introduced at ISMB/ECCB 2007 in Vienna (Lengauer et al., 2007), and immediately became one of the most popular features of the conference. All original research papers that had been published in peer-review journals between 1 March, 2015, and the submission deadline of 29 March 2016, were eligible to be presented as a Highlight talk. After thorough consideration and discussion, the ECCB 2016 theme (co-)chairs selected 24 proposals out of 71 submissions, mainly on the criteria of compatibility with the ECCB objectives, wide impact in the life sciences and the potential for attracting a large audience to the conference. The Applications Track features 11 presentations, which were selected out of 31 submissions. At the time of writing, three Sponsored Talks will be delivered as part of the Applications Track by respectively The Hyve, Keygene and Data Computing. Sponsored Talks enable sponsors to showcase their innovations in computational biology and to highlight their scientific value. Finally, the ELIXIR track proved to be a very popular addition, with 50 high-quality submissions for only 12 presentation slots. Submissions to the Poster Track were evaluated based on a 250-word abstract and will be shown at the conference along the five main conference themes. In line with the new Application and ELIXIR tracks, there will also be special Application and ELIXIR poster tracks. A dedicated Education Poster Track was created to devote in-depth attention to education that is so crucial for the next generation of bioinformaticians. All poster abstracts are available on the conference web site. We also arranged with F1000Research (http://f1000research.com/) to publish the posters via a new ECCB2016 channel; submission is on a voluntary basis. At least sixteen exhibitor booths will be open throughout the conference in the central conference hall, which is well connected to the other activities at ECCB: The Hyve, EMBL-EBI, TimeLogic, Springer, ISCB Student Council, Goblet, ELIXIR, ISCB, Oxford University Press, ENPICOM, Data Computing, CRC Press, SIB, ELIXIR Denmark and Cambridge University Press. They will be presenting the latest scientific literature in the field of computational biology, bioinformatics, data stewardship, modeling and simulation, as well as new hardware, software and technology developments. The Hague Tourist Office, and the four organizing institutions (DTL, Netherlands Bioinformatics and System Biology Research School (BioSB), VU University Amsterdam and Delft University of Technology) will also be represented. During the weekend before the conference, a satellite meeting, 14 workshops and nine tutorials will take place. The Student Council of the International Society for Computational Biology (ISCB) organizes its 4th European Student Council Symposium (ESCS), chaired by Annika Jacobsen (Vrije Universiteit Amsterdam) and Kevin Schwahn (Universität Potsdam, Max Planck Institute of Molecular Plant Physiology). ESCS highlights will be published in F1000Research via the ISCB Student Council channel. The ECCB 2016 organizing committee congratulates all these dynamic young scientists for their enthusiasm, which is essential to the future of research in computational biology. The 14 workshops preceding the ECCB 2016 main meeting were selected out of a total of 29 applications, showing how popular ECCB has become as a venue for dissemination of computational biology research. The workshops, running all but one for a single day, provide participants with an informal setting to discuss technical issues, exchange research ideas, and to share practical experiences on a range of focused or emerging topics in computational biology. Taken together, the workshops demonstrate how extensively technologies have found their way into large-scale practical applications: • (W1) The 10th International Workshop on Machine Learning in Systems Biology, organised by (Juho Rousu, Aalto University, Finland), Dick de Ridder Wageningen University, The Netherlands), Harri Lähdesmäki (Aalto University, Finland) and Aalt—Jan van Dijk (Wageningen University, The Netherlands), is a two-day workshop with its own proceedings track, where accepted papers will be published in BMC Bioinformatics, • (W2) Network Inference: New Methods and New Data, organised by Anagha Joshi (Roslin Institute, University of Edinburgh, UK), Tom Michoel (Roslin Institute, University of Edinburgh, UK) and Eric Bonnet (Centre National de Génotypage, CEA, Paris, France) • (W3) Getting the Most out of Your Methods and Algorithms: A Workshop on How to Use Existing Datasets to Gain Novel Insight, chaired by Morris Swertz and Lude Franke (both at University Medical Centre Groningen, The Netherlands), • (W4) Digital Pathology Meets Bioinformatics, organised by Yves Sucaet (Vrije Universiteit Brussel, Belgium), Jeroen Van der Laak (UMC Radboud, Nijmegen, The Netherlands), Marius Nap (HistoGeneX, Belgium and Rigshospitalet Copenhagen, Denmark), Zev Leifer (New York College of Podiatric Medicine, USA), Yukako Yagi (Harvard Medical School, Cambridge and Massachusetts General Hospital, Boston, USA) and Raphaël Marée (Université de Liège, Belgium), • (W5) RepSeq 2016: Immune Repertoire Sequencing—Bioinformatics and Applications in Hematology and Immunology, organised by Jack Bartram (University College London, UK), Eva Froňková (Charles University Prague, Czech Republic), Mathieu Giraud, (CNRS, Lille, France—program co-chair), Peter N. Robinson (Charité Berlin, Germany), Mikaël Salson (Université de Lille, France—proceedings chair), Mikhail Shugay (Shemyakin and Ovchinnikov Institute of Bioorganic Chemistry, Moscow, Russia—program co-chair) and Andrew P. Stubbs (Erasmus MC, Rotterdam, The Netherlands), • (W6) FAIR Data and Data Stewardship, organised by Erik Schultes, Mark Thompson Marco Roos (all three at Leiden University Medical Centre, The Netherlands), Mark Wilkinson (Universidad Politecnica de Madrid, Spain), Luiz Olavo Bonino da Silva Santos (Dutch Techcentre for Life Sciences and Vrije Universiteit Amsterdam, The Netherlands), • (W7) Challenges and Approaches in Comprehensive and Informative Complex Network Analysis for Precision Medicine, organised by Igor Jurisica (University of Toronto, Canada), Natasa Przulj (University College London, UK) and Tijana Milenkovic (University of Notre Dame, USA), • (W8) Computing a Tissue: Modeling Multicellular Systems, organised by Walter de Back (TU Dresden, Germany), Sara Montagna (University of Bologna, Italy) and Roeland Merks (Center for Mathematics and Computer Science (CWI), Amsterdam and Leiden University, The Netherlands), • (W9) Computational Pan-Genomics, organised by Zamin Iqbal (University of Oxford, UK), Tobias Marschall (Saarland University and Max Planck Institute for Informatics, Saarbrücken, Germany) and Benedict Paten (3UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA), • (W10) BioNetVisA: from Biological Network Reconstruction to data visualisation and analysis in Molecular Biology and Medicine, organised by Inna Kuperstein, Emmanuel Barillot, Andrei Zinovyev (all three at Institut Curie, Paris, France), Hiroaki Kitano (Okinawa Institute of Science and Technology Graduate University, RIKEN Center for Integrative Medical Sciences, Japan), Minoru Kanehisa (Kyoto University, Japan), Samik Ghosh (Systems Biology Institute, Tokyo, Japan), Nicolas Le Novère (Babraham Institute, Cambridge, UK), Robin Haw (Ontario Institute for Cancer Research, Canada), Alfonso Valencia (Spanish National Bioinformatics Institute, Madrid, Spain) and Lodewyk Wessels (Netherlands Cancer Institute, Amsterdam, and Technical University Delft, The Netherlands), • (W11) Recent Computational Advances in Metagenomics, organised by Sophie Schbath, Valentin Loux and Mahendra Mariadassou (all at INRA, Jouy-en-Josas, France), • (W12) BioExcel: Advanced Simulations for Biomolecular Research, a SIG workshop organised by Rossen Apostolov (KTH, Stockholm, Sweden)), Alexandre Bonvin (Utrecht University, The Netherlands), Cath Brooksbank (EMBL-EBI, Hinxton, UK) and Ian Harrow (Ian Harrow Consulting, Whitstable, UK), • (W13) Clinical Bioinformatics as a Service, organised by Niko Beerenwinkel (ETH Zurich, and SIB Swiss Institute of Bioinformatics, Switzerland), Wolfgang Huber (EMBL, Heidelberg), Simon Tavaré (Cancer Research UK Cambridge Institute, UK) and Daniel Stekhoven and SIB Swiss Institute of Bioinformatics, Switzerland), • Computational Challenges of Data organised by University Medical Center, The Netherlands), (UMC Utrecht, The Netherlands) and Hans The Netherlands). • Data Analysis with and organised by Jeroen and The Netherlands), • Genome organised by (University of Spain) and (University of Spain), • and organised by (University of Spain), • Analysis and with organised by United • and for Modeling Biological Systems, organised by (University of Germany) and (University of Germany), • Analysis and for Data, organised by Medical Center, Amsterdam and University of Amsterdam, The Netherlands), • A to Research on organised by and (both at University Pompeu Fabra, Barcelona, Spain), • and their in Genome Data, organised by N. (University of Cambridge, UK), University of Science and and (University of UK), • into and its Applications in Bioinformatics and in Network organised by and (both at Hans Institute, and University Hospital, It is to thank all the and that are ECCB 2016 a and high-quality conference. of we are to the committee composed of the theme (co-)chairs and reviewers for their crucial and dedicated We are to the ECCB committee for their and to the of the conference. In the and by Committee and ECCB ECCB Yves ECCB ECCB and ECCB 2012) were Yves and for and The and of ISCB Society for Computational Biology) in the about ECCB 2016 at the international and for were We thank all provided to the conference. In to the co-organising ELIXIR, we gratefully acknowledge sponsors The Netherlands and USA) and the National Centre for Research (CNRS, We also as where also the opening will be the of The Hague, the and The Bioinformatics Centre the Poster We are to the Oxford University for the ECCB 2016 special issue. The F1000Research is for and the ECCB 2016 posters on their web The ECCB 2016 science has been in the of the efforts have been van (Netherlands Bioinformatics and Systems Biology Research Techcentre for Life ECCB 2016 for the selection (TU Delft, The Netherlands), ECCB 2016 for out the workshop selection (EMBL, Germany), ECCB 2016 Applications Track for and the new Applications Andrew and (all at ELIXIR Hinxton, UK) for taking care in of the first ELIXIR Applications and (University of Finland), ECCB 2016 Poster Track for and the ECCB 2016 to the of the conference, beyond the call of and we a in particular and van both at the Techcentre for Life Sciences for their and in this are also to and van of by Netherlands), ECCB 2016 (Dutch Techcentre for Life had a in the while van is more for also in Finally, all these efforts be the many participants from all over the of will to the conference in of scientific contributions, applications, or poster presentations and all for there and for to science at ECCB 2016 in The
Jaap Heringa, Marcel J. T. Reinders, Sanne Abeln, Jeroen de Ridder
Bioinform.2
2015 Proteny: discovering and visualizing statistically significant syntenic clusters at the proteome level
abstract
BACKGROUND: With more and more genomes being sequenced, detecting synteny between genomes becomes more and more important. However, for microorganisms the genomic divergence quickly becomes large, resulting in different codon usage and shuffling of gene order and gene elements such as exons. RESULTS: We present Proteny, a methodology to detect synteny between diverged genomes. It operates on the amino acid sequence level to be insensitive to codon usage adaptations and clusters groups of exons disregarding order to handle diversity in genomic ordering between genomes. Furthermore, Proteny assigns significance levels to the syntenic clusters such that they can be selected on statistical grounds. Finally, Proteny provides novel ways to visualize results at different scales, facilitating the exploration and interpretation of syntenic regions. We test the performance of Proteny on a standard ground truth dataset, and we illustrate the use of Proteny on two closely related genomes (two different strains of Aspergillus niger) and on two distant genomes (two species of Basidiomycota). In comparison to other tools, we find that Proteny finds clusters with more true homologies in fewer clusters that contain more genes, i.e. Proteny is able to identify a more consistent synteny. Further, we show how genome rearrangements, assembly errors, gene duplications and the conservation of specific genes can be easily studied with Proteny. AVAILABILITY AND IMPLEMENTATION: Proteny is freely available at the Delft Bioinformatics Lab website http://bioinformatics.tudelft.nl/dbl/software. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thies Gehrmann, Marcel J. T. Reinders
Bioinform.2
2015 Integration of gene expression and DNA-methylation profiles improves molecular subtype classification in acute myeloid leukemia
abstract
BACKGROUND: Acute Myeloid Leukemia (AML) is characterized by various cytogenetic and molecular abnormalities. Detection of these abnormalities is important in the risk-classification of patients but requires laborious experimentation. Various studies showed that gene expression profiles (GEP), and the gene signatures derived from GEP, can be used for the prediction of subtypes in AML. Similarly, successful prediction was also achieved by exploiting DNA-methylation profiles (DMP). There are, however, no studies that compared classification accuracy and performance between GEP and DMP, neither are there studies that integrated both types of data to determine whether predictive power can be improved. APPROACH: Here, we used 344 well-characterized AML samples for which both gene expression and DNA-methylation profiles are available. We created three different classification strategies including early, late and no integration of these datasets and used them to predict AML subtypes using a logistic regression model with Lasso regularization. RESULTS: We illustrate that both gene expression and DNA-methylation profiles contain distinct patterns that contribute to discriminating AML subtypes and that an integration strategy can exploit these patterns to achieve synergy between both data types. We show that concatenation of features from both data sets, i.e. early integration, improves the predictive power compared to classifiers trained on GEP or DMP alone. A more sophisticated strategy, i.e. the late integration strategy, employs a two-layer classifier which outperforms the early integration strategy. CONCLUSION: We demonstrate that prediction of known cytogenetic and molecular abnormalities in AML can be further improved by integrating GEP and DMP profiles.
Erdogan Taskesen, Sepideh Babaei, Marcel J. T. Reinders, Jeroen de Ridder
BMC Bioinform.3
2015 An integrated approach of gene expression and DNA-methylation profiles of WNT signaling genes uncovers novel prognostic markers in Acute Myeloid Leukemia
abstract
BACKGROUND: The wingless-Int (WNT) pathway has an essential role in cell regulation of hematopoietic stem cells (HSC). For Acute Myeloid Leukemia (AML), the malignant counterpart of HSC, currently only a selective number of genes of the WNT pathway are analyzed by using either gene expression or DNA-methylation profiles for the identification of prognostic markers and potential candidate targets for drug therapy. It is known that mRNA expression is controlled by DNA-methylation and that specific patterns can infer the ability to differentiate biological differences, thus a combined analysis using all WNT annotated genes could provide more insight in the WNT signaling. APPROACH: We created a computational approach that integrates gene expression and DNA promoter methylation profiles. The approach represents the continuous gene expression and promoter methylation profiles with nine discrete mutually exclusive scenarios. The scenario representation allows for a refinement of patient groups by a more powerful statistical analysis, and the construction of a co-expression network. We focused on 268 WNT annotated signaling genes that are derived from the molecular signature database. RESULTS: Using the scenarios we identified seven prognostic markers for overall survival and event-free survival. Three genes are novel prognostic markers; two with favorable outcome (PSMD2, PPARD) and one with unfavorable outcome (XPNPEP). The remaining four genes (LEF1, SFRP2, RUNX1, and AXIN2) were previously identified but we could refine the patient groups. Three AML risk groups were further analyzed and the co-expression network showed that only the good risk group harbors frequent promoter hypermethylation and significantly correlated interactions with proteasome family members. CONCLUSION: Our results provide novel insights in WNT signaling in AML, we discovered new and previously identified prognostic markers and a refinement of the patient groups.
Erdogan Taskesen, Frank J. T. Staal, Marcel J. T. Reinders
BMC Bioinform.3
2015 Hi-C Chromatin Interaction Networks Predict Co-expression in the Mouse Cortex
abstract
The three dimensional conformation of the genome in the cell nucleus influences important biological processes such as gene expression regulation. Recent studies have shown a strong correlation between chromatin interactions and gene co-expression. However, predicting gene co-expression from frequent long-range chromatin interactions remains challenging. We address this by characterizing the topology of the cortical chromatin interaction network using scale-aware topological measures. We demonstrate that based on these characterizations it is possible to accurately predict spatial co-expression between genes in the mouse cortex. Consistent with previous findings, we find that the chromatin interaction profile of a gene-pair is a good predictor of their spatial co-expression. However, the accuracy of the prediction can be substantially improved when chromatin interactions are described using scale-aware topological measures of the multi-resolution chromatin interaction network. We conclude that, for co-expression prediction, it is necessary to take into account different levels of chromatin interactions ranging from direct interaction between genes (i.e. small-scale) to chromatin compartment interactions (i.e. large-scale).
Sepideh Babaei, Ahmed Mahfouz, Marc Hulsman, Boudewijn P. F. Lelieveldt, Jeroen de Ridder, Marcel J. T. Reinders
PLoS Comput. Biol.6
2015 Unbiased Quantitative Models of Protein Translation Derived from Ribosome Profiling Data
abstract
Translation of RNA to protein is a core process for any living organism. While for some steps of this process the effect on protein production is understood, a holistic understanding of translation still remains elusive. In silico modelling is a promising approach for elucidating the process of protein synthesis. Although a number of computational models of the process have been proposed, their application is limited by the assumptions they make. Ribosome profiling (RP), a relatively new sequencing-based technique capable of recording snapshots of the locations of actively translating ribosomes, is a promising source of information for deriving unbiased data-driven translation models. However, quantitative analysis of RP data is challenging due to high measurement variance and the inability to discriminate between the number of ribosomes measured on a gene and their speed of translation. We propose a solution in the form of a novel multi-scale interpretation of RP data that allows for deriving models with translation dynamics extracted from the snapshots. We demonstrate the usefulness of this approach by simultaneously determining for the first time per-codon translation elongation and per-gene translation initiation rates of Saccharomyces cerevisiae from RP data for two versions of the Totally Asymmetric Exclusion Process (TASEP) model of translation. We do this in an unbiased fashion, by fitting the models using only RP data with a novel optimization scheme based on Monte Carlo simulation to keep the problem tractable. The fitted models match the data significantly better than existing models and their predictions show better agreement with several independent protein abundance datasets than existing models. Results additionally indicate that the tRNA pool adaptation hypothesis is incomplete, with evidence suggesting that tRNA post-transcriptional modifications and codon context may play a role in determining codon elongation rates.
Alexey A. Gritsenko, Marc Hulsman, Marcel J. T. Reinders, Dick de Ridder
PLoS Comput. Biol.3
2014 The Effect of Aggregating Subtype Performances Depends Strongly on the Performance Measure Used
abstract
For some classification tasks the data can be partitioned into disjoint subsets based on some attribute, for example a disease subtype. It then seems logical to train a classifier with the same classes as the original classification problem for each subtype separately, such that the performance per subtype is optimized. Unfortunately, the influence of the subtype performances on the aggregated overall performance depends strongly on the performance measure used and can be very counterintuitive. We show that for some performance measures (e.g., classification accuracy, precision, recall, Fi) the aggregated performance is a simple linear combination of subtype performances. In these cases, improving the performance of a subtype-specific classifier implies that the overall performance improves. However, for other performance measures (e.g., balanced accuracy rate, area under the ROC curve) and also for performance measures in survival analysis (concordance index), additional cross terms appear in the aggregation of the subtype performances. These cross terms are heavily dependent on both the overall class imbalance and the subtype class imbalances. For these measures, improving subtype performances may actually result in a decrease of the overall performance.
David M. J. Tax, Herman M. J. Sontrop, Marcel J. T. Reinders, Perry D. Moerland
ICPR3
2014 SPiCE: a web-based tool for sequence-based protein classification and exploration
abstract
BACKGROUND: Amino acid sequences and features extracted from such sequences have been used to predict many protein properties, such as subcellular localization or solubility, using classifier algorithms. Although software tools are available for both feature extraction and classifier construction, their application is not straightforward, requiring users to install various packages and to convert data into different formats. This lack of easily accessible software hampers quick, explorative use of sequence-based classification techniques by biologists. RESULTS: We have developed the web-based software tool SPiCE for exploring sequence-based features of proteins in predefined classes. It offers data upload/download, sequence-based feature calculation, data visualization and protein classifier construction and testing in a single integrated, interactive environment. To illustrate its use, two example datasets are included showing the identification of differences in amino acid composition between proteins yielding low and high production levels in fungi and low and high expression levels in yeast, respectively. CONCLUSIONS: SPiCE is an easy-to-use online tool for extracting and exploring sequence-based features of sets of proteins, allowing non-experts to apply advanced classification techniques. The tool is available at http://helix.ewi.tudelft.nl/spice.
Bastiaan A. van den Berg, Marcel J. T. Reinders, Johannes A. Roubos, Dick de Ridder
BMC Bioinform.2
2013 Pattern recognition in bioinformatics
abstract
Pattern recognition is concerned with the development of systems that learn to solve a given problem using a set of example instances, each represented by a number of features. These problems include clustering, the grouping of similar instances; classification, the task of assigning a discrete label to a given instance; and dimensionality reduction, combining or selecting features to arrive at a more useful representation. The use of statistical pattern recognition algorithms in bioinformatics is pervasive. Classification and clustering are often applied to high-throughput measurement data arising from microarray, mass spectrometry and next-generation sequencing experiments for selecting markers, predicting phenotype and grouping objects or genes. Less explicitly, classification is at the core of a wide range of tools such as predictors of genes, protein function, functional or genetic interactions, etc., and used extensively in systems biology. A course on pattern recognition (or machine learning) should therefore be at the core of any bioinformatics education program. In this review, we discuss the main elements of a pattern recognition course, based on material developed for courses taught at the BSc, MSc and PhD levels to an audience of bioinformaticians, computer scientists and life scientists. We pay attention to common problems and pitfalls encountered in applications and in interpretation of the results obtained.
Dick de Ridder, Jeroen de Ridder, Marcel J. T. Reinders
Briefings Bioinform.3
2013 Exploring variation-aware contig graphs for (comparative) metagenomics using MaryGold
abstract
MOTIVATION: Although many tools are available to study variation and its impact in single genomes, there is a lack of algorithms for finding such variation in metagenomes. This hampers the interpretation of metagenomics sequencing datasets, which are increasingly acquired in research on the (human) microbiome, in environmental studies and in the study of processes in the production of foods and beverages. Existing algorithms often depend on the use of reference genomes, which pose a problem when a metagenome of a priori unknown strain composition is studied. In this article, we develop a method to perform reference-free detection and visual exploration of genomic variation, both within a single metagenome and between metagenomes. RESULTS: We present the MaryGold algorithm and its implementation, which efficiently detects bubble structures in contig graphs using graph decomposition. These bubbles represent variable genomic regions in closely related strains in metagenomic samples. The variation found is presented in a condensed Circos-based visualization, which allows for easy exploration and interpretation of the found variation. We validated the algorithm on two simulated datasets containing three respectively seven Escherichia coli genomes and showed that finding allelic variation in these genomes improves assemblies. Additionally, we applied MaryGold to publicly available real metagenomic datasets, enabling us to find within-sample genomic variation in the metagenomes of a kimchi fermentation process, the microbiome of a premature infant and in microbial communities living on acid mine drainage. Moreover, we used MaryGold for between-sample variation detection and exploration by comparing sequencing data sampled at different time points for both of these datasets. AVAILABILITY: MaryGold has been written in C++ and Python and can be downloaded from http://bioinformatics.tudelft.nl/software
Jurgen F. Nijkamp, Mihai Pop, Marcel J. T. Reinders, Dick de Ridder
Bioinform.3
2013 Detecting recurrent gene mutation in interaction network context using multi-scale graph diffusion
abstract
BACKGROUND: Delineating the molecular drivers of cancer, i.e. determining cancer genes and the pathways which they deregulate, is an important challenge in cancer research. In this study, we aim to identify pathways of frequently mutated genes by exploiting their network neighborhood encoded in the protein-protein interaction network. To this end, we introduce a multi-scale diffusion kernel and apply it to a large collection of murine retroviral insertional mutagenesis data. The diffusion strength plays the role of scale parameter, determining the size of the network neighborhood that is taken into account. As a result, in addition to detecting genes with frequent mutations in their genomic vicinity, we find genes that harbor frequent mutations in their interaction network context. RESULTS: We identify densely connected components of known and putatively novel cancer genes and demonstrate that they are strongly enriched for cancer related pathways across the diffusion scales. Moreover, the mutations in the clusters exhibit a significant pattern of mutual exclusion, supporting the conjecture that such genes are functionally linked. Using multi-scale diffusion kernel, various infrequently mutated genes are found to harbor significant numbers of mutations in their interaction network neighborhood. Many of them are well-known cancer genes. CONCLUSIONS: The results demonstrate the importance of defining recurrent mutations while taking into account the interaction network context. Importantly, the putative cancer genes and networks detected in this study are found to be significant at different diffusion scales, confirming the necessity of a multi-scale analysis.
Sepideh Babaei, Marc Hulsman, Marcel J. T. Reinders, Jeroen de Ridder
BMC Bioinform.3
2012 Enabling large genomic data transfers using nation-wide and international dynamic lightpaths
abstract
The recent advances made in high throughput genomic sequencing allow researchers to accurately determine the genetic make-up of an individual. Sharing this data across research institutes has proven to be challenging as the amount of data and available bandwidth cause large delays. Here, we present a network of dynamic lightpaths dedicated to the life sciences which connects research groups within the Netherlands to each other, to compute and storage providers and to commercial partners.
Jan Bot, Migiel de Vos, Sander Boele, Marcel J. T. Reinders, Joost N. Kok
eScience4
2012 GRASS: a generic algorithm for scaffolding next-generation sequencing assemblies
abstract
MOTIVATION: The increasing availability of second-generation high-throughput sequencing (HTS) technologies has sparked a growing interest in de novo genome sequencing. This in turn has fueled the need for reliable means of obtaining high-quality draft genomes from short-read sequencing data. The millions of reads usually involved in HTS experiments are first assembled into longer fragments called contigs, which are then scaffolded, i.e. ordered and oriented using additional information, to produce even longer sequences called scaffolds. Most existing scaffolders of HTS genome assemblies are not suited for using information other than paired reads to perform scaffolding. They use this limited information to construct scaffolds, often preferring scaffold length over accuracy, when faced with the tradeoff. RESULTS: We present GRASS (GeneRic ASsembly Scaffolder)-a novel algorithm for scaffolding second-generation sequencing assemblies capable of using diverse information sources. GRASS offers a mixed-integer programming formulation of the contig scaffolding problem, which combines contig order, distance and orientation in a single optimization objective. The resulting optimization problem is solved using an expectation-maximization procedure and an unconstrained binary quadratic programming approximation of the original problem. We compared GRASS with existing HTS scaffolders using Illumina paired reads of three bacterial genomes. Our algorithm constructs a comparable number of scaffolds, but makes fewer errors. This result is further improved when additional data, in the form of related genome sequences, are used. AVAILABILITY: GRASS source code is freely available from http://code.google.com/p/tud-scaffolding/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexey A. Gritsenko, Jurgen F. Nijkamp, Marcel J. T. Reinders, Dick de Ridder
Bioinform.3
2012 De novo detection of copy number variation by co-assembly
abstract
MOTIVATION: Comparing genomes of individual organisms using next-generation sequencing data is, until now, mostly performed using a reference genome. This is challenging when the reference is distant and introduces bias towards the exact sequence present in the reference. Recent improvements in both sequencing read length and efficiency of assembly algorithms have brought direct comparison of individual genomes by de novo assembly, rather than through a reference genome, within reach. RESULTS: Here, we develop and test an algorithm, named Magnolya, that uses a Poisson mixture model for copy number estimation of contigs assembled from sequencing data. We combine this with co-assembly to allow de novo detection of copy number variation (CNV) between two individual genomes, without mapping reads to a reference genome. In co-assembly, multiple sequencing samples are combined, generating a single contig graph with different traversal counts for the nodes and edges between the samples. In the resulting 'coloured' graph, the contigs have integer copy numbers; this negates the need to segment genomic regions based on depth of coverage, as required for mapping-based detection methods. Magnolya is then used to assign integer copy numbers to contigs, after which CNV probabilities are easily inferred. The copy number estimator and CNV detector perform well on simulated data. Application of the algorithms to hybrid yeast genomes showed allotriploid content from different origin in the wine yeast Y12, and extensive CNV in aneuploid brewing yeast genomes. Integer CNV was also accurately detected in a short-term laboratory-evolved yeast strain.
Jurgen F. Nijkamp, Marcel van den Broek, Jan-Maarten A. Geertman, Marcel J. T. Reinders, Jean-Marc Daran, Dick de Ridder
Bioinform.4
2011 CytoscapeRPC: a plugin to create, modify and query Cytoscape networks from scripting languages
abstract
SUMMARY: CytoscapeRPC is a plugin for Cytoscape which allows users to create, query and modify Cytoscape networks from any programming language which supports XML-RPC. This enables them to access Cytoscape functionality and visualize their data interactively without leaving the programming environment with which they are familiar. AVAILABILITY: Install through the Cytoscape plugin manager or visit the web page: http://wiki.nbic.nl/index.php/CytoscapeRPC for the user tutorial and download. CONTACT: [email protected]; [email protected].
Jan Bot, Marcel J. T. Reinders
Bioinform.2
2011 Human-inspired search for redundancy in automatic sign language recognition
abstract
Human perception of sign language can serve as inspiration for the improvement of automatic recognition systems. Experiments with human signers show that sign language signs contain redundancy over time. In this article, experiments are conducted to investigate whether comparable redundancies also exist for an automatic sign language recognition system. Such redundancies could be exploited, for example, by reserving more processing resources for the more informative phases of a sign, or by discarding uninformative phases. In the experiments, an automatic system is trained and tested on isolated fragments of sign language signs. The stimuli used were similar to those of the human signer experiments, allowing us to compare the results. The experiments show that redundancy over time exists for the automatic recognizer. The central phase of a sign is the most informative phase, and the first half of a sign is sufficient to achieve a recognition performance similar to that of the entire sign. These findings concur with the results of the human signer studies. However, there are differences as well, most notably the fact that human signers score better on the early phases of a sign than the automatic system. The results can be used to improve the automatic recognizer, by using only the most informative phases of a sign as input.
Gineke A. ten Holt, Andrea J. van Doorn, Marcel J. T. Reinders, Emile A. Hendriks, Huib de Ridder
ACM Trans. Appl. Percept.3
2011 Predicting Metabolic Fluxes Using Gene Expression Differences As Constraints
abstract
A standard approach to estimate intracellular fluxes on a genome-wide scale is flux-balance analysis (FBA), which optimizes an objective function subject to constraints on (relations between) fluxes. The performance of FBA models heavily depends on the relevance of the formulated objective function and the completeness of the defined constraints. Previous studies indicated that FBA predictions can be improved by adding regulatory on/off constraints. These constraints were imposed based on either absolute or relative gene expression values. We provide a new algorithm that directly uses regulatory up/down constraints based on gene expression data in FBA optimization (tFBA). Our assumption is that if the activity of a gene drastically changes from one condition to the other, the flux through the reaction controlled by that gene will change accordingly. We allow these constraints to be violated, to account for posttranscriptional control and noise in the data. These up/down constraints are less stringent than the on/off constraints as previously proposed. Nevertheless, we obtain promising predictions, since many up/down constraints can be enforced. The potential of the proposed method, tFBA, is demonstrated through the analysis of fluxes in yeast under nine different cultivation conditions, between which approximately 5,000 regulatory up/down constraints can be defined. We show that changes in gene expression are predictive for changes in fluxes. Additionally, we illustrate that flux distributions obtained with tFBA better fit transcriptomics data than previous methods. Finally, we compare tFBA and FBA predictions to show that our approach yields more biologically relevant results.
Rogier J. P. van Berlo, Dick de Ridder, Jean-Marc Daran, Pascale A. S. Daran-Lapujade, Bas Teusink, Marcel J. T. Reinders
IEEE ACM Trans. Comput. Biol. Bioinform.6
2010 Finding Wormholes with Flickr Geotags
Maarten Clements, Pavel Serdyukov, Arjen P. de Vries, Marcel J. T. Reinders
ECIR4
2010 Using flickr geotags to predict user travel behaviour
abstract
We propose a method to predict a user's favourite locations in a city, based on his Flickr geotags in other cities. We define a similarity between the geotag distributions of two users based on a Gaussian kernel convolution. The geotags of the most similar users are then combined to rerank the popular locations in the target city personalised for this user.
Maarten Clements, Pavel Serdyukov, Arjen P. de Vries, Marcel J. T. Reinders
SIGIR4
2010 Integrating genome assemblies with MAIA
abstract
MOTIVATION: De novo assembly of a eukaryotic genome with next-generation sequencing data is still a challenging task. Over the past few years several assemblers have been developed, often suitable for one specific type of sequencing data. The number of known genomes is expanding rapidly, therefore it becomes possible to use multiple reference genomes for assembly projects. We introduce an assembly integrator that makes use of all available data, i.e. multiple de novo assemblies and mappings against multiple related genomes, by optimizing a weighted combination of criteria. RESULTS: The developed algorithm was applied on the de novo sequencing of the Saccharomyces cerevisiae CEN.PK 113-7D strain. Using Solexa and 454 read data, two de novo and three comparative assemblies were constructed and subsequently integrated, yielding 29 contigs, covering more than 12 Mbp; a drastic improvement compared with the single assemblies. AVAILABILITY: MAIA is available as a Matlab package and can be downloaded from http://bioinformatics.tudelft.nl.
Jurgen F. Nijkamp, Wynand Winterbach, Marcel van den Broek, Jean-Marc Daran, Marcel J. T. Reinders, Dick de Ridder
Bioinform.5
2010 Inferring combinatorial association logic networks in multimodal genome-wide screens
abstract
MOTIVATION: We propose an efficient method to infer combinatorial association logic networks from multiple genome-wide measurements from the same sample. We demonstrate our method on a genetical genomics dataset, in which we search for Boolean combinations of multiple genetic loci that associate with transcript levels. RESULTS: Our method provably finds the global solution and is very efficient with runtimes of up to four orders of magnitude faster than the exhaustive search. This enables permutation procedures for determining accurate false positive rates and allows selection of the most parsimonious model. When applied to transcript levels measured in myeloid cells from 24 genotyped recombinant inbred mouse strains, we discovered that nine gene clusters are putatively modulated by a logical combination of trait loci rather than a single locus. A literature survey supports and further elucidates one of these findings. Due to our approach, optimal solutions for multi-locus logic models and accurate estimates of the associated false discovery rates become feasible. Our algorithm, therefore, offers a valuable alternative to approaches employing complex, albeit suboptimal optimization strategies to identify complex models. AVAILABILITY: The MATLAB code of the prototype implementation is available on: http://bioinformatics.tudelft.nl/ or http://bioinformatics.nki.nl/.
Jeroen de Ridder, Alice Gerrits, Jan Bot, Gerald de Haan, Marcel J. T. Reinders, Lodewyk F. A. Wessels
Bioinform.5
2010 Delineation of amplification, hybridization and location effects in microarray data yields better-quality normalization
abstract
BACKGROUND: Oligonucleotide arrays have become one of the most widely used high-throughput tools in biology. Due to their sensitivity to experimental conditions, normalization is a crucial step when comparing measurements from these arrays. Normalization is, however, far from a solved problem. Frequently, we encounter datasets with significant technical effects that currently available methods are not able to correct. RESULTS: We show that by a careful decomposition of probe specific amplification, hybridization and array location effects, a normalization can be performed that allows for a much improved analysis of these data. Identification of the technical sources of variation between arrays has allowed us to build statistical models that are used to estimate how the signal of individual probes is affected, based on their properties. This enables a model-based normalization that is probe-specific, in contrast with the signal intensity distribution normalization performed by many current methods. Next to this, we propose a novel way of handling background correction, enabling the use of background information to weight probes during summarization. Testing of the proposed method shows a much improved detection of differentially expressed genes over earlier proposed methods, even when tested on (experimentally tightly controlled and replicated) spike-in datasets. CONCLUSIONS: When a limited number of arrays are available, or when arrays are run in different batches, technical effects have a large influence on the measured expression of genes. We show that a detailed modelling and correction of these technical effects allows for an improved analysis in these situations.
Marc Hulsman, Anouk Mentink, Eugene P. van Someren, Koen J. Dechering, Jan de Boer, Marcel J. T. Reinders
BMC Bioinform.6
2010 HAT: Hypergeometric Analysis of Tiling-arrays with application to promoter-GeneChip data
abstract
BACKGROUND: Tiling-arrays are applicable to multiple types of biological research questions. Due to its advantages (high sensitivity, resolution, unbiased), the technology is often employed in genome-wide investigations. A major challenge in the analysis of tiling-array data is to define regions-of-interest, i.e., contiguous probes with increased signal intensity (as a result of hybridization of labeled DNA) in a region. Currently, no standard criteria are available to define these regions-of-interest as there is no single probe intensity cut-off level, different regions-of-interest can contain various numbers of probes, and can vary in genomic width. Furthermore, the chromosomal distance between neighboring probes can vary across the genome among different arrays. RESULTS: We have developed Hypergeometric Analysis of Tiling-arrays (HAT), and first evaluated its performance for tiling-array datasets from a Chromatin Immunoprecipitation study on chip (ChIP-on-chip) for the identification of genome-wide DNA binding profiles of transcription factor Cebpa (used for method comparison). Using this assay, we can refine the detection of regions-of-interest by illustrating that regions detected by HAT are more highly enriched for expected motifs in comparison with an alternative detection method (MAT). Subsequently, data from a retroviral insertional mutagenesis screen were used to examine the performance of HAT among different applications of tiling-array datasets. In both studies, detected regions-of-interest have been validated with (q)PCR. CONCLUSIONS: We demonstrate that HAT has increased specificity for analysis of tiling-array data in comparison with the alternative method, and that it accurately detects regions-of-interest in two different applications of tiling-arrays. HAT has several advantages over previous methods: i) as there is no single cut-off level for probe-intensity, HAT can detect regions-of-interest at various thresholds, ii) it can detect regions-of-interest of any size, iii) it is independent of probe-resolution across the genome, and across tiling-array platforms and iv) it employs a single user defined parameter: the significance level. Regions-of-interest are detected by computing the hypergeometric-probability, while controlling the Family Wise Error. Furthermore, the method does not require experimental replicates, common regions-of-interest are indicated, a sequence-of-interest can be examined for every detected region-of-interest, and flanking genes can be reported.
Erdogan Taskesen, Renee Beekman, Jeroen de Ridder, Bas J. Wouters, Justine K. Peeters, Ivo P. Touw, Marcel J. T. Reinders, Ruud Delwel
BMC Bioinform.7
2010 The influence of personalization on tag query length in social media search
Maarten Clements, Arjen P. de Vries, Marcel J. T. Reinders
Inf. Process. Manag.3
2010 Personalization of tagging systems
Jun Wang 0012, Maarten Clements, Jie Yang 0015, Arjen P. de Vries, Marcel J. T. Reinders
Inf. Process. Manag.5
2010 Identification of Networks of Co-Occurring, Tumor-Related DNA Copy Number Changes Using a Genome-Wide Scoring Approach
abstract
Tumorigenesis is a multi-step process in which normal cells transform into malignant tumors following the accumulation of genetic mutations that enable them to evade the growth control checkpoints that would normally suppress their growth or result in apoptosis. It is therefore important to identify those combinations of mutations that collaborate in cancer development and progression. DNA copy number alterations (CNAs) are one of the ways in which cancer genes are deregulated in tumor cells. We hypothesized that synergistic interactions between cancer genes might be identified by looking for regions of co-occurring gain and/or loss. To this end we developed a scoring framework to separate truly co-occurring aberrations from passenger mutations and dominant single signals present in the data. The resulting regions of high co-occurrence can be investigated for between-region functional interactions. Analysis of high-resolution DNA copy number data from a panel of 95 hematological tumor cell lines correctly identified co-occurring recombinations at the T-cell receptor and immunoglobulin loci in T- and B-cell malignancies, respectively, showing that we can recover truly co-occurring genomic alterations. In addition, our analysis revealed networks of co-occurring genomic losses and gains that are enriched for cancer genes. These networks are also highly enriched for functional relationships between genes. We further examine sub-networks of these networks, core networks, which contain many known cancer genes. The core network for co-occurring DNA losses we find seems to be independent of the canonical cancer genes within the network. Our findings suggest that large-scale, low-intensity copy number alterations may be an important feature of cancer development or maintenance by affecting gene dosage of a large interconnected network of functionally related genes.
Christiaan Klijn, Jan Bot, David J. Adams, Marcel J. T. Reinders, Lodewyk F. A. Wessels, Jos Jonkers
PLoS Comput. Biol.4
2010 The task-dependent effect of tags and ratings on social media access
abstract
Recently, online social networks have emerged that allow people to share their multimedia files, retrieve interesting content, and discover like-minded people. These systems often provide the possibility to annotate the content with tags and ratings. Using a random walk through the social annotation graph, we have combined these annotations into a retrieval model that effectively balances the personal preferences and opinions of like-minded users into a single relevance ranking for either content, tags, or people. We use this model to identify the influence of different annotation methods and system design aspects on common ranking tasks in social content systems. Our results show that a combination of rating and tagging information can improve tasks like search and recommendation. The optimal influence of both sources on the ranking is highly dependent on the retrieval task and system design. Results on content search and tag suggestion indicate that the profile created by a user's annotations can be used effectively to adapt the ranking to personal preferences. The random walk reduces sparsity problems by smoothly integrating indirectly related concepts in the relevance ranking, which is especially valuable for cold-start users or individual tagging systems like YouTube and Flickr.
Maarten Clements, Arjen P. de Vries, Marcel J. T. Reinders
ACM Trans. Inf. Syst.3
2009 Exploiting Positive and Negative Graded Relevance Assessments for Content Recommendation
Maarten Clements, Arjen P. de Vries, Marcel J. T. Reinders
WAW3
2009 Metabolite and reaction inference based on enzyme specificities
abstract
MOTIVATION: Many enzymes are not absolutely specific, or even promiscuous: they can catalyze transformations of more compounds than the traditional ones as listed in, e.g. KEGG. This information is currently only available in databases, such as the BRENDA enzyme activity database. In this article, we propose to model enzyme aspecificity by predicting whether an input compound is likely to be transformed by a certain enzyme. Such a predictor has many applications, for example, to complete reconstructed metabolic networks, to aid in metabolic engineering or to help identify unknown peaks in mass spectra. RESULTS: We have developed a system for metabolite and reaction inference based on enzyme specificities (MaRIboES). It employs structural and stereochemistry similarity measures and molecular fingerprints to generalize enzymatic reactions based on data available in BRENDA. Leave-one-out cross-validation shows that 80% of known reactions are predicted well. Application to the yeast glycolytic and pentose phosphate pathways predicts a large number of known and new reactions, often leading to the formation of novel compounds, as well as a number of interesting bypasses and cross-links. AVAILABILITY: Matlab and C++ code is freely available at https://gforge.nbic.nl/projects/mariboes/
Marco J. L. de Groot, Rogier J. P. van Berlo, Wouter A. van Winden, Peter J. T. Verheijen, Marcel J. T. Reinders, Dick de Ridder
Bioinform.5
2009 Fewer permutations, more accurate P-values
abstract
MOTIVATION: Permutation tests have become a standard tool to assess the statistical significance of an event under investigation. The statistical significance, as expressed in a P-value, is calculated as the fraction of permutation values that are at least as extreme as the original statistic, which was derived from non-permuted data. This empirical method directly couples both the minimal obtainable P-value and the resolution of the P-value to the number of permutations. Thereby, it imposes upon itself the need for a very large number of permutations when small P-values are to be accurately estimated. This is computationally expensive and often infeasible. RESULTS: A method of computing P-values based on tail approximation is presented. The tail of the distribution of permutation values is approximated by a generalized Pareto distribution. A good fit and thus accurate P-value estimates can be obtained with a drastically reduced number of permutations when compared with the standard empirical way of computing P-values. AVAILABILITY: The Matlab code can be obtained from the corresponding author on request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Theo A. Knijnenburg, Lodewyk F. A. Wessels, Marcel J. T. Reinders, Ilya Shmulevich
Bioinform.3
2009 Analysis of mass spectrometry data using sub-spectra
abstract
BACKGROUND: Spectra resulting from Surface-Enhanced Laser Desorption/Ionisation (SELDI) mass spectrometry measurements are constructed by combining sub-spectra, each of which are the result of a single firing of the laser responsible for the process of desorption/ionisation. These firings are performed at different locations of the spot on which the sample is analysed. The final spectrum is then constructed by summing over all these sub-spectra. This process is sub-optimal in that it can average out peaks from peptides that are present in low abundance or are unevenly distributed across the spot, particularly because the amount of noise varies considerably between sub-spectra. This argues for analysing sub-spectra separately and combining results afterwards. RESULTS: Here, we propose to analyse these sub-spectra one-by-one and combine the results using a framework which includes a significance test. This allows one to, for the first time, attach a confidence measure to detected peaks, based on the signal strength of a peak across sub-spectra. In a comparison with three other approaches the sub-spectral approach achieves a higher sensitivity and a low FDR. We further introduce the notion of peak-bags, which provide rich information about the sub-spectral contributions to a given peak. CONCLUSION: The proposed procedure offers better control over the process of distinguishing signal from noise, resulting in an improved performance over other available methods. Moreover, our method provides an implicit deconvolution of peaks, yielding insight in the actual shape of a peak, potentially aiding in a deeper understanding of peak distribution. AVAILABILITY: Implementations of the algorithm in R are available upon request.
Wouter Meuleman, Judith Y. M. N. Engwegen, Marie-Christine W. Gast, Lodewyk F. A. Wessels, Marcel J. T. Reinders
BMC Bioinform.5
2009 A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability
abstract
BACKGROUND: Large discrepancies in signature composition and outcome concordance have been observed between different microarray breast cancer expression profiling studies. This is often ascribed to differences in array platform as well as biological variability. We conjecture that other reasons for the observed discrepancies are the measurement error associated with each feature and the choice of preprocessing method. Microarray data are known to be subject to technical variation and the confidence intervals around individual point estimates of expression levels can be wide. Furthermore, the estimated expression values also vary depending on the selected preprocessing scheme. In microarray breast cancer classification studies, however, these two forms of feature variability are almost always ignored and hence their exact role is unclear. RESULTS: We have performed a comprehensive sensitivity analysis of microarray breast cancer classification under the two types of feature variability mentioned above. We used data from six state of the art preprocessing methods, using a compendium consisting of eight different datasets, involving 1131 hybridizations, containing data from both one and two-color array technology. For a wide range of classifiers, we performed a joint study on performance, concordance and stability. In the stability analysis we explicitly tested classifiers for their noise tolerance by using perturbed expression profiles that are based on uncertainty information directly related to the preprocessing methods. Our results indicate that signature composition is strongly influenced by feature variability, even if the array platform and the stratification of patient samples are identical. In addition, we show that there is often a high level of discordance between individual class assignments for signatures constructed on data coming from different preprocessing schemes, even if the actual signature composition is identical. CONCLUSION: Feature variability can have a strong impact on breast cancer signature composition, as well as the classification of individual patient samples. We therefore strongly recommend that feature variability is considered in analyzing data from microarray breast cancer expression profiling experiments.
Herman M. J. Sontrop, Perry D. Moerland, René van den Ham, Marcel J. T. Reinders, Wim F. J. Verhaegh
BMC Bioinform.4
2009 Knowledge driven decomposition of tumor expression profiles
abstract
BACKGROUND: Tumors have been hypothesized to be the result of a mixture of oncogenic events, some of which will be reflected in the gene expression of the tumor. Based on this hypothesis a variety of data-driven methods have been employed to decompose tumor expression profiles into component profiles, hypothetically linked to these events. Interpretation of the resulting data-driven components is often done by post-hoc comparison to, for instance, functional groupings of genes into gene sets. None of the data-driven methods allow the incorporation of that type of knowledge directly into the decomposition. RESULTS: We present a linear model which uses knowledge driven, pre-defined components to perform the decomposition. We solve this decomposition model in a constrained linear least squares fashion. From a variety of options, a lasso-based solution to the model performs best in linking single gene perturbation data to mouse data. Moreover, we show the decomposition of expression profiles from human breast cancer samples into single gene perturbation profiles and gene sets that are linked to the hallmarks of cancer. For these breast cancer samples we were able to discern several links between clinical parameters, and the decomposition weights, providing new insights into the biology of these tumors. Lastly, we show that the order in which the Lasso regularization shrinks the weights, unveils consensus patterns within clinical subgroups of the breast cancer samples. CONCLUSION: The proposed lasso-based constrained least squares decomposition provides a stable and relevant relation between samples and knowledge-based components, and is thus a viable alternative to data-driven methods. In addition, the consensus order of component importance within clinical subgroups provides a better molecular characterization of the subtypes.
Martin H. van Vliet, Lodewyk F. A. Wessels, Marcel J. T. Reinders
BMC Bioinform.3
2009 Evolutionary Optimization of Kernel Weights Improves Protein Complex Comembership Prediction
abstract
In recent years, more and more high-throughput data sources useful for protein complex prediction have become available (e.g., gene sequence, mRNA expression, and interactions). The integration of these different data sources can be challenging. Recently, it has been recognized that kernel-based classifiers are well suited for this task. However, the different kernels (data sources) are often combined using equal weights. Although several methods have been developed to optimize kernel weights, no large-scale example of an improvement in classifier performance has been shown yet. In this work, we employ an evolutionary algorithm to determine weights for a larger set of kernels by optimizing a criterion based on the area under the ROC curve. We show that setting the right kernel weights can indeed improve performance. We compare this to the existing kernel weight optimization methods (i.e., (regularized) optimization of the SVM criterion or aligning the kernel with an ideal kernel) and find that these do not result in a significant performance improvement and can even cause a decrease in performance. Results also show that an expert approach of assigning high weights to features with high individual performance is not necessarily the best strategy.
Marc Hulsman, Marcel J. T. Reinders, Dick de Ridder
IEEE ACM Trans. Comput. Biol. Bioinform.2
2008 Metabolic Pathway Alignment (M-Pal) Reveals Diversity and Alternatives in Conserved Networks
Yunlei Li, Dick de Ridder, Marco J. L. de Groot, Marcel J. T. Reinders
APBC4
2008 A learning environment for sign language
abstract
We have developed a prototype for a learning environment for deaf and hard of hearing children. This demonstration consists of hands-on experience with the prototype. In total, there are three exercises: 1) an introduction of all pictures and corresponding signs, 2) multiple choice sign-to-picture and 3) performing the sign that corresponds to the picture shown on the screen. The live recognition from a wide-angle stereo camera provides immediate feedback for the third exercise where the sign must be performed.
Jeroen Lichtenauer, Gineke A. ten Holt, Emile A. Hendriks, Marcel J. T. Reinders, Abbie Vanhoutte, Irene Kamp, Jeroen Arendsen, Ans J. van Doorn, Huib de Ridder, Elmar Wenners, Mariëlle Elzenaar, Gerard Spaai, Connie Fortgens, Marjan Bruins
FG4
2008 Learning to recognize a sign from a single example
abstract
We present a method to automatically construct a sign language classifier for a previously unseen sign. The only required input of a new sign is one example, performed by a sign language tutor. The method works by comparing the measurements of the new sign to signs that have been trained on a large number of persons. The parameters of the respective trained classifier models are used to construct a classification model for the new sign. We show that the performance of a classifier constructed from an instructed sign is significantly better than that of dynamic time warping (DTW) with the same sign. Using only a single example, the proposed method has a performance comparable to a regular training with five examples, while being more stable because of the larger source of information.
Jeroen Lichtenauer, Emile A. Hendriks, Marcel J. T. Reinders
FG3
2008 Combinatorial influence of environmental parameters on transcription factor activity
abstract
MOTIVATION: Cells receive a wide variety of environmental signals, which are often processed combinatorially to generate specific genetic responses. Changes in transcript levels, as observed across different environmental conditions, can, to a large extent, be attributed to changes in the activity of transcription factors (TFs). However, in unraveling these transcription regulation networks, the actual environmental signals are often not incorporated into the model, simply because they have not been measured. The unquantified heterogeneity of the environmental parameters across microarray experiments frustrates regulatory network inference. RESULTS: We propose an inference algorithm that models the influence of environmental parameters on gene expression. The approach is based on a yeast microarray compendium of chemostat steady-state experiments. Chemostat cultivation enables the accurate control and measurement of many of the key cultivation parameters, such as nutrient concentrations, growth rate and temperature. The observed transcript levels are explained by inferring the activity of TFs in response to combinations of cultivation parameters. The interplay between activated enhancers and repressors that bind a gene promoter determine the possible up- or downregulation of the gene. The model is translated into a linear integer optimization problem. The resulting regulatory network identifies the combinatorial effects of environmental parameters on TF activity and gene expression. AVAILABILITY: The Matlab code is available from the authors upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Theo A. Knijnenburg, Lodewyk F. A. Wessels, Marcel J. T. Reinders
ISMB3
2008 Detecting synonyms in social tagging systems to improve content retrieval
abstract
Collaborative tagging used in online social content systems is naturally characterized by many synonyms, causing low precision retrieval. We propose a mechanism based on user preference profiles to identify synonyms that can be used to retrieve more relevant documents by expanding the user's query. Using a popular online book catalog we discuss the effectiveness of our method over usual similarity based expansion methods.
Maarten Clements, Arjen P. de Vries, Marcel J. T. Reinders
SIGIR3
2008 Comparison of normalisation methods for surface-enhanced laser desorption and ionisation (SELDI) time-of-flight (TOF) mass spectrometry data
abstract
BACKGROUND: Mass spectrometry for biological data analysis is an active field of research, providing an efficient way of high-throughput proteome screening. A popular variant of mass spectrometry is SELDI, which is often used to measure sample populations with the goal of developing (clinical) classifiers. Unfortunately, not only is the data resulting from such measurements quite noisy, variance between replicate measurements of the same sample can be high as well. Normalisation of spectra can greatly reduce the effect of this technical variance and further improve the quality and interpretability of the data. However, it is unclear which normalisation method yields the most informative result. RESULTS: In this paper, we describe the first systematic comparison of a wide range of normalisation methods, using two objectives that should be met by a good method. These objectives are minimisation of inter-spectra variance and maximisation of signal with respect to class separation. The former is assessed using an estimation of the coefficient of variation, the latter using the classification performance of three types of classifiers on real-world datasets representing two-class diagnostic problems. To obtain a maximally robust evaluation of a normalisation method, both objectives are evaluated over multiple datasets and multiple configurations of baseline correction and peak detection methods. Results are assessed for statistical significance and visualised to reveal the performance of each normalisation method, in particular with respect to using no normalisation. The normalisation methods described have been implemented in the freely available MASDA R-package. CONCLUSION: In the general case, normalisation of mass spectra is beneficial to the quality of data. The majority of methods we compared performed significantly better than the case in which no normalisation was used. We have shown that normalisation methods that scale spectra by a factor based on the dispersion (e.g., standard deviation) of the data clearly outperform those where a factor based on the central location (e.g., mean) is used. Additional improvements in performance are obtained when these factors are estimated locally, using a sliding window within spectra, instead of globally, over full spectra. The underperforming category of methods using a globally estimated factor based on the central location of the data includes the method used by the majority of SELDI users.
Wouter Meuleman, Judith Y. M. N. Engwegen, Marie-Christine W. Gast, Jos H. Beijnen, Marcel J. T. Reinders, Lodewyk F. A. Wessels
BMC Bioinform.5
2008 TRIBLER: a social-based peer-to-peer system
abstract
Abstract Most current peer‐to‐peer (P2P) file‐sharing systems treat their users as anonymous, unrelated entities, and completely disregard any social relationships between them. However, social phenomena such as friendship and the existence of communities of users with similar tastes or interests may well be exploited in such systems in order to increase their usability and performance. In this paper we present a novel social‐based P2P file‐sharing paradigm that exploits social phenomena by maintaining social networks and using these in content discovery, content recommendation, and downloading. Based on this paradigm's main concepts such as taste buddies and friends, we have designed and implemented the TRIBLER P2P file‐sharing system as a set of extensions to BitTorrent. We present and discuss the design of TRIBLER, and we show evidence that TRIBLER enables fast content discovery and recommendation at a low additional overhead, and a significant improvement in download performance. Copyright © 2007 John Wiley & Sons, Ltd.
Johan A. Pouwelse, Pawel Garbacki, Jun Wang 0012, Arno Bakker, Jie Yang 0015, Alexandru Iosup, Dick H. J. Epema, Marcel J. T. Reinders, Maarten van Steen, Henk J. Sips
Concurr. Comput. Pract. Exp.8
2008 Probabilistic relevance ranking for collaborative filtering
Jun Wang 0012, Stephen E. Robertson, Arjen P. de Vries, Marcel J. T. Reinders
Inf. Retr.4
2008 Personalization on a peer-to-peer television system
Jun Wang 0012, Johan A. Pouwelse, Jenneke Fokker, Arjen P. de Vries, Marcel J. T. Reinders
Multim. Tools Appl.5
2008 Sign Language Recognition by Combining Statistical DTW and Independent Classification
abstract
To recognize speech, handwriting or sign language, many hybrid approaches have been proposed that combine Dynamic Time Warping (DTW) or Hidden Markov Models (HMM) with discriminative classifiers. However, all methods rely directly on the likelihood models of DTW/HMM. We hypothesize that time warping and classification should be separated because of conflicting likelihood modelling demands. To overcome these restrictions, we propose to use Statistical DTW (SDTW) only for time warping, while classifying the warped features with a different method. Two novel statistical classifiers are proposed (CDFD and Q-DFFM), both using a selection of discriminative features (DF), and are shown to outperform HMM and SDTW. However, we have found that combining likelihoods of multiple models in a second classification stage degrades performance of the proposed classifiers, while improving performance with HMM and SDTW. A proof-of-concept experiment, combining DFFM mappings of multiple SDTW models with SDTW likelihoods, shows that also for model-combining, hybrid classification can provide significant improvement over SDTW. Although recognition is mainly based on 3D hand motion features, these results can be expected to generalize to recognition with more detailed measurements such as hand/body pose and facial expression.
Jeroen Lichtenauer, Emile A. Hendriks, Marcel J. T. Reinders
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Integration of prior knowledge of measurement noise in kernel density classification
Yunlei Li, Dick de Ridder, Robert P. W. Duin, Marcel J. T. Reinders
Pattern Recognit.4
2008 Erratum to "Classification in the presence of class noise using a probabilistic kernel fisher method": [Pattern Recognition 40 (12) 3349-3357]
Yunlei Li, Lodewyk F. A. Wessels, Dick de Ridder, Marcel J. T. Reinders
Pattern Recognit.4
2008 Unified relevance models for rating prediction in collaborative filtering
abstract
Collaborative filtering aims at predicting a user's interest for a given item based on a collection of user profiles. This article views collaborative filtering as a problem highly related to information retrieval, drawing an analogy between the concepts of users and items in recommender systems and queries and documents in text retrieval. We present a probabilistic user-to-item relevance framework that introduces the concept of relevance into the related problem of collaborative filtering. Three different models are derived, namely, auser-based, anitem-based, and aunified relevance model, and we estimate their rating predictions from three sources: the user's own ratings for different items, other users' ratings for the same item, and ratings from different but similar users for other but similar items. To reduce the data sparsity encountered when estimating the probability density function of the relevance variable, we apply the nonparametric (data-driven) density estimation technique known as theParzen-window method(or kernel-based density estimation). Using a Gaussian window function, the similarity between users and/or items would, however, be based on Euclidean distance. Because the collaborative filtering literature has reported improved prediction accuracy when using cosine similarity, we generalize the Parzen-window method by introducing aprojection kernel. Existing user-based and item-based approaches correspond to two simplified instantiations of our framework. User-based and item-based collaborative filterings represent only a partial view of the prediction problem, where the unified relevance model brings these partial views together under the same umbrella. Experimental results complement the theoretical insights with improved recommendation accuracy. The unified model is more robust to data sparsity because the different types of ratings are used in concert.
Jun Wang 0012, Arjen P. de Vries, Marcel J. T. Reinders
ACM Trans. Inf. Syst.3
2007 Sign language detection using 3D visual cues
abstract
A 3D visual hand gesture recognition method is proposed that detects correctly performed signs from stereo camera input. Hand tracking is based on skin detection with an adaptive chrominance model to get high accuracy. Informative high level motion properties are extracted to simplify the classification task. Each example is mapped onto a fixed reference sign by Dynamic Time Warping, to get precise time correspondences. The classification is done by combining weak classifiers based on robust statistics. Each base classifier assumes a uniform distribution of a single feature, determined by winsorization on the noisy training set. The operating point of the classifier is determined by stretching the uniform distributions of the base classifiers instead of changing the threshold on the total posterior likelihood. In a cross validation with 120 signs performed by 70 different persons, 95% of the test signs were correctly detected at a false positive rate of 5%.
Jeroen Lichtenauer, Gineke A. ten Holt, Emile A. Hendriks, Marcel J. T. Reinders
AVSS4
2007 SIRAC: Supervised Identification of Regions of Aberration in aCGH datasets
abstract
BACKGROUND: Array comparative genome hybridization (aCGH) provides information about genomic aberrations. Alterations in the DNA copy number may cause the cell to malfunction, leading to cancer. Therefore, the identification of DNA amplifications or deletions across tumors may reveal key genes involved in cancer and improve our understanding of the underlying biological processes associated with the disease. RESULTS: We propose a supervised algorithm for the analysis of aCGH data and the identification of regions of chromosomal alteration (SIRAC). We first determine the DNA-probes that are important to distinguish the classes of interest, and then evaluate in a systematic and robust scheme if these relevant DNA-probes are closely located, i.e. form a region of amplification/deletion. SIRAC does not need any preprocessing of the aCGH datasets, and requires only few, intuitive parameters. CONCLUSION: We illustrate the features of the algorithm with the use of a simple artificial dataset. The results on two breast cancer datasets show promising outcomes that are in agreement with previous findings, but SIRAC better pinpoints the dissimilarities between the classes of interest.
Carmen Lai, Hugo M. Horlings, Marc J. van de Vijver, Eric H. van Beers, Petra M. Nederlof, Lodewyk F. A. Wessels, Marcel J. T. Reinders
BMC Bioinform.7
2007 Classification in the presence of class noise using a probabilistic Kernel Fisher method
Yunlei Li, Lodewyk F. A. Wessels, Dick de Ridder, Marcel J. T. Reinders
Pattern Recognit.4
2006 A User-Item Relevance Model for Log-Based Collaborative Filtering
Jun Wang 0012, Arjen P. de Vries, Marcel J. T. Reinders
ECIR3
2006 Unifying user-based and item-based collaborative filtering approaches by similarity fusion
abstract
Memory-based methods for collaborative filtering predict new ratings by averaging (weighted) ratings between, respectively, pairs of similar users or items. In practice, a large number of ratings from similar users or similar items are not available, due to the sparsity inherent to rating data. Consequently, prediction quality can be poor. This paper re-formulates the memory-based collaborative filtering problem in a generative probabilistic framework, treating individual user-item ratings as predictors of missing ratings. The final rating is estimated by fusing predictions from three sources: predictions based on ratings of the same item by other users, predictions based on different item ratings made by the same user, and, third, ratings predicted based on data from other but similar users rating other but similar items. Existing user-based and item-based approaches correspond to the two simple cases of our framework. The complete model is however more robust to data sparsity, because the different types of ratings are used in concert, while additional ratings from similar users towards similar items are employed as a background model to smooth the predictions. Experiments demonstrate that the proposed methods are indeed more robust against data sparsity and give better recommendations.
Jun Wang 0012, Arjen P. de Vries, Marcel J. T. Reinders
SIGIR3
2006 Maximum significance clustering of oligonucleotide microarrays
abstract
MOTIVATION: Affymetrix high-density oligonucleotide microarrays measure the expression of DNA transcripts using probesets, i.e. multiple probes per transcript. Usually, these multiple measurements are transformed into a single probeset expression level before data analysis proceeds; any information on variability is lost. In this paper we demonstrate how individual probe measurements can be used in a statistic for differential expression. Furthermore, we show how this statistic can serve as a criterion for clustering microarrays. RESULTS: A novel clustering algorithm using this maximum significance criterion is demonstrated to be more efficient with the measured data than competing techniques for dealing with repeated measurements, especially when the sample size is small.
Dick de Ridder, Frank J. T. Staal, Jacques J. M. van Dongen, Marcel J. T. Reinders
Bioinform.4
2006 Least absolute regression network analysis of the murine osteoblast differentiation network
abstract
Abstract Motivation: We propose a reverse engineering scheme to discover genetic regulation from genome-wide transcription data that monitors the dynamic transcriptional response after a change in cellular environment. The interaction network is estimated by solving a linear model using simultaneous shrinking of the least absolute weights and the prediction error. Results: The proposed scheme has been applied to the murine C2C12 cell-line stimulated to undergo osteoblast differentiation. Results show that our method discovers genetic interactions that display significant enrichment of co-citation in literature. More detailed study showed that the inferred network exhibits properties and hypotheses that are consistent with current biological knowledge. Availability: Software is freely available for academic use as a Matlab package called GENLAB: Contact: [email protected] Supplementary information: Additional data, results and figures can be found at
Eugene P. van Someren, B. L. T. Vaes, W. T. Steegenga, Anneke M. Sijbers, Koen J. Dechering, Marcel J. T. Reinders
Bioinform.6
2006 A comparison of univariate and multivariate gene selection techniques for classification of cancer datasets
abstract
BACKGROUND: Gene selection is an important step when building predictors of disease state based on gene expression data. Gene selection generally improves performance and identifies a relevant subset of genes. Many univariate and multivariate gene selection approaches have been proposed. Frequently the claim is made that genes are co-regulated (due to pathway dependencies) and that multivariate approaches are therefore per definition more desirable than univariate selection approaches. Based on the published performances of all these approaches a fair comparison of the available results can not be made. This mainly stems from two factors. First, the results are often biased, since the validation set is in one way or another involved in training the predictor, resulting in optimistically biased performance estimates. Second, the published results are often based on a small number of relatively simple datasets. Consequently no generally applicable conclusions can be drawn. RESULTS: In this study we adopted an unbiased protocol to perform a fair comparison of frequently used multivariate and univariate gene selection techniques, in combination with a ränge of classifiers. Our conclusions are based on seven gene expression datasets, across several cancer types. CONCLUSION: Our experiments illustrate that, contrary to several previous studies, in five of the seven datasets univariate selection approaches yield consistently better results than multivariate approaches. The simplest multivariate selection approach, the Top Scoring method, achieves the best results on the remaining two datasets. We conclude that the correlation structures, if present, are difficult to extract due to the small number of samples, and that consequently, overly-complex gene selection algorithms that attempt to extract these structures are prone to overtraining.
Carmen Lai, Marcel J. T. Reinders, Laura J. van't Veer, Lodewyk F. A. Wessels
BMC Bioinform.2
2006 The effect of oligonucleotide microarray data pre-processing on the analysis of patient-cohort studies
abstract
BACKGROUND: Intensity values measured by Affymetrix microarrays have to be both normalized, to be able to compare different microarrays by removing non-biological variation, and summarized, generating the final probe set expression values. Various pre-processing techniques, such as dChip, GCRMA, RMA and MAS have been developed for this purpose. This study assesses the effect of applying different pre-processing methods on the results of analyses of large Affymetrix datasets. By focusing on practical applications of microarray-based research, this study provides insight into the relevance of pre-processing procedures to biology-oriented researchers. RESULTS: Using two publicly available datasets, i.e., gene-expression data of 285 patients with Acute Myeloid Leukemia (AML, Affymetrix HG-U133A GeneChip) and 42 samples of tumor tissue of the embryonal central nervous system (CNS, Affymetrix HuGeneFL GeneChip), we tested the effect of the four pre-processing strategies mentioned above, on (1) expression level measurements, (2) detection of differential expression, (3) cluster analysis and (4) classification of samples. In most cases, the effect of pre-processing is relatively small compared to other choices made in an analysis for the AML dataset, but has a more profound effect on the outcome of the CNS dataset. Analyses on individual probe sets, such as testing for differential expression, are affected most; supervised, multivariate analyses such as classification are far less sensitive to pre-processing. CONCLUSION: Using two experimental datasets, we show that the choice of pre-processing method is of relatively minor influence on the final analysis outcome of large microarray studies whereas it can have important effects on the results of a smaller study. The data source (platform, tissue homogeneity, RNA quality) is potentially of bigger importance than the choice of pre-processing method.
Roel G. W. Verhaak, Frank J. T. Staal, Peter J. M. Valk, Bob Löwenberg, Marcel J. T. Reinders, Dick de Ridder
BMC Bioinform.5
2006 Detecting Statistically Significant Common Insertion Sites in Retroviral Insertional Mutagenesis Screens
abstract
Retroviral insertional mutagenesis screens, which identify genes involved in tumor development in mice, have yielded a substantial number of retroviral integration sites, and this number is expected to grow substantially due to the introduction of high-throughput screening techniques. The data of various retroviral insertional mutagenesis screens are compiled in the publicly available Retroviral Tagged Cancer Gene Database (RTCGD). Integrally analyzing these screens for the presence of common insertion sites (CISs, i.e., regions in the genome that have been hit by viral insertions in multiple independent tumors significantly more than expected by chance) requires an approach that corrects for the increased probability of finding false CISs as the amount of available data increases. Moreover, significance estimates of CISs should be established taking into account both the noise, arising from the random nature of the insertion process, as well as the bias, stemming from preferential insertion sites present in the genome and the data retrieval methodology. We introduce a framework, the kernel convolution (KC) framework, to find CISs in a noisy and biased environment using a predefined significance level while controlling the family-wise error (FWE) (the probability of detecting false CISs). Where previous methods use one, two, or three predetermined fixed scales, our method is capable of operating at any biologically relevant scale. This creates the possibility to analyze the CISs in a scale space by varying the width of the CISs, providing new insights in the behavior of CISs across multiple scales. Our method also features the possibility of including models for background bias. Using simulated data, we evaluate the KC framework using three kernel functions, the Gaussian, triangular, and rectangular kernel function. We applied the Gaussian KC to the data from the combined set of screens in the RTCGD and found that 53% of the CISs do not reach the significance threshold in this combined setting. Still, with the FWE under control, application of our method resulted in the discovery of eight novel CISs, which each have a probability less than 5% of being false detections.
Jeroen de Ridder, Anthony Uren, Jaap Kool, Marcel J. T. Reinders, Lodewyk F. A. Wessels
PLoS Comput. Biol.4
2006 Artifacts of Markov blanket filtering based on discretized features in small sample size applications
Theo A. Knijnenburg, Marcel J. T. Reinders, Lodewyk F. A. Wessels
Pattern Recognit. Lett.2
2006 Random subspace method for multivariate feature selection
Carmen Lai, Marcel J. T. Reinders, Lodewyk F. A. Wessels
Pattern Recognit. Lett.2
2005 Isophote Properties as Features for Object Detection
abstract
Usually, object detection is performed directly on (normalized) gray values or gray primitives like gradients or Haar-like features. In that case the learning of relationships between gray primitives, that describe the structure of the object, is the complete responsibility of the classifier. We propose to apply more knowledge about the image structure in the preprocessing step, by computing local isophote directions and curvatures, in order to supply the classifier with much more informative image structure features. However, a periodic feature space, like orientation, is unsuited for common classification methods. Therefore, we split orientation into two more suitable components. Experiments show that the isophote features result in better detection performance than intensities, gradients or Haar-like features.
Jeroen Lichtenauer, Emile A. Hendriks, Marcel J. T. Reinders
CVPR (2)3
2005 Self-organizing distributed collaborative filtering
abstract
We propose a fully decentralized collaborative filtering approach that is self-organizing and operates in a distributed way. The relevances between downloading files (items) are stored locally at these items in so called item-based buddy tables and are updated each time that the items are downloaded. We then propose to use the language model to build recommendations for the different users based on the buddy tables of those items a user has downloaded previously. We have tested and compared our distributed collaborative filtering approach to centralized collaborative filtering and showed that it has similar performance. It is therefore a promising technique to facilitate recommendations in peer-to-peer networks.
Jun Wang 0012, Marcel J. T. Reinders, Reginald L. Lagendijk, Johan A. Pouwelse
SIGIR2
2005 A protocol for building and evaluating predictors of disease state based on microarray data
abstract
MOTIVATION: Microarray gene expression data are increasingly employed to identify sets of marker genes that accurately predict disease development and outcome in cancer. Many computational approaches have been proposed to construct such predictors. However, there is, as yet, no objective way to evaluate whether a new approach truly improves on the current state of the art. In addition no 'standard' computational approach has emerged which enables robust outcome prediction. RESULTS: An important contribution of this work is the description of a principled training and validation protocol, which allows objective evaluation of the complete methodology for constructing a predictor. We review the possible choices of computational approaches, with specific emphasis on predictor choice and reporter selection strategies. Employing this training-validation protocol, we evaluated different reporter selection strategies and predictors on six gene expression datasets of varying degrees of difficulty. We demonstrate that simple reporter selection strategies (forward filtering and shrunken centroids) work surprisingly well and outperform partial least squares in four of the six datasets. Similarly, simple predictors, such as the nearest mean classifier, outperform more complex classifiers. Our training-validation protocol provides a robust methodology to evaluate the performance of new computational approaches and to objectively compare outcome predictions on different datasets.
Lodewyk F. A. Wessels, Marcel J. T. Reinders, Augustinus A. M. Hart, Cor J. Veenman, Hongyue Dai, Yudong D. He, Laura J. van't Veer
Bioinform.2
2005 The Nearest Subclass Classifier: A Compromise between the Nearest Mean and Nearest Neighbor Classifier
abstract
We present the Nearest Subclass Classifier (NSC), which is a classification algorithm that unifies the flexibility of the nearest neighbor classifier with the robustness of the nearest mean classifier. The algorithm is based on the Maximum Variance Cluster algorithm and, as such, it belongs to the class of prototype-based classifiers. The variance constraint parameter of the cluster algorithm serves to regularize the classifier, that is, to prevent overfitting. With a low variance constraint value, the classifier turns into the nearest neighbor classifier and, with a high variance parameter, it becomes the nearest mean classifier with the respective properties. In other words, the number of prototypes ranges from the whole training set to only one per class. In the experiments, we compared the NSC with regard to its performance and data set compression ratio to several other prototype-based methods. On several data sets, the NSC performed similarly to the k-nearest neighbor classifier, which is a well-established classifier in many domains. Also concerning storage requirements and classification speed, the NSC has favorable properties, so it gives a good compromise between classification performance and efficiency.
Cor J. Veenman, Marcel J. T. Reinders
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Edge-Based Image Restoration
abstract
In this paper, we propose a new image inpainting algorithm that relies on explicit edge information. The edge information is used both for the reconstruction of a skeleton image structure in the missing areas, as well as for guiding the interpolation that follows. The structure reconstruction part exploits different properties of the edges, such as the colors of the objects they separate, an estimate of how well one edge continues into another one, and the spatial order of the edges with respect to each other. In order to preserve both sharp and smooth edges, the areas delimited by the recovered structure are interpolated independently, and the process is guided by the direction of the nearby edges. The novelty of our approach lies primarily in exploiting explicitly the constraint enforced by the numerical interpretation of the sequential order of edges, as well as in the pixel filling method which takes into account the proximity and direction of edges. Extensive experiments are carried out in order to validate and compare the algorithm both quantitatively and qualitatively. They show the advantages of our algorithm and its readily application to real world cases.
Andrei Rares, Marcel J. T. Reinders, Jan Biemond
IEEE Trans. Image Process.2
2004 Video content representation on tiny devices
abstract
The perceptual satisfaction of a user watching video on a tiny mobile device is constrained by the display capability and network bandwidth. To maximize the user's perceptual satisfaction in this constrained environment, we propose a new method to represent the video content adaptively in real-time on tiny devices according to the user's attention. First, a sampling based dynamic attention model is proposed to obtain and maintain the user's attention in the video streams. Second, based on the most attended regions and sequences extracted, the attention based representation is introduced to achieve a higher perceptual satisfaction on a small display. Experiments with users show the effectiveness of our proposed method in a video surveillance application.
Jun Wang 0012, Marcel J. T. Reinders, Reginald L. Lagendijk, Jasper Lindenberg, Mohan Kankanhalli
ICME2
2003 Establishing motion correspondence using extended temporal scope
Cor J. Veenman, Marcel J. T. Reinders, Eric Backer
Artif. Intell.2
2003 Motion tracking as a constrained optimization problem
Cor J. Veenman, Marcel J. T. Reinders, Eric Backer
Pattern Recognit.2
2003 Multi-criterion optimization for genetic network modeling
Eugene P. van Someren, Lodewyk F. A. Wessels, Eric Backer, Marcel J. T. Reinders
Signal Process.4
2003 A cellular coevolutionary algorithm for image segmentation
abstract
Clustering is inherently a difficult problem, both with respect to the definition of adequate models as well as to the optimization of the models. We present a model for the cluster problem that does not need knowledge about the number of clusters a priori. This property is among others useful in the image segmentation domain, which we especially address. Further, we propose a cellular coevolutionary algorithm for the optimization of the model. Within this scheme multiple agents are placed in a regular two-dimensional (2-D) grid representing the image, which imposes neighboring relations on them. The agents cooperatively consider pixel migration from one agent to the other in order to improve the homogeneity of the ensemble of the image regions they represent. If the union of the regions of neighboring agents is homogeneous then the agents form alliances. On the other hand, if an agent discovers a deviant subject, it isolates the subject. In the experiments we show the effectiveness of the proposed method and compare it to other segmentation algorithms. The efficiency can easily be improved by exploiting the intrinsic parallelism of the proposed method.
Cor J. Veenman, Marcel J. T. Reinders, Eric Backer
IEEE Trans. Image Process.2
2002 Image sequence restoration in the presence of pathological motion and severe artifacts
abstract
Incorrect motion vectors represent the main reason why current image sequence restoration schemes fail. We present a scheme that 1) identifies areas that are likely to contain wrong motion vectors, 2) finds artifacts within these areas, and 3) restores these artifacts: Although temporal information is commonly used in nowadays restoration systems [7, 12], in this particular case the artifact restoration cannot rely on it due to the uncertain motion information. Hence, the restoration needs to rely on spatial information alone. The novel spatial restoration algorithm that we introduce here, performs a non-linear interpolation that preserves the edges surrounding the artifact area. In this way, the general structure of the image is reconstructed. The restoration algorithm is demonstrated on both artificial and real life examples, and the advantages of the proposed edge-based restoration are highlighted. In addition, results are shown for the proposed complete image sequence restoration scheme.
Andrei Rares, Marcel J. T. Reinders, Jan Biemond
ICASSP2
2002 A spatiotemporal image sequence restoration algorithm
abstract
Temporal restoration of image sequences may fail because of erroneously estimated motion. Motion estimation failures may occur for several reasons -complicated motion, severe artefacts, or a combination of both. We present an algorithm that tries to overcome at least the problems related to severe artefacts. It consists of a spatial restoration, which is supposed to recover the general image structure within the missing areas, followed by a temporal restoration. Performing spatial restoration first has the advantage that motion estimation during temporal restoration can rely on the spatially restored frames. The efficiency of the spatial restoration scheme is demonstrated for both artificial and real life examples. Restoration results when using only temporal restoration are presented for the real life sequences. Finally, we compare both restoration methods with the results of the proposed spatiotemporal algorithm and show its superior behaviour.
Andrei Rares, Marcel J. T. Reinders, Reginald L. Lagendijk, Jan Biemond
ICIP (2)2
2002 Scale-invariant segmentation of dynamic contrast-enhanced perfusion MR images with inherent scale selection
abstract
Abstract Selection of the best set of scales is problematic when developing signal‐driven approaches for pixel‐based image segmentation. Often, different possibly conflicting criteria need to be fulfilled in order to obtain the best trade‐off between uncertainty (variance) and location accuracy. The optimal set of scales depends on several factors: the noise level present in the image material, the prior distribution of the different types of segments, the class‐conditional distributions associated with each type of segment as well as the actual size of the (connected) segments. We analyse, theoretically and through experiments, the possibility of using the overall and class‐conditional error rates as criteria for selecting the optimal sampling of the linear and morphological scale spaces. It is shown that the overall error rate is optimized by taking the prior class distribution in the image material into account. However, a uniform (ignorant) prior distribution ensures constant class‐conditional error rates. Consequently, we advocate for a uniform prior class distribution when an uncommitted, scale‐invariant segmentation approach is desired. Experiments with a neural net classifier developed for segmentation of dynamic magnetic resonance (MR) images, acquired with a paramagnetic tracer, support the theoretical results. Furthermore, the experiments show that the addition of spatial features to the classifier, extracted from the linear or morphological scale spaces, improves the segmentation result compared to a signal‐driven approach based solely on the dynamic MR signal. The segmentation results obtained from the two types of features are compared using two novel quality measures that characterize spatial properties of labelled images. Copyright © 2002 John Wiley & Sons, Ltd.
Jasper P. Janssen, Michael Egmont-Petersen, Emile A. Hendriks, Marcel J. T. Reinders, Rob J. van der Geest, P. C. W. Hogendoorn, Johan H. C. Reiber
Comput. Animat. Virtual Worlds4
2002 A Maximum Variance Cluster Algorithm
abstract
We present a partitional cluster algorithm that minimizes the sum-of-squared-error criterion while imposing a hard constraint on the cluster variance. Conceptually, hypothesized clusters act in parallel and cooperate with their neighboring clusters in order to minimize the criterion and to satisfy the variance constraint. In order to enable the demarcation of the cluster neighborhood without crucial parameters, we introduce the notion of foreign cluster samples. Finally, we demonstrate a new method for cluster tendency assessment based on varying the variance constraint parameter.
Cor J. Veenman, Marcel J. T. Reinders, Eric Backer
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Facial Landmarks Localization Based on Fuzzy and Gabor Wavelet Graph Matching
abstract
Proposes a method that automatically finds human faces as well as its landmark points in color images based on a fuzzy analysis. The proposed approach first uses color information to detect face candidate regions and then uses a fuzzy analysis of the color, shape, symmetry and interior facial features. A deformable Gabor wavelet graph matching is used to locate the facial landmark points describing the face. The latter allows for size and orientation variation since the search for landmark points allows for affine transformations as well as local deformations of the Gabor wavelet graph. The search is performed using a genetic algorithm that is essential because it effectively searches the solution space. Results based on the proposed method are included to verify the effectiveness of the proposed approach.
Resmana Lim, Marcel J. T. Reinders
FUZZ-IEEE2
2001 Complex event classification in degraded image sequences
abstract
We address the problem of motion estimation failure in degraded image sequences. This failure is caused by some complex events that take place in the image. As a result, the sequence of operations that rely on the motion estimation process, such as motion compensation and motion picture restoration, fail as well. The statistical analysis of the complex event areas indicates that it is possible to discriminate between the complex events resulting from complicated object motion and the ones resulting from image artefacts. An analysis scheme based on segment matching is proposed for the task of classifying the detected complex event areas.
Andrei Rares, Marcel J. T. Reinders, Jan Biemond
ICIP (1)2
2001 Determination of Binding Amino Acids Based on Random Peptide Array Screening Data
Peter J. van der Veen, Lodewyk F. A. Wessels, J. W. Slootstra, R. H. Meloen, Marcel J. T. Reinders, Hans Hellendoorn
WABI5
2001 Utility map reconstruction
abstract
In this paper we describe a new system for the reconstruction of geometric map elements as developed in the Dutch TopSpin-PNEM project 'Knowledge-based conversion of utility maps'. In the knowledge base map conversion process vectorized map elements coming from an image interpretation system must be reconstructed in a geometrically accurate topographic base map. This is done using a map correspondence between the utility map and the base map based on mutual topography and relational information supplied by a map interpretation module. The map interpretation module uses geometrical inference to deduce the relational information, while applying a novel conflict resolution strategy. The map elements are accurately reconstructed, relationally and geometrically, using the map correspondence, the relations and several object-specific functions.
J. P. De Knecht, John G. M. Schavemaker, Marcel J. T. Reinders, Albert M. Vossepoel
Int. J. Geogr. Inf. Sci.3
2001 Resolving Motion Correspondence for Densely Moving Points
abstract
Studies the motion correspondence problem for which a diversity of qualitative and statistical solutions exist. We concentrate on qualitative modeling, especially in situations where assignment conflicts arise either because multiple features compete for one detected point or because multiple detected points fit a single feature point. We leave out the possibility of point track initiation and termination because that principally conflicts with allowing for temporary point occlusion. We introduce individual, combined, and global motion models and fit existing qualitative solutions in this framework. Additionally, we present a tracking algorithm that satisfies these-possibly constrained-models in a greedy matching sense, including an effective way to handle detection errors and occlusion. The performance evaluation shows that the proposed algorithm outperforms existing greedy matching algorithms. Finally, we describe an extension to the tracker that enables automatic initialization of the point tracks. Several experiments show that the extended algorithm is efficient, hardly sensitive to its few parameters, and qualitatively better than other algorithms, including the presumed optimal statistical multiple hypothesis tracker.
Cor J. Veenman, Marcel J. T. Reinders, Eric Backer
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Object Tracking by Adaptive Modeling
abstract
This paper addresses the problem of object tracking in image sequences. The approach taken is based upon adaptive statistical models. An object selected in a frame by a user is tracked throughout the sequence by using a blob-like description of its features. The object features are continuously updated by using the on-line version of the expectation-maximization algorithm. The proposed object description results in a flexible representation.
Andrei Rares, Marcel J. T. Reinders
ICIP2
2000 Detecting Generic Low-Level Features in Images
abstract
Traditional definitions of edges and corners are limited in the sense that they restrict their application. In this paper we generalize the concepts of edges and corners to 1D features and 2D features, respectively. The phase congruency theory is chosen as a mathematical system to model the generalized concepts. The basics of the phase congruency theory along with the local energy model are described. The extraction schemes of 2D features from 2D gray level images are studied and discussed as a demonstration of our generalization idea. The performance evaluation and the preliminary improvement show that this kind of generalization on features and the modeling are successful and promising.
Bang Jun Lei, Emile A. Hendriks, Marcel J. T. Reinders
ICPR3
2000 Eye Activity Detection and Recognition Using Morphological Scale-Space Decomposition
abstract
Automatic recovery of eye gestures from image sequences is one of the important topics for face recognition and model-based coding of videophone sequences. Usually, complicated models of the eye and its motion are used. In this paper an eye gesture parameter estimation is described. A previously published automatic eye detection/tracking algorithm, based on template matching, is used for the eye pose detection. The eye gesture analysis is realised with a mathematical morphology scale-space approach, forming spatio-temporal curves out of scale measurement statistics. The resulting curves provide a direct measure of the eye gesture, which can then be used as an eye animation parameter. Experimental results demonstrate the efficiency and robustness of the method.
Ilse Ravyse, Hichem Sahli, Jan Cornelis 0001, Marcel J. T. Reinders
ICPR4
2000 Linear Modeling of Genetic Networks from Experimental Data
Eugene P. van Someren, Lodewyk F. A. Wessels, Marcel J. T. Reinders
ISMB3
2000 Competitive Segmentation: A Struggle for Image Space
Cor J. Veenman, Marcel J. T. Reinders, Eric Backer
PPSN2
2000 Image sharpening by morphological filtering
John G. M. Schavemaker, Marcel J. T. Reinders, Jan J. Gerbrands, Eric Backer
Pattern Recognit.2
2000 On-line detection of red blood cell shape using deformable templates
P. J. H. Bronkorsta, Marcel J. T. Reinders, Emile A. Hendriks, J. Grimbergen, Rob M. Heethaar, G. J. Brakenhoff
Pattern Recognit. Lett.2
1999 Information processing for intelligent molecular diagnosis
Lodewyk F. A. Wessels, Eugene P. van Someren, Marcel J. T. Reinders
Pattern Recognit. Lett.3
1998 A Fast and Robust Point Tracking Algorithm
Cor J. Veenman, Emile A. Hendriks, Marcel J. T. Reinders
ICIP (3)3
1996 Locating Facial Features in Image Sequences using Neural Networks
abstract
The paper describes a method for the automatic location of facial features, such as eyes, nose and mouth, in image sequences using a neural network approach. It is shown that by modeling the feature sought as a structural assembly of micro-features, and by using a probabilistic interpretation of neural network outputs, it is possible to construct a location system that is more robust than a location system which uses the feature as a single entity. With this micro-feature approach, not only can the position of the features be found but shape of the features can be obtained as well.
Marcel J. T. Reinders, R. W. C. Koch, Jan J. Gerbrands
FG1
1995 Facial feature localization and adaptation of a generic face model for model-based coding
Marcel J. T. Reinders, P. J. L. van Beek, Bülent Sankur, Jan C. A. van der Lubbe
Signal Process. Image Commun.1
1993 Tracking of global motion and facial expressions of a human face in image sequences
abstract
We present a system in which the global motion (3D rotation and translation) and the local motion (facial expressions) of the face of a talking person are estimated automatically from an input image sequence. First, the shape of the facial features, such as eyes and mouth, are robustly extracted from the images. Then, based on the extracted shape of the facial features, the global motion and facial expressions are estimated. No human assistance is necessary throughout the process. The system relies on the use of a priori knowledge about the scene and facial motions. In the feature extraction scheme this a priori knowledge is modeled by a descriptive tree of the scene and geometric shape representations of the facial features. The global motion can be easily obtained from these facial features. In the local motion estimation scheme the a priori knowledge is modeled in a belief network, in which knowledge about muscle actuators is represented, e.g., the interactions between the muscle actuators and their visible manifestations. A few experiments are included to illustrate the system.
Marcel J. T. Reinders, F. A. Odijk, Jan C. A. van der Lubbe, Jan J. Gerbrands
VCIP1
1992 Transformation of a general 3D facial model to an actual scene face
abstract
Accurate modeling and adaptation of scene objects play a crucial role in model-based coding schemes. In this paper a method is presented which transforms a general model of a face to the actual face viewed in the scene. The transformation is based on the facial contours (face border, eyes and mouth) and consists of an affine transformation (for all vertices) and local transformations for each individual vertex. Local transformations are executed taking into consideration the elastic properties of the model.>
Marcel J. T. Reinders, Bülent Sankur, Jan C. A. van der Lubbe
ICPR (3)1