VLDB 2026 Research / reviewers in the wild / expert
Yvan Saeys
dblp:s/YvanSaeys
· DBLP profile ↗
59ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-0415-1506ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Theory of computation · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable analysis of whole slide spatial proteomics with HarpyabstractMOTIVATION: Current spatial proteomics data analysis workflows are limited in efficiency and scalability when applied to gigapixel sized datasets. Moreover, they often lack extensive quality control tools and exhibit limited interoperability with existing spatial omics analysis ecosystems. RESULTS: We introduce Harpy, a new Python workflow capable of accelerated processing of large spatial proteomics datasets. We demonstrate the utility of Harpy on four datasets and show that it can rapidly apply state-of-the-art segmentation and feature extraction via parallel processing. Each analysis step is accompanied by appropriate quality control steps. Scalable clustering of cells and pixels allows identification of cell types, processed up to 27 times faster than previously reported. Processing and visualization can be performed locally or on high-performance computing servers. Additionally, Harpy integrates well with existing spatial single-cell analysis tools in the Python and R software ecosystem. AVAILABILITY AND IMPLEMENTATION: Harpy is available on GitHub at https://github.com/saeyslab/harpy and archived on Zenodo at https://doi.org/10.5281/zenodo.15546703. Benjamin Rombaut 0001, Arne Defauw, Frank Vernaillen, Julien Mortier, Evelien Van Hamme, Sofie Van Gassen, Ruth Seurinck, Yvan Saeys |
Bioinform. | 8 |
| 2026 | Feature subset weighting for distance-based supervised learning
Adnan Theerens, Yvan Saeys, Chris Cornelis |
Pattern Recognit. | 2 |
| 2025 | Evaluation of out-of-distribution detection methods for data shifts in single-cell transcriptomicsabstractAutomatic cell-type annotation methods assign cell-type labels to new, unlabeled datasets by leveraging relationships from a reference RNA-seq atlas. However, new datasets may include labels absent from the reference dataset or exhibit feature distributions that diverge from it. These scenarios can significantly affect the reliability of cell type predictions, a factor often overlooked in current automatic annotation methods. The field of out-of-distribution detection (OOD), primarily focused on computer vision, addresses the identification of instances that differ from the training distribution. Therefore, the implementation of OOD methods in the context of novel cell type annotation and data shift detection for single-cell transcriptomics may enhance annotation accuracy and trustworthiness. We evaluate six OOD detection methods: LogitNorm, MC dropout, Deep Ensembles, Energy-based OOD, Deep NN, and Posterior networks, for their annotation and OOD detection performance in both synthetical and real-life application settings. We show that OOD detection methods can accurately identify novel cell types and demonstrate potential to detect significant data shifts in non-integrated datasets. Moreover, we find that integration of the OOD datasets does not interfere with OOD detection of novel cell types. Lauren Theunissen, Thomas Mortier, Yvan Saeys, Willem Waegeman |
Briefings Bioinform. | 3 |
| 2024 | Efficient cytometry analysis with FlowSOM in Python boosts interoperability with other single-cell toolsabstractMOTIVATION: We describe a new Python implementation of FlowSOM, a clustering method for cytometry data. RESULTS: This implementation is faster than the original version in R, better adapted to work with single-cell omics data including integration with current single-cell data structures and includes all the original visualizations, such as the star and pie plot. AVAILABILITY AND IMPLEMENTATION: The FlowSOM Python implementation is freely available on GitHub: https://github.com/saeyslab/FlowSOM_Python. Artuur Couckuyt, Benjamin Rombaut 0001, Yvan Saeys, Sofie Van Gassen |
Bioinform. | 3 |
| 2024 | Uncertainty-aware single-cell annotation with a hierarchical reject optionabstractMOTIVATION: Automatic cell type annotation methods assign cell type labels to new datasets by extracting relationships from a reference RNA-seq dataset. However, due to the limited resolution of gene expression features, there is always uncertainty present in the label assignment. To enhance the reliability and robustness of annotation, most machine learning methods address this uncertainty by providing a full reject option, i.e. when the predicted confidence score of a cell type label falls below a user-defined threshold, no label is assigned and no prediction is made. As a better alternative, some methods deploy hierarchical models and consider a so-called partial rejection by returning internal nodes of the hierarchy as label assignment. However, because a detailed experimental analysis of various rejection approaches is missing in the literature, there is currently no consensus on best practices. RESULTS: We evaluate three annotation approaches (i) full rejection, (ii) partial rejection, and (iii) no rejection for both flat and hierarchical probabilistic classifiers. Our findings indicate that hierarchical classifiers are superior when rejection is applied, with partial rejection being the preferred rejection approach, as it preserves a significant amount of label information. For optimal rejection implementation, the rejection threshold should be determined through careful examination of a method's rejection behavior. Without rejection, flat and hierarchical annotation perform equally well, as long as the cell type hierarchy accurately captures transcriptomic relationships. AVAILABILITY AND IMPLEMENTATION: Code is freely available at https://github.com/Latheuni/Hierarchical_reject and https://doi.org/10.5281/zenodo.10697468. Lauren Theunissen, Thomas Mortier, Yvan Saeys, Willem Waegeman |
Bioinform. | 3 |
| 2024 | Evaluating feature attribution methods in the image domainabstractAbstract Feature attribution maps are a popular approach to highlight the most important pixels in an image for a given prediction of a model. Despite a recent growth in popularity and available methods, the objective evaluation of such attribution maps remains an open problem. Building on previous work in this domain, we investigate existing quality metrics and propose new variants of metrics for the evaluation of attribution maps. We confirm a recent finding that different quality metrics seem to measure different underlying properties of attribution maps, and extend this finding to a larger selection of attribution methods, quality metrics, and datasets. We also find that metric results on one dataset do not necessarily generalize to other datasets, and methods with desirable theoretical properties do not necessarily outperform computationally cheaper alternatives in practice. Based on these findings, we propose a general benchmarking approach to help guide the selection of attribution methods for a given use case. Implementations of attribution metrics and our experiments are available online ( https://github.com/arnegevaert/benchmark-general-imaging ). Graphical abstract Arne Gevaert, Axel-Jan Rousseau, Thijs Becker, Dirk Valkenborg, Tijl De Bie, Yvan Saeys |
Mach. Learn. | 6 |
| 2024 | An Introduction to Adversarially Robust Deep LearningabstractThe widespread success of deep learning in solving machine learning problems has fueled its adoption in many fields, from speech recognition to drug discovery and medical imaging. However, deep learning systems are extremely fragile: imperceptibly small modifications to their input data can cause the models to produce erroneous output. It is very easy to generate such adversarial perturbations even for state-of-the-art models, yet immunization against them has proven exceptionally challenging. Despite over a decade of research on this problem, our solutions are still far from satisfactory and many open problems remain. In this work, we survey some of the most important contributions in the field of adversarial robustness. We pay particular attention to the reasons why past attempts at improving robustness have been insufficient, and we identify several promising areas for future research. Jonathan Peck, Bart Goossens, Yvan Saeys |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Efficient Approximation of Asymmetric Shapley Values Using Functional DecompositionabstractAbstract Asymmetric Shapley values (ASVs) are an extension of Shapley values that allow a user to incorporate partial causal knowledge into the explanation process. Unfortunately, computing ASVs requires sampling permutations, which quickly becomes computationally expensive. We propose A-PDD-SHAP, an algorithm that employs a functional decomposition approach to approximate ASVs at a speed orders of magnitude faster compared to permutation sampling, which significantly reduces the amortized complexity of computing ASVs when many explanations are needed. Apart from this, once the A-PDD-SHAP model is trained, it can be used to compute both symmetric and asymmetric Shapley values without having to re-train or re-sample, allowing for very efficient comparisons between different types of explanations. Arne Gevaert, Anna Saranti, Andreas Holzinger, Yvan Saeys |
CD-MAKE | 4 |
| 2022 | The Curse Revisited: When are Distances Informative for the Ground Truth in Noisy High-Dimensional Data?abstractDistances between data points are widely used in machine learning applications. Yet, when corrupted by noise, these distances—and thus the models based upon them—may lose their usefulness in high dimensions. Indeed, the small marginal effects of the noise may then accumulate quickly, shifting empirical closest and furthest neighbors away from the ground truth. In this paper, we exactly characterize such effects in noisy high-dimensional data using an asymptotic probabilistic expression. Previously, it has been argued that neighborhood queries become meaningless and unstable when distance concentration occurs, which means that there is a poor relative discrimination between the furthest and closest neighbors in the data. However, we conclude that this is not necessarily the case when we decompose the data in a ground truth—which we aim to recover—and noise component. More specifically, we derive that under particular conditions, empirical neighborhood relations affected by noise are still likely to be truthful even when distance concentration occurs. We also include thorough empirical verification of our results, as well as interesting experiments in which our derived ‘phase shift’ where neighbors become random or not turns out to be identical to the phase shift where common dimensionality reduction methods perform poorly or well for recovering low-dimensional reconstructions of high-dimensional data with dense noise. Robin Vandaele, Bo Kang, Tijl De Bie, Yvan Saeys |
AISTATS | 4 |
| 2022 | Distilling Deep RL Models Into Interpretable Neuro-Fuzzy SystemsabstractDeep Reinforcement Learning uses a deep neural network to encode a policy, which achieves very good performance in a wide range of applications but is widely regarded as a black box model. A more interpretable alternative to deep networks is given by neuro-fuzzy controllers. Unfortunately, neuro-fuzzy controllers often need a large number of rules to solve relatively simple tasks, making them difficult to interpret. In this work, we present an algorithm to distill the policy from a deep Q-network into a compact neuro-fuzzy controller. This allows us to train compact neuro-fuzzy controllers through distillation to solve tasks that they are unable to solve directly, combining the flexibility of deep reinforcement learning and the interpretability of compact rule bases. We demonstrate the algorithm on three well-known environments from OpenAI Gym, where we nearly match the performance of a DQN agent using only 2 to 6 fuzzy rules. Arne Gevaert, Jonathan Peck, Yvan Saeys |
FUZZ-IEEE | 3 |
| 2022 | Topologically Regularized Data Embeddings
Robin Vandaele, Bo Kang, Jefrey Lijffijt, Tijl De Bie, Yvan Saeys |
ICLR | 5 |
| 2021 | A study on the calibration of fingerprint classifiersabstractFingerprint classification is a frequent approach to deal with very large scale databases in fingerprint recognition. In the last few years, several proposals based on Convolutional Neural Networks have pushed state of the art results even further. However, it has also been proven that such networks are prone to be overconfident in the predictions of the classes, which may have an impact on their performance. This paper aims to study the problem from a systematic point of view. First, it is determined that the most common network to classify fingerprints does suffer from badly calibrated predictions. Second, two calibration methods (temperature scaling and Dirichlet calibration) are applied to correct for this tendency. Third, a modified search strategy is proposed, which makes use of the calibrated class probabilities predicted by the classifier to further reduce the penetration rate and avoid the negative impact of impostor input fingerprints. Fourth, all the proposals are evaluated on five datasets, which combine synthetic and real fingerprints of different qualities. Dirichlet calibration led to improved predicted class probabilities, which in turn allowed for further reduction of the penetration, while maintaining a good trade-off with respect to the false rejection rate. Daniel Peralta, Maxim Lippeveld, Yvan Saeys |
IEEE BigData | 4 |
| 2021 | Stable topological signatures for metric trees through graph approximationsabstractThe rising field of Topological Data Analysis (TDA) provides a new approach to learning from data through persistence diagrams, which are topological signatures that quantify topological properties of data in a comparable manner. For point clouds, these diagrams are often derived from the Vietoris-Rips filtration—based on the metric equipped on the data—which allows one to deduce topological patterns such as components and cycles of the underlying space. In metric trees these diagrams often fail to capture other crucial topological properties, such as the present leaves and multifurcations. Prior methods and results for persistent homology attempting to overcome this issue mainly target Rips graphs, which are often unfavorable in case of non-uniform density across our point cloud. We therefore introduce a new theoretical foundation for learning a wider variety of topological patterns through any given graph. Given particular powerful functions defining persistence diagrams to summarize topological patterns, including the normalized centrality or eccentricity, we prove a new stability result, explicitly bounding the bottleneck distance between the true and empirical diagrams for metric trees. This bound is tight if the metric distortion obtained through the graph and its maximal edge-weight are small. Through a case study of gene expression data, we demonstrate that our newly introduced diagrams provide novel quality measures and insights into cell trajectory inference. Robin Vandaele, Bastian Rieck, Yvan Saeys, Tijl De Bie |
Pattern Recognit. Lett. | 3 |
| 2020 | Graph Approximations to Geodesics on Metric GraphsabstractIn machine learning, high-dimensional point clouds are often assumed to be sampled from a topological space of which the intrinsic dimension is significantly lower than the representation dimension. Proximity graphs, such as the Rips graph or kNN graph, are often used as an intermediate representation to learn or visualize topological and geometrical properties of this space. The key idea behind this approach is that distances on the graph preserve the geodesic distances on the unknown space well, and as such, can be used to infer local and global geometric patterns of this space. Prior results provide us with conditions under which these distances are well-preserved for geodesically convex, smooth, compact manifolds. Yet, proximity graphs are ideal representations for a much broader class of spaces, such as metric graphs, i.e., graphs embedded in the Euclidean space. It turns out-as proven in this paper-that these existing conditions cannot be straightforwardly adapted to these spaces. In this work, we provide novel, flexible, and insightful characteristics and results for topological pattern recognition of metric graphs to bridge this gap. Robin Vandaele, Yvan Saeys, Tijl De Bie |
ICPR | 2 |
| 2020 | TinGa: fast and flexible trajectory inference with Growing Neural GasabstractMOTIVATION: During the last decade, trajectory inference (TI) methods have emerged as a novel framework to model cell developmental dynamics, most notably in the area of single-cell transcriptomics. At present, more than 70 TI methods have been published, and recent benchmarks showed that even state-of-the-art methods only perform well for certain trajectory types but not others. RESULTS: In this work, we present TinGa, a new TI model that is fast and flexible, and that is based on Growing Neural Graphs. We performed an extensive comparison of TinGa to five state-of-the-art methods for TI on a set of 250 datasets, including both synthetic as well as real datasets. Overall, TinGa improves the state-of-the-art by producing accurate models (comparable to or an improvement on the state-of-the-art) on the whole spectrum of data complexity, from the simplest linear datasets to the most complex disconnected graphs. In addition, TinGa obtained the fastest execution times, showing that our method is thus one of the most versatile methods up to date. AVAILABILITY AND IMPLEMENTATION: R scripts for running TinGa, comparing it to top existing methods and generating the figures of this article are available at https://github.com/Helena-todd/TinGa. Helena Todorov, Robrecht Cannoodt, Wouter Saelens, Yvan Saeys |
Bioinform. | 4 |
| 2020 | Detecting adversarial manipulation using inductive Venn-ABERS predictorsabstractInductive Venn-ABERS predictors (IVAPs) are a type of probabilistic predictors with the theoretical guarantee that their predictions are perfectly calibrated. In this paper, we propose to exploit this calibration property for the detection of adversarial examples in binary classification tasks. By rejecting predictions if the uncertainty of the IVAP is too high, we obtain an algorithm that is both accurate on the original test set and resistant to adversarial examples. This robustness is observed on adversarials for the underlying model as well as adversarials that were generated by taking the IVAP into account. The method appears to offer competitive robustness compared to the state-of-the-art in adversarial defense yet it is computationally much more tractable. Jonathan Peck, Bart Goossens, Yvan Saeys |
Neurocomputing | 3 |
| 2020 | Mining Topological Structure in Graphs through Forest RepresentationsabstractWe consider the problem of inferring simplified topological substructures—which we term backbones—in metric and non-metric graphs. Intuitively, these are subgraphs with ‘few’ nodes, multifurcations, and cycles, that model the topology of the original graph well. We present a multistep procedure for inferring these backbones. First, we encode local (geometric) information of each vertex in the original graph by means of the boundary coefficient (BC) to identify ‘core’ nodes in the graph. Next, we construct a forest representation of the graph, termed an f-pine, that connects every node of the graph to a local ‘core’ node. The final backbone is then inferred from the f-pine through CLOF (Constrained Leaves Optimal subForest), a novel graph optimization problem we introduce in this paper. On a theoretical level, we show that CLOF is NP-hard for general graphs. However, we prove that CLOF can be efficiently solved for forest graphs, a surprising fact given that CLOF induces a nontrivial monotone submodular set function maximization problem on tree graphs. This result is the basis of our method for mining backbones in graphs through forest representation. We qualitatively and quantitatively confirm the applicability, effectiveness, and scalability of our method for discovering backbones in a variety of graph-structured data, such as social networks, earthquake locations scattered across the Earth, and high-dimensional cell trajectory data. Robin Vandaele, Yvan Saeys, Tijl De Bie |
J. Mach. Learn. Res. | 2 |
| 2019 | Detecting adversarial examples with inductive Venn-ABERS predictors
Jonathan Peck, Bart Goossens, Yvan Saeys |
ESANN | 3 |
| 2019 | Weight selection strategies for ordered weighted average based fuzzy rough sets
Sarah Vluymans, Neil Mac Parthaláin, Chris Cornelis, Yvan Saeys |
Inf. Sci. | 4 |
| 2019 | Mining the Enriched Subgraphs for Specific Vertices in a Biological GraphabstractIn this paper, we present a subgroup discovery method to find subgraphs in a graph that are associated with a given set of vertices. The association between a subgraph pattern and a set of vertices is defined by its significant enrichment based on a Bonferroni-corrected hypergeometric probability value. This interestingness measure requires a dedicated pruning procedure to limit the number of subgraph matches that must be calculated. The presented mining algorithm to find associated subgraph patterns in large graphs is therefore designed to efficiently traverse the search space. We demonstrate the operation of this method by applying it on three biological graph data sets and show that we can find associated subgraphs for a biologically relevant set of vertices and that the found subgraphs themselves are biologically interesting. Pieter Meysman, Yvan Saeys, Ehsan Sabaghian, Wout Bittremieux, Yves Van de Peer, Bart Goethals, Kris Laukens |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | Local Topological Data Analysis to Uncover the Global Structure of Data Approaching Graph-Structured Topologies
Robin Vandaele, Tijl De Bie, Yvan Saeys |
ECML/PKDD (2) | 3 |
| 2018 | SpliceRover: interpretable convolutional neural networks for improved splice site predictionabstractMotivation: During the last decade, improvements in high-throughput sequencing have generated a wealth of genomic data. Functionally interpreting these sequences and finding the biological signals that are hallmarks of gene function and regulation is currently mostly done using automated genome annotation platforms, which mainly rely on integrated machine learning frameworks to identify different functional sites of interest, including splice sites. Splicing is an essential step in the gene regulation process, and the correct identification of splice sites is a major cornerstone in a genome annotation system. Results: In this paper, we present SpliceRover, a predictive deep learning approach that outperforms the state-of-the-art in splice site prediction. SpliceRover uses convolutional neural networks (CNNs), which have been shown to obtain cutting edge performance on a wide variety of prediction tasks. We adapted this approach to deal with genomic sequence inputs, and show it consistently outperforms already existing approaches, with relative improvements in prediction effectiveness of up to 80.9% when measured in terms of false discovery rate. However, a major criticism of CNNs concerns their 'black box' nature, as mechanisms to obtain insight into their reasoning processes are limited. To facilitate interpretability of the SpliceRover models, we introduce an approach to visualize the biologically relevant information learnt. We show that our visualization approach is able to recover features known to be important for splice site prediction (binding motifs around the splice site, presence of polypyrimidine tracts and branch points), as well as reveal new features (e.g. several types of exclusion patterns near splice sites). Availability and implementation: SpliceRover is available as a web service. The prediction tool and instructions can be found at http://bioit2.irc.ugent.be/splicerover/. Supplementary information: Supplementary data are available at Bioinformatics online. Jasper Zuallaert, Fréderic Godin, Mijung Kim, Arne Soete, Yvan Saeys, Wesley De Neve |
Bioinform. | 5 |
| 2018 | On the use of convolutional neural networks for robust classification of multiple fingerprint capturesabstractFingerprint classification is one of the most common approaches to accelerate the identification in large databases of fingerprints. Fingerprints are grouped into disjoint classes, so that an input fingerprint is compared only with those belonging to the predicted class, reducing the penetration rate of the search. The classification procedure usually starts by the extraction of features from the fingerprint image, frequently based on visual characteristics. In this work, we propose an approach to fingerprint classification using convolutional neural networks, which avoid the necessity of an explicit feature extraction process by incorporating the image processing within the training of the classifier. Furthermore, such an approach is able to predict a class even for low-quality fingerprints that are rejected by commonly used algorithms, such as FingerCode. The study gives special importance to the robustness of the classification for different impressions of the same fingerprint, aiming to minimize the penetration in the database. In our experiments, convolutional neural networks yielded better accuracy and penetration rate than state-of-the-art classifiers based on explicit feature extraction. The tested networks also improved on the runtime, as a result of the joint optimization of both feature extraction and classification. Daniel Peralta, Isaac Triguero, Salvador García 0001, Yvan Saeys, José Manuel Benítez 0001, Francisco Herrera |
Int. J. Intell. Syst. | 4 |
| 2018 | Multi-label classification using a fuzzy rough neighborhood consensus
Sarah Vluymans, Chris Cornelis, Francisco Herrera, Yvan Saeys |
Inf. Sci. | 4 |
| 2018 | Dynamic affinity-based classification of multi-class imbalanced data with one-versus-one decomposition: a fuzzy rough set approach
Sarah Vluymans, Alberto Fernández 0001, Yvan Saeys, Chris Cornelis, Francisco Herrera |
Knowl. Inf. Syst. | 3 |
| 2017 | Interpretable convolutional neural networks for effective translation initiation site predictionabstractThanks to rapidly evolving sequencing techniques, the amount of genomic data at our disposal is growing increasingly large. Determining the gene structure is a fundamental requirement to effectively interpret gene function and regulation. An important part in that determination process is the identification of translation initiation sites. In this paper, we propose a novel approach for automatic prediction of translation initiation sites, leveraging convolutional neural networks that allow for automatic feature extraction. Our experimental results demonstrate that we are able to improve the state-of-the-art approaches with a decrease of 75.2% in false positive rate and with a decrease of 24.5% in error rate on chosen datasets. Furthermore, an in-depth analysis of the decision-making process used by our predictive model shows that our neural network implicitly learns biologically relevant features from scratch, without any prior knowledge about the problem at hand, such as the Kozak consensus sequence, the influence of stop and start codons in the sequence and the presence of donor splice site patterns. In summary, our findings yield a better understanding of the internal reasoning of a convolutional neural network when applying such a neural network to genomic data. Jasper Zuallaert, Mijung Kim, Yvan Saeys, Wesley De Neve |
BIBM | 3 |
| 2017 | Lower bounds on the robustness to adversarial perturbationsabstractThe input-output mappings learned by state-of-the-art neural networks are significantly discontinuous. It is possible to cause a neural network used for image recognition to misclassify its input by applying very specific, hardly perceptible perturbations to the input, called adversarial perturbations. Many hypotheses have been proposed to explain the existence of these peculiar samples as well as several methods to mitigate them. A proven explanation remains elusive, however. In this work, we take steps towards a formal characterization of adversarial perturbations by deriving lower bounds on the magnitudes of perturbations necessary to change the classification of neural networks. The bounds are experimentally verified on the MNIST and CIFAR-10 data sets. Jonathan Peck, Joris Roels, Bart Goossens, Yvan Saeys |
NIPS | 4 |
| 2017 | Distributed incremental fingerprint identification with reduced database penetration rate using a hierarchical classification based on feature fusion and selection
Daniel Peralta, Isaac Triguero, Salvador García 0001, Yvan Saeys, José Manuel Benítez 0001, Francisco Herrera |
Knowl. Based Syst. | 4 |
| 2016 | Decreasing Time Consumption of Microscopy Image Segmentation Through Parallel Processing on the GPU
Joris Roels, Jonas De Vylder, Yvan Saeys, Bart Goossens, Wilfried Philips |
ACIVS | 3 |
| 2016 | Fuzzy rough sets for self-labelling: An exploratory analysisabstractSemi-supervised learning incorporates aspects of both supervised and unsupervised learning. In semi-supervised classification, only some data instances have associated class labels, while others are unlabelled. One particular group of semi-supervised classification approaches are those known as self-labelling techniques, which attempt to assign class labels to the unlabelled data instances. This is achieved by using the class predictions based upon the information of the labelled part of the data. In this paper, the applicability and suitability of fuzzy rough set theory for the task of self-labelling is investigated. An important preparatory experimental study is presented that evaluates how accurately different fuzzy rough set models can predict the classes of unlabelled data instances for semi-supervised classification. The predictions are made either by considering only the labelled data instances or by involving the unlabelled data instances as well. A stability analysis of the predictions also helps to provide further insight into the characteristics of the different fuzzy rough models. Our study shows that the ordered weighted average based fuzzy rough model performs best in terms of both accuracy and stability. Our conclusions offer a solid foundation and rationale that will allow the construction of a fuzzy rough self-labelling technique. They also provide an understanding of the applicability of fuzzy rough sets for the task of semi-supervised classification in general. Sarah Vluymans, Neil Mac Parthaláin, Chris Cornelis, Yvan Saeys |
FUZZ-IEEE | 4 |
| 2016 | Machine Learning Challenges for Single Cell Data
Sofie Van Gassen, Tom Dhaene, Yvan Saeys |
ECML/PKDD (3) | 3 |
| 2016 | Netter: re-ranking gene network inference predictions using structural network propertiesabstractBACKGROUND: Many algorithms have been developed to infer the topology of gene regulatory networks from gene expression data. These methods typically produce a ranking of links between genes with associated confidence scores, after which a certain threshold is chosen to produce the inferred topology. However, the structural properties of the predicted network do not resemble those typical for a gene regulatory network, as most algorithms only take into account connections found in the data and do not include known graph properties in their inference process. This lowers the prediction accuracy of these methods, limiting their usability in practice. RESULTS: We propose a post-processing algorithm which is applicable to any confidence ranking of regulatory interactions obtained from a network inference method which can use, inter alia, graphlets and several graph-invariant properties to re-rank the links into a more accurate prediction. To demonstrate the potential of our approach, we re-rank predictions of six different state-of-the-art algorithms using three simple network properties as optimization criteria and show that Netter can improve the predictions made on both artificially generated data as well as the DREAM4 and DREAM5 benchmarks. Additionally, the DREAM5 E.coli. community prediction inferred from real expression data is further improved. Furthermore, Netter compares favorably to other post-processing algorithms and is not restricted to correlation-like predictions. Lastly, we demonstrate that the performance increase is robust for a wide range of parameter settings. Netter is available at http://bioinformatics.intec.ugent.be. CONCLUSIONS: Network inference from high-throughput data is a long-standing challenge. In this work, we present Netter, which can further refine network predictions based on a set of user-defined graph properties. Netter is a flexible system which can be applied in unison with any method producing a ranking from omics data. It can be tailored to specific prior knowledge by expert users but can also be applied in general uses cases. Concluding, we believe that Netter is an interesting second step in the network inference process to further increase the quality of prediction. Joeri Ruyssinck, Piet Demeester, Tom Dhaene, Yvan Saeys |
BMC Bioinform. | 4 |
| 2016 | EPRENNID: An evolutionary prototype reduction based ensemble for nearest neighbor classification of imbalanced data
Sarah Vluymans, Isaac Triguero, Chris Cornelis, Yvan Saeys |
Neurocomputing | 4 |
| 2016 | Fuzzy rough classifiers for class imbalanced multi-instance data
Sarah Vluymans, Dánel Sánchez Tarragó, Yvan Saeys, Chris Cornelis, Francisco Herrera |
Pattern Recognit. | 3 |
| 2016 | Fuzzy Multi-Instance ClassifiersabstractMulti-instance learning is a setting in supervised learning where the data consist of bags of instances. Samples in the dataset are groups of individual instances. In classification problems, a decision value is assigned to the entire bag, and the classification of an unseen bag involves the prediction of the decision value based on the instances it contains. In this paper, we develop a framework for multi-instance classifiers based on fuzzy set theory. Fuzzy sets have been used in many machine learning applications, but so far not in the classification of multi-instance data. We explore its untapped potential here. We interpret the classes as fuzzy sets and determine membership degrees of unseen bags to these sets based on the available training data. In doing so, we develop a framework of classifiers that extract the required membership degrees either at the level of instances (instance-based) or at the level of bags (bag-based). We offer an extensive analysis of the different settings within the proposed framework. We experimentally compare our proposal to state-of-the-art multi-instance classifiers, and based on two evaluation measures, our methods are shown to perform very well. Sarah Vluymans, Dánel Sánchez Tarragó, Yvan Saeys, Chris Cornelis, Francisco Herrera |
IEEE Trans. Fuzzy Syst. | 3 |
| 2015 | Evolutionary undersampling for imbalanced big data classificationabstractClassification techniques in the big data scenario are in high demand in a wide variety of applications. The huge increment of available data may limit the applicability of most of the standard techniques. This problem becomes even more difficult when the class distribution is skewed, the topic known as imbalanced big data classification. Evolutionary undersampling techniques have shown to be a very promising solution to deal with the class imbalance problem. However, their practical application is limited to problems with no more than tens of thousands of instances. In this contribution we design a parallel model to enable evolutionary undersampling methods to deal with large-scale problems. To do this, we rely on a MapReduce scheme that distributes the functioning of these kinds of algorithms in a cluster of computing elements. Moreover, we develop a windowing approach for class imbalance data in order to speed up the undersampling process without losing accuracy. In our experiments we test the capabilities of the proposed scheme with several data sets with up to 4 million instances. The results show promising scalability abilities for evolutionary undersampling within the proposed framework. Isaac Triguero, Mikel Galar, Sarah Vluymans, Chris Cornelis, Humberto Bustince, Francisco Herrera, Yvan Saeys |
CEC | 7 |
| 2015 | Fuzzy Rough Set Prototype Selection for RegressionabstractInstance selection methods are a class of preprocessing techniques that have been widely studied in machine learning to remove redundant or noisy instances from a training set. The main focus of such prior efforts has been on the selection of suitable training instances to perform a classification task for crisp class labels. In this paper, we propose a novel instance selection technique termed Fuzzy Rough Set Prototype Selection for Regression (FRPS-R) for solving regression problems, where the outcome is continuous. We use concepts from fuzzy rough set theory and extend the currently well-known fuzzy rough set prototype selection technique to model the quality of all available elements and then use a wrapper approach to select an optimal subset of high-quality instances; thereby generalizing the idea. Our experimental evaluation shows that the application of our proposed instance selection technique can significantly improve the predictive performance of the weighted k-nearest neighbor regression algorithm, in particular when noise is present in the original training set. Sarah Vluymans, Yvan Saeys, Chris Cornelis, Ankur Teredesai, Martine De Cock |
FUZZ-IEEE | 2 |
| 2015 | Applications of Fuzzy Rough Set Theory in Machine Learning: a SurveyabstractData used in machine learning applications is prone to contain both vague and incomplete information. Many authors have proposed to use fuzzy rough set theory in the development of new techniques tackling these characteristics. Fuzzy sets deal with vague data, while rough sets allow to model incomp lete information. As such, the hybrid setting of the two paradigms is an ideal candidate tool to confront the separate challenges. In this paper, we present a thorough review on the use of fuzzy rough sets in machine learning applications. We recall their integration in preprocessing methods and consider learning algorithms in the supervised, unsupervised and semi-supervised domains and outline future challenges. Throughout the paper, we highlight the interaction between theoretical advances on fuzzy rough sets and practical machine learning tools that take advantage of them. Sarah Vluymans, Lynn D'eer, Yvan Saeys, Chris Cornelis |
Fundam. Informaticae | 3 |
| 2014 | Complex Aggregates over Clusters of Elements
Celine Vens, Sofie Van Gassen, Tom Dhaene, Yvan Saeys |
ILP | 4 |
| 2012 | Statistical interpretation of machine learning-based feature importance scores for biomarker discoveryabstractMOTIVATION: Univariate statistical tests are widely used for biomarker discovery in bioinformatics. These procedures are simple, fast and their output is easily interpretable by biologists but they can only identify variables that provide a significant amount of information in isolation from the other variables. As biological processes are expected to involve complex interactions between variables, univariate methods thus potentially miss some informative biomarkers. Variable relevance scores provided by machine learning techniques, however, are potentially able to highlight multivariate interacting effects, but unlike the p-values returned by univariate tests, these relevance scores are usually not statistically interpretable. This lack of interpretability hampers the determination of a relevance threshold for extracting a feature subset from the rankings and also prevents the wide adoption of these methods by practicians. RESULTS: We evaluated several, existing and novel, procedures that extract relevant features from rankings derived from machine learning approaches. These procedures replace the relevance scores with measures that can be interpreted in a statistical way, such as p-values, false discovery rates, or family wise error rates, for which it is easier to determine a significance level. Experiments were performed on several artificial problems as well as on real microarray datasets. Although the methods differ in terms of computing times and the tradeoff, they achieve in terms of false positives and false negatives, some of them greatly help in the extraction of truly relevant biomarkers and should thus be of great practical interest for biologists and physicians. As a side conclusion, our experiments also clearly highlight that using model performance as a criterion for feature selection is often counter-productive. AVAILABILITY AND IMPLEMENTATION: Python source codes of all tested methods, as well as the MATLAB scripts used for data simulation, can be found in the Supplementary Material. Vân Anh Huynh-Thu, Yvan Saeys, Louis Wehenkel, Pierre Geurts |
Bioinform. | 2 |
| 2011 | A greedy, graph-based algorithm for the alignment of multiple homologous gene listsabstractMOTIVATION: Many comparative genomics studies rely on the correct identification of homologous genomic regions using accurate alignment tools. In such case, the alphabet of the input sequences consists of complete genes, rather than nucleotides or amino acids. As optimal multiple sequence alignment is computationally impractical, a progressive alignment strategy is often employed. However, such an approach is susceptible to the propagation of alignment errors in early pairwise alignment steps, especially when dealing with strongly diverged genomic regions. In this article, we present a novel accurate and efficient greedy, graph-based algorithm for the alignment of multiple homologous genomic segments, represented as ordered gene lists. RESULTS: Based on provable properties of the graph structure, several heuristics are developed to resolve local alignment conflicts that occur due to gene duplication and/or rearrangement events on the different genomic segments. The performance of the algorithm is assessed by comparing the alignment results of homologous genomic segments in Arabidopsis thaliana to those obtained by using both a progressive alignment method and an earlier graph-based implementation. Especially for datasets that contain strongly diverged segments, the proposed method achieves a substantially higher alignment accuracy, and proves to be sufficiently fast for large datasets including a few dozens of eukaryotic genomes. AVAILABILITY: http://bioinformatics.psb.ugent.be/software. The algorithm is implemented as a part of the i-ADHoRe 3.0 package. Jan Fostier, Sebastian Proost, Bart Dhoedt, Yvan Saeys, Piet Demeester, Yves Van de Peer, Klaas Vandepoele |
Bioinform. | 4 |
| 2011 | High-Precision Bio-molecular Event Extraction from Text Using Parallel Binary ClassifiersabstractWe have developed a machine learning framework to accurately extract complex genetic interactions from text. Employing type-specific classifiers, this framework processes research articles to extract various biological events. Subsequently, the algorithm identifies regulation events that take other events as arguments, allowing a nested structure of predictions. All predictions are merged into an integrated network, useful for visualization and for deduction of new biological knowledge. In this paper, we discuss several design choices for an event-based extraction framework. These detailed studies help improving on existing systems, which is illustrated by the relative performance gain of 10% of our system compared to the official results in the recent BioNLP’09 Shared Task. Our framework now achieves state-of-the-art performance with 37.43 recall, 54.81 precision and 44.48 F-score. We further present the first study of feature selection for bio-molecular event extraction from text. While producing more cost-effective models, feature selection can also lead to a better insight into the complexity of the challenge. Finally, this paper tries to bridge the gap between theoretical relation extraction from text and experimental work on bio-molecular interactions by discussing interesting opportunities to employ event-based text mining tools for real-life tasks such as hypothesis generation, database curation and knowledge discovery. Sofie Van Landeghem, Bernard De Baets, Yves Van de Peer, Yvan Saeys |
Comput. Intell. | 4 |
| 2011 | Peakbin Selection in Mass Spectrometry Data Using a Consensus Approach with Estimation of Distribution AlgorithmsabstractProgress is continuously being made in the quest for stable biomarkers linked to complex diseases. Mass spectrometers are one of the devices for tackling this problem. The data profiles they produce are noisy and unstable. In these profiles, biomarkers are detected as signal regions (peaks), where control and disease samples behave differently. Mass spectrometry (MS) data generally contain a limited number of samples described by a high number of features. In this work, we present a novel class of evolutionary algorithms, estimation of distribution algorithms (EDA), as an efficient peak selector in this MS domain. There is a trade-of f between the reliability of the detected biomarkers and the low number of samples for analysis. For this reason, we introduce a consensus approach, built upon the classical EDA scheme, that improves stability and robustness of the final set of relevant peaks. An entire data workflow is designed to yield unbiased results. Four publicly available MS data sets (two MALDI-TOF and another two SELDI-TOF) are analyzed. The results are compared to the original works, and a new plot (peak frequential plot) for graphically inspecting the relevant peaks is introduced. A complete online supplementary page, which can be found at http://www.sc.ehu.es/ccwbayes/members/ruben/ms, includes extended info and results, in addition to Matlab scripts and references. Rubén Armañanzas, Yvan Saeys, Iñaki Inza, Miguel García-Torres, Concha Bielza, Yves Van de Peer, Pedro Larrañaga |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Towards an ASR-free objective analysis of pathological speechabstractNowadays, intelligibility is a popular measure of the severity of the articulatory deficiencies of a pathological speaker. Usually, this measure is obtained by means of a perceptual test, consisting of nonconventional and/or nonconnected words. In previous work, we developed a system incorporating two Automatic Speech Recognizers (ASR) that could fairly accurately estimate phoneme intelligibility (PI). In the present paper, we propose a novel method that aims to assess the running speech intelligibility (RSI) as a more relevant indicator of the communication efficiency of a speaker in a natural setting. The proposed method computes a phonological characterization of the speaker by means of a statistical analysis of frame-level phonological features. Important is that this analysis requires no knowledge of what the speaker was supposed to say. The new characterization is demonstrated to predict PI and to provide valuable information about the nature and severity of the pathology. Catherine Middag, Yvan Saeys, Jean-Pierre Martens |
INTERSPEECH | 2 |
| 2010 | Robust biomarker identification for cancer diagnosis with ensemble feature selection methodsabstractMOTIVATION: Biomarker discovery is an important topic in biomedical applications of computational biology, including applications such as gene and SNP selection from high-dimensional data. Surprisingly, the stability with respect to sampling variation or robustness of such selection processes has received attention only recently. However, robustness of biomarkers is an important issue, as it may greatly influence subsequent biological validations. In addition, a more robust set of markers may strengthen the confidence of an expert in the results of a selection method. RESULTS: Our first contribution is a general framework for the analysis of the robustness of a biomarker selection algorithm. Secondly, we conducted a large-scale analysis of the recently introduced concept of ensemble feature selection, where multiple feature selections are combined in order to increase the robustness of the final set of selected features. We focus on selection methods that are embedded in the estimation of support vector machines (SVMs). SVMs are powerful classification models that have shown state-of-the-art performance on several diagnosis and prognosis tasks on biological data. Their feature selection extensions also offered good results for gene selection tasks. We show that the robustness of SVMs for biomarker discovery can be substantially increased by using ensemble feature selection techniques, while at the same time improving upon classification performances. The proposed methodology is evaluated on four microarray datasets showing increases of up to almost 30% in robustness of the selected biomarkers, along with an improvement of approximately 15% in classification performance. The stability improvement with ensemble methods is particularly noticeable for small signature sizes (a few tens of genes), which is most relevant for the design of a diagnosis or prognosis model from a gene signature. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Abeel, Thibault Helleputte, Yves Van de Peer, Pierre Dupont, Yvan Saeys |
Bioinform. | 5 |
| 2010 | Discriminative and informative features for biomolecular text mining with ensemble feature selectionabstractMOTIVATION: In the field of biomolecular text mining, black box behavior of machine learning systems currently limits understanding of the true nature of the predictions. However, feature selection (FS) is capable of identifying the most relevant features in any supervised learning setting, providing insight into the specific properties of the classification algorithm. This allows us to build more accurate classifiers while at the same time bridging the gap between the black box behavior and the end-user who has to interpret the results. RESULTS: We show that our FS methodology successfully discards a large fraction of machine-generated features, improving classification performance of state-of-the-art text mining algorithms. Furthermore, we illustrate how FS can be applied to gain understanding in the predictions of a framework for biomolecular event extraction from text. We include numerous examples of highly discriminative features that model either biological reality or common linguistic constructs. Finally, we discuss a number of insights from our FS analyses that will provide the opportunity to considerably improve upon current text mining tools. AVAILABILITY: The FS algorithms and classifiers are available in Java-ML (http://java-ml.sf.net). The datasets are publicly available from the BioNLP'09 Shared Task web site (http://www-tsujii.is.s.u-tokyo.ac.jp/GENIA/SharedTask/). Sofie Van Landeghem, Thomas Abeel, Yvan Saeys, Yves Van de Peer |
Bioinform. | 3 |
| 2010 | Highlights of the BioTM 2010 workshop on advances in bio text miningabstractRecently, the application of text mining (TM) and natural language processing (NLP) techniques to the biological and medical sciences has received increasing interest. In addition to many new workshops and conferences arising in this domain, recently also a number of community-wide tasks were conducted to benchmark text mining techniques on specific challenges (e.g. BioCreative, BioNLP Shared Task, ...) Thomas Abeel, Sofie Van Landeghem, Roser Morante, Vincent Van Asch, Yves Van de Peer, Walter Daelemans, Yvan Saeys |
BMC Bioinform. | 7 |
| 2009 | Toward a gold standard for promoter prediction evaluationabstractMOTIVATION: Promoter prediction is an important task in genome annotation projects, and during the past years many new promoter prediction programs (PPPs) have emerged. However, many of these programs are compared inadequately to other programs. In most cases, only a small portion of the genome is used to evaluate the program, which is not a realistic setting for whole genome annotation projects. In addition, a common evaluation design to properly compare PPPs is still lacking. RESULTS: We present a large-scale benchmarking study of 17 state-of-the-art PPPs. A multi-faceted evaluation strategy is proposed that can be used as a gold standard for promoter prediction evaluation, allowing authors of promoter prediction software to compare their method to existing methods in a proper way. This evaluation strategy is subsequently used to compare the chosen promoter predictors, and an in-depth analysis on predictive performance, promoter class specificity, overlap between predictors and positional bias of the predictions is conducted. AVAILABILITY: We provide the implementations of the four protocols, as well as the datasets required to perform the benchmarks to the academic community free of charge on request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Abeel, Yves Van de Peer, Yvan Saeys |
Bioinform. | 3 |
| 2009 | Java-ML: A Machine Learning Library
Thomas Abeel, Yves Van de Peer, Yvan Saeys |
J. Mach. Learn. Res. | 3 |
| 2008 | ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profilesabstractMOTIVATION: More and more genomes are being sequenced, and to keep up with the pace of sequencing projects, automated annotation techniques are required. One of the most challenging problems in genome annotation is the identification of the core promoter. Because the identification of the transcription initiation region is such a challenging problem, it is not yet a common practice to integrate transcription start site prediction in genome annotation projects. Nevertheless, better core promoter prediction can improve genome annotation and can be used to guide experimental work. RESULTS: Comparing the average structural profile based on base stacking energy of transcribed, promoter and intergenic sequences demonstrates that the core promoter has unique features that cannot be found in other sequences. We show that unsupervised clustering by using self-organizing maps can clearly distinguish between the structural profiles of promoter sequences and other genomic sequences. An implementation of this promoter prediction program, called ProSOM, is available and has been compared with the state-of-the-art. We propose an objective, accurate and biologically sound validation scheme for core promoter predictors. ProSOM performs at least as well as the software currently available, but our technique is more balanced in terms of the number of predicted sites and the number of false predictions, resulting in a better all-round performance. Additional tests on the ENCODE regions of the human genome show that 98% of all predictions made by ProSOM can be associated with transcriptionally active regions, which demonstrates the high precision. AVAILABILITY: Predictions for the human genome, the validation datasets and the program (ProSOM) are available upon request. Thomas Abeel, Yvan Saeys, Pierre Rouzé, Yves Van de Peer |
ISMB | 2 |
| 2008 | Robust Feature Selection Using Ensemble Feature Selection Techniques
Yvan Saeys, Thomas Abeel, Yves Van de Peer |
ECML/PKDD (2) | 1 |
| 2008 | FunSiP: a modular and extensible classifier for the prediction of functional sites in DNAabstractMOTIVATION: Many problems in genome annotation are tackled by using a classification model to predict functional sites such as splice sites, translation start sites or stop codons. Locating the correct position of these sites remains one of the most important but also one of the most difficult issues in the structural annotation of genomes. Most of the software currently in use is written for a very specific problem, thereby limiting the possibilities for reuse. SUMMARY: We developed a software platform that uses a very general approach towards the classification of functional sites in DNA sequences. The program uses an ab initio approach towards the identification of these sites, and extends SpliceMachine, a previously developed splice site predictor that shows state-of-the-art performance for both donor and acceptor splice site recognition in the human and Arabidopsis thaliana genome. AVAILABILITY: The program is developed as a stand-alone Java application, and is available as GPLv3 open-source software. The program, source and documentation can be obtained from the 'Software' section at http://bioinformatics.psb.ugent.be/. SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online. Michiel Van Bel, Yvan Saeys, Yves Van de Peer |
Bioinform. | 2 |
| 2007 | A review of feature selection techniques in bioinformaticsabstractAbstract Feature selection techniques have become an apparent need in many bioinformatics applications. In addition to the large pool of techniques that have already been developed in the machine learning and data mining fields, specific applications in bioinformatics have led to a wealth of newly proposed techniques. In this article, we make the interested reader aware of the possibilities of feature selection, providing a basic taxonomy of feature selection techniques, and discussing their use, variety and potential in a number of both common as well as upcoming bioinformatics applications. Contact: [email protected] Supplementary information: http://bioinformatics.psb.ugent.be/supplementary_data/yvsae/fsreview Yvan Saeys, Iñaki Inza, Pedro Larrañaga |
Bioinform. | 1 |
| 2007 | In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protistsabstractMOTIVATION: Prediction of the coding potential for stretches of DNA is crucial in gene calling and genome annotation, where it is used to identify potential exons and to position their boundaries in conjunction with functional sites, such as splice sites and translation initiation sites. The ability to discriminate between coding and non-coding sequences relates to the structure of coding sequences, which are organized in codons, and by their biased usage. For statistical reasons, the longer the sequences, the easier it is to detect this codon bias. However, in many eukaryotic genomes, where genes harbour many introns, both introns and exons might be small and hard to distinguish based on coding potential. RESULTS: Here, we present novel approaches that specifically aim at a better detection of coding potential in short sequences. The methods use complementary sequence features, combined with identification of which features are relevant in discriminating between coding and non-coding sequences. These newly developed methods are evaluated on different species, representative of four major eukaryotic kingdoms, and extensively compared to state-of-the-art Markov models, which are often used for predicting coding potential. The main conclusions drawn from our analyses are that (1) combining complementary sequence features clearly outperforms current Markov models for coding potential prediction in short sequence fragments, (2) coding potential prediction benefits from length-specific models, and these models are not necessarily the same for different sequence lengths and (3) comparing the results across several species indicates that, although our combined method consistently performs extremely well, there are important differences across genomes. SUPPLEMENTARY DATA: http://bioinformatics.psb.ugent.be/. Yvan Saeys, Pierre Rouzé, Yves Van de Peer |
Bioinform. | 1 |
| 2007 | Validating module network learning algorithms using simulated dataabstractBACKGROUND: In recent years, several authors have used probabilistic graphical models to learn expression modules and their regulatory programs from gene expression data. Despite the demonstrated success of such algorithms in uncovering biologically relevant regulatory relations, further developments in the area are hampered by a lack of tools to compare the performance of alternative module network learning strategies. Here, we demonstrate the use of the synthetic data generator SynTReN for the purpose of testing and comparing module network learning algorithms. We introduce a software package for learning module networks, called LeMoNe, which incorporates a novel strategy for learning regulatory programs. Novelties include the use of a bottom-up Bayesian hierarchical clustering to construct the regulatory programs, and the use of a conditional entropy measure to assign regulators to the regulation program nodes. Using SynTReN data, we test the performance of LeMoNe in a completely controlled situation and assess the effect of the methodological changes we made with respect to an existing software package, namely Genomica. Additionally, we assess the effect of various parameters, such as the size of the data set and the amount of noise, on the inference performance. RESULTS: Overall, application of Genomica and LeMoNe to simulated data sets gave comparable results. However, LeMoNe offers some advantages, one of them being that the learning process is considerably faster for larger data sets. Additionally, we show that the location of the regulators in the LeMoNe regulation programs and their conditional entropy may be used to prioritize regulators for functional validation, and that the combination of the bottom-up clustering strategy with the conditional entropy-based assignment of regulators improves the handling of missing or hidden regulators. CONCLUSION: We show that data simulators such as SynTReN are very well suited for the purpose of developing, testing and improving module network algorithms. We used SynTReN data to develop and test an alternative module network learning strategy, which is incorporated in the software package LeMoNe, and we provide evidence that this alternative strategy has several advantages with respect to existing methods. Tom Michoel, Steven Maere, Eric Bonnet, Anagha Joshi, Yvan Saeys, Tim Van den Bulcke, Koenraad Van Leemput, Piet van Remortel, Martin Kuiper, Kathleen Marchal, Yves Van de Peer |
BMC Bioinform. | 5 |
| 2006 | Feature Extraction Using Clustering of Protein
Isis Bonet, Yvan Saeys, Ricardo del Corazón Grau-Ábalo, María Matilde García Lorenzo, Robersy Sanchez, Yves Van de Peer |
CIARP | 2 |
| 2005 | SpliceMachine: predicting splice sites from high-dimensional local context representationsabstractMOTIVATION: In this age of complete genome sequencing, finding the location and structure of genes is crucial for further molecular research. The accurate prediction of intron boundaries largely facilitates the correct prediction of gene structure in nuclear genomes. Many tools for localizing these boundaries on DNA sequences have been developed and are available to researchers through the internet. Nevertheless, these tools still make many false positive predictions. RESULTS: This manuscript presents a novel publicly available splice site prediction tool named SpliceMachine that (i) shows state-of-the-art prediction performance on Arabidopsis thaliana and human sequences, (ii) performs a computationally fast annotation and (iii) can be trained by the user on its own data. AVAILABILITY: Results, figures and software are available at http://www.bioinformatics.psb.ugent.be/supplementary_data/ CONTACT: [email protected]; [email protected]. Sven Degroeve, Yvan Saeys, Bernard De Baets, Pierre Rouzé, Yves Van de Peer |
Bioinform. | 2 |
| 2004 | Digging into Acceptor Splice Site Prediction: An Iterative Feature Selection Approach
Yvan Saeys, Sven Degroeve, Yves Van de Peer |
PKDD | 1 |
| 2004 | Feature selection for splice site prediction: A new method using EDA-based feature rankingabstractBACKGROUND: The identification of relevant biological features in large and complex datasets is an important step towards gaining insight in the processes underlying the data. Other advantages of feature selection include the ability of the classification system to attain good or even better solutions using a restricted subset of features, and a faster classification. Thus, robust methods for fast feature selection are of key importance in extracting knowledge from complex biological data. RESULTS: In this paper we present a novel method for feature subset selection applied to splice site prediction, based on estimation of distribution algorithms, a more general framework of genetic algorithms. From the estimated distribution of the algorithm, a feature ranking is derived. Afterwards this ranking is used to iteratively discard features. We apply this technique to the problem of splice site prediction, and show how it can be used to gain insight into the underlying biological process of splicing. CONCLUSION: We show that this technique proves to be more robust than the traditional use of estimation of distribution algorithms for feature selection: instead of returning a single best subset of features (as they normally do) this method provides a dynamical view of the feature selection process, like the traditional sequential wrapper methods. However, the method is faster than the traditional techniques, and scales better to datasets described by a large number of features. Yvan Saeys, Sven Degroeve, Dirk Aeyels, Pierre Rouzé, Yves Van de Peer |
BMC Bioinform. | 1 |