EDBT 2026 Demo / reviewers in the wild / expert
Yves Moreau
dblp:52/5372
· DBLP profile ↗
76ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0002-4647-6560ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 52 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | JINet: easy and secure private data analysis for everyoneabstractBACKGROUND: The barriers to effective data analysis are sometimes insurmountable. Concerns ranging from privacy, security, and complexity can prevent researchers from using existing data analysis tools. RESULTS: JINet is a web browser-based platform intended to democratise access to advanced clinical and genomic data analysis software. It hosts numerous data analysis applications that are run in the safety of each User's web browser, without the data ever leaving their machine. CONCLUSIONS: JINet promotes collaboration, standardisation and reproducibility by sharing scripts rather than data and creating a self-sustaining community around it in which Users and data analysis tools Developers interact thanks to JINet's interoperability primitives. Giada Lalli, James Collier, Yves Moreau, Daniele Raimondi |
BMC Bioinform. | 3 |
| 2024 | Atom-Level Optical Chemical Structure Recognition with Limited SupervisionabstractIdentifying the chemical structure from a graphical representation, or image, of a molecule is a challenging pattern recognition task that would greatly benefit drug development. Yet, existing methods for chemical structure recognition do not typically generalize well, and show diminished effectiveness when confronted with domains where data is sparse, or costly to generate, such as hand-drawn molecule images. To address this limitation, we propose a new chemical structure recognition tool that delivers state-of-the-art performance and can adapt to new domains with a limited number of data samples and supervision. Unlike previous approaches, our method provides atom-level localization, and can therefore segment the image into the different atoms and bonds. Our model is the first model to perform OCSR with atom-level entity detection with only SMILES supervision. Through rigorous and extensive benchmarking, we demonstrate the preeminence of our chemical structure recognition approach in terms of data efficiency, accuracy, and atom-level entity prediction. Martijn Oldenhof, Edward De Brouwer, Adam Arany, Yves Moreau |
CVPR | 4 |
| 2024 | Model Based Clustering of Time Series Utilizing Expert ODEs
András Formanek, Edward De Brouwer, Péter Antal, Yves Moreau, Adam Arany |
ICANN (4) | 4 |
| 2024 | MetDecode: methylation-based deconvolution of cell-free DNA for noninvasive multi-cancer typingabstractMOTIVATION: Circulating-cell free DNA (cfDNA) is widely explored as a noninvasive biomarker for cancer screening and diagnosis. The ability to decode the cells of origin in cfDNA would provide biological insights into pathophysiological mechanisms, aiding in cancer characterization and directing clinical management and follow-up. RESULTS: We developed a DNA methylation signature-based deconvolution algorithm, MetDecode, for cancer tissue origin identification. We built a reference atlas exploiting de novo and published whole-genome methylation sequencing data for colorectal, breast, ovarian, and cervical cancer, and blood-cell-derived entities. MetDecode models the contributors absent in the atlas with methylation patterns learnt on-the-fly from the input cfDNA methylation profiles. In addition, our model accounts for the coverage of each marker region to alleviate potential sources of noise. In-silico experiments showed a limit of detection down to 2.88% of tumor tissue contribution in cfDNA. MetDecode produced Pearson correlation coefficients above 0.95 and outperformed other methods in simulations (P < 0.001; T-test; one-sided). In plasma cfDNA profiles from cancer patients, MetDecode assigned the correct tissue-of-origin in 84.2% of cases. In conclusion, MetDecode can unravel alterations in the cfDNA pool components by accurately estimating the contribution of multiple tissues, while supplied with an imperfect reference atlas. AVAILABILITY AND IMPLEMENTATION: MetDecode is available at https://github.com/JorisVermeeschLab/MetDecode. Antoine Passemiers, Stefania Tuveri, Dhanya Sudhakaran, Tatjana Jatsenko, Tina Laga, Kevin Punie, Sigrid Hatse, Sabine Tejpar, An Coosemans, Els Van Nieuwenhuysen, Dirk Timmerman, Giuseppe Floris, Anne-Sophie Van Rompuy, Xavier Sagaert, Antonia C. Testa, Daniela Ficherova, Daniele Raimondi, Frederic Amant, Liesbeth Lenaerts, Yves Moreau, Joris Robert Vermeesch |
Bioinform. | 20 |
| 2023 | Industry-Scale Orchestrated Federated Learning for Drug DiscoveryabstractTo apply federated learning to drug discovery we developed a novel platform in the context of European Innovative Medicines Initiative (IMI) project MELLODDY (grant n°831472), which was comprised of 10 pharmaceutical companies, academic research labs, large industrial companies and startups. The MELLODDY platform was the first industry-scale platform to enable the creation of a global federated model for drug discovery without sharing the confidential data sets of the individual partners. The federated model was trained on the platform by aggregating the gradients of all contributing partners in a cryptographic, secure way following each training iteration. The platform was deployed on an Amazon Web Services (AWS) multi-account architecture running Kubernetes clusters in private subnets. Organisationally, the roles of the different partners were codified as different rights and permissions on the platform and administrated in a decentralized way. The MELLODDY platform generated new scientific discoveries which are described in a companion paper. Martijn Oldenhof, Gergely Ács, Balazs Pejo, Ansgar Schuffenhauer, Nicholas Holway, Noé Sturm, Arne Dieckmann, Oliver Fortmeier, Eric Boniface, Clément Mayer, Arnaud Gohier, Peter Schmidtke, Ritsuya Niwayama, Dieter Kopecky, Lewis H. Mervin, Prakash Chandra Rathi, Lukas Friedrich, András Formanek, Péter Antal, Jordon Rahaman, Adam Zalewski, Wouter Heyndrickx, Ezron Oluoch, Manuel Stößel, Michal Vanco, David Endico, Fabien Gelus, Thaïs de Boisfossé, Adrien Darbier, Ashley Nicollet, Matthieu Blottière, Maria Telenczuk, Van Tien Nguyen, Thibaud Martinez, Camille Boillet, Kelvin Moutet, Alexandre Picosson, Aurélien Gasser, Inal Djafar, Antoine Simon, Adam Arany, Jaak Simm, Yves Moreau, Ola Engkvist, Hugo Ceulemans, Camille Marini, Mathieu Galtier |
AAAI | 43 |
| 2023 | Weakly Supervised Knowledge Transfer with Probabilistic Logical Reasoning for Object Detection
Martijn Oldenhof, Adam Arany, Yves Moreau, Edward De Brouwer |
ICLR | 3 |
| 2023 | Nonlinear data fusion over Entity-Relation graphs for Drug-Target Interaction predictionabstractMOTIVATION: The prediction of reliable Drug-Target Interactions (DTIs) is a key task in computer-aided drug design and repurposing. Here, we present a new approach based on data fusion for DTI prediction built on top of the NXTfusion library, which generalizes the Matrix Factorization paradigm by extending it to the nonlinear inference over Entity-Relation graphs. RESULTS: We benchmarked our approach on five datasets and we compared our models against state-of-the-art methods. Our models outperform most of the existing methods and, simultaneously, retain the flexibility to predict both DTIs as binary classification and regression of the real-valued drug-target affinity, competing with models built explicitly for each task. Moreover, our findings suggest that the validation of DTI methods should be stricter than what has been proposed in some previous studies, focusing more on mimicking real-life DTI settings where predictions for previously unseen drugs, proteins, and drug-protein pairs are needed. These settings are exactly the context in which the benefit of integrating heterogeneous information with our Entity-Relation data fusion approach is the most evident. AVAILABILITY AND IMPLEMENTATION: All software and data are available at https://github.com/eugeniomazzone/CPI-NXTFusion and https://pypi.org/project/NXTfusion/. Eugenio Mazzone, Yves Moreau, Piero Fariselli, Daniele Raimondi |
Bioinform. | 2 |
| 2022 | Topological Graph Neural Networks
Max Horn, Edward De Brouwer, Michael Moor, Yves Moreau, Bastian Rieck, Karsten M. Borgwardt |
ICLR | 4 |
| 2022 | Fast and accurate inference of gene regulatory networks through robust precision matrix estimationabstractMOTIVATION: Transcriptional regulation mechanisms allow cells to adapt and respond to external stimuli by altering gene expression. The possible cell transcriptional states are determined by the underlying gene regulatory network (GRN), and reliably inferring such network would be invaluable to understand biological processes and disease progression. RESULTS: In this article, we present a novel method for the inference of GRNs, called PORTIA, which is based on robust precision matrix estimation, and we show that it positively compares with state-of-the-art methods while being orders of magnitude faster. We extensively validated PORTIA using the DREAM and MERLIN+P datasets as benchmarks. In addition, we propose a novel scoring metric that builds on graph-theoretical concepts. AVAILABILITY AND IMPLEMENTATION: The code and instructions for data acquisition and full reproduction of our results are available at https://github.com/AntoinePassemiers/PORTIA-Manuscript. PORTIA is available on PyPI as a Python package (portia-grn). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antoine Passemiers, Yves Moreau, Daniele Raimondi |
Bioinform. | 2 |
| 2021 | Latent Convergent Cross Mapping
Edward De Brouwer, Adam Arany, Jaak Simm, Yves Moreau |
ICLR | 4 |
| 2021 | In silico prediction of in vitro protein liquid-liquid phase separation experiments outcomes with multi-head neural attentionabstractMOTIVATION: Proteins able to undergo liquid-liquid phase separation (LLPS) in vivo and in vitro are drawing a lot of interest, due to their functional relevance for cell life. Nevertheless, the proteome-scale experimental screening of these proteins seems unfeasible, because besides being expensive and time-consuming, LLPS is heavily influenced by multiple environmental conditions such as concentration, pH and temperature, thus requiring a combinatorial number of experiments for each protein. RESULTS: To overcome this problem, we propose a neural network model able to predict the LLPS behavior of proteins given specified experimental conditions, effectively predicting the outcome of in vitro experiments. Our model can be used to rapidly screen proteins and experimental conditions searching for LLPS, thus reducing the search space that needs to be covered experimentally. We experimentally validate Droppler's prediction on the TAR DNA-binding protein in different experimental conditions, showing the consistency of its predictions. AVAILABILITY AND IMPLEMENTATION: A python implementation of Droppler is available at https://bitbucket.org/grogdrinker/droppler. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Emiel Michiels, Donya Pakravan, Anna Bratek-Skicki, Ludo Van Den Bosch, Yves Moreau, Frederic Rousseau 0001, Joost Schymkowitz |
Bioinform. | 7 |
| 2021 | A novel method for data fusion over entity-relation graphs and its application to protein-protein interaction predictionabstractMOTIVATION: Modern bioinformatics is facing increasingly complex problems to solve, and we are indeed rapidly approaching an era in which the ability to seamlessly integrate heterogeneous sources of information will be crucial for the scientific progress. Here, we present a novel non-linear data fusion framework that generalizes the conventional matrix factorization paradigm allowing inference over arbitrary entity-relation graphs, and we applied it to the prediction of protein-protein interactions (PPIs). Improving our knowledge of PPI networks at the proteome scale is indeed crucial to understand protein function, physiological and disease states and cell life in general. RESULTS: We devised three data fusion-based models for the proteome-level prediction of PPIs, and we show that our method outperforms state of the art approaches on common benchmarks. Moreover, we investigate its predictions on newly published PPIs, showing that this new data has a clear shift in its underlying distributions and we thus train and test our models on this extended dataset. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Jaak Simm, Adam Arany, Yves Moreau |
Bioinform. | 4 |
| 2020 | Insight into the protein solubility driving forces with neural attentionabstractProtein solubility is a key aspect for many biotechnological, biomedical and industrial processes, such as the production of active proteins and antibodies. In addition, understanding the molecular determinants of the solubility of proteins may be crucial to shed light on the molecular mechanisms of diseases caused by aggregation processes such as amyloidosis. Here we present SKADE, a novel Neural Network protein solubility predictor and we show how it can provide novel insight into the protein solubility mechanisms, thanks to its neural attention architecture. First, we show that SKADE positively compares with state of the art tools while using just the protein sequence as input. Then, thanks to the neural attention mechanism, we use SKADE to investigate the patterns learned during training and we analyse its decision process. We use this peculiarity to show that, while the attention profiles do not correlate with obvious sequence aspects such as biophysical properties of the aminoacids, they suggest that N- and C-termini are the most relevant regions for solubility prediction and are predictive for complex emergent properties such as aggregation-prone regions involved in beta-amyloidosis and contact density. Moreover, SKADE is able to identify mutations that increase or decrease the overall solubility of the protein, allowing it to be used to perform large scale in-silico mutagenesis of proteins in order to maximize their solubility. Daniele Raimondi, Gabriele Orlando, Piero Fariselli, Yves Moreau |
PLoS Comput. Biol. | 4 |
| 2019 | GRU-ODE-Bayes: Continuous Modeling of Sporadically-Observed Time SeriesabstractModeling real-world multidimensional time series can be particularly challenging when these are sporadically observed (i.e., sampling is irregular both in time and across dimensions)—such as in the case of clinical patient data. To address these challenges, we propose (1) a continuous-time version of the Gated Recurrent Unit, building upon the recent Neural Ordinary Differential Equations (Chen et al., 2018), and (2) a Bayesian update network that processes the sporadic observations. We bring these two ideas together in our GRU-ODE-Bayes method. We then demonstrate that the proposed method encodes a continuity prior for the latent process and that it can exactly represent the Fokker-Planck dynamics of complex processes driven by a multidimensional stochastic differential equation. Additionally, empirical evaluation shows that our method outperforms the state of the art on both synthetic data and real-world data with applications in healthcare and climate forecast. What is more, the continuity prior is shown to be well suited for low number of samples settings. Edward De Brouwer, Jaak Simm, Adam Arany, Yves Moreau |
NeurIPS | 4 |
| 2019 | GRNBoost2 and Arboreto: efficient and scalable inference of gene regulatory networksabstractSUMMARY: Inferring a Gene Regulatory Network (GRN) from gene expression data is a computationally expensive task, exacerbated by increasing data sizes due to advances in high-throughput gene profiling technology, such as single-cell RNA-seq. To equip researchers with a toolset to infer GRNs from large expression datasets, we propose GRNBoost2 and the Arboreto framework. GRNBoost2 is an efficient algorithm for regulatory network inference using gradient boosting, based on the GENIE3 architecture. Arboreto is a computational framework that scales up GRN inference algorithms complying with this architecture. Arboreto includes both GRNBoost2 and an improved implementation of GENIE3, as a user-friendly open source Python package. AVAILABILITY AND IMPLEMENTATION: Arboreto is available under the 3-Clause BSD license at http://arboreto.readthedocs.io. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Moerman 0002, Sara Aibar Santos, Carmen Bravo González-Blas, Jaak Simm, Yves Moreau, Jan Aerts, Stein Aerts |
Bioinform. | 5 |
| 2019 | Computational identification of prion-like RNA-binding proteins that form liquid phase-separated condensatesabstractMOTIVATION: Eukaryotic cells contain different membrane-delimited compartments, which are crucial for the biochemical reactions necessary to sustain cell life. Recent studies showed that cells can also trigger the formation of membraneless organelles composed by phase-separated proteins to respond to various stimuli. These condensates provide new ways to control the reactions and phase-separation proteins (PSPs) are thus revolutionizing how cellular organization is conceived. The small number of experimentally validated proteins, and the difficulty in discovering them, remain bottlenecks in PSPs research. RESULTS: Here we present PSPer, the first in-silico screening tool for prion-like RNA-binding PSPs. We show that it can prioritize PSPs among proteins containing similar RNA-binding domains, intrinsically disordered regions and prions. PSPer is thus suitable to screen proteomes, identifying the most likely PSPs for further experimental investigation. Moreover, its predictions are fully interpretable in the sense that it assigns specific functional regions to the predicted proteins, providing valuable information for experimental investigation of targeted mutations on these regions. Finally, we show that it can estimate the ability of artificially designed proteins to form condensates (r=-0.87), thus providing an in-silico screening tool for protein design experiments. AVAILABILITY AND IMPLEMENTATION: PSPer is available at bio2byte.com/psp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriele Orlando, Daniele Raimondi, Francesco Tabaro, Francesco Codicè, Yves Moreau, Wim F. Vranken |
Bioinform. | 5 |
| 2019 | Fast semi-supervised discriminant analysis for binary classification of large data sets
Joris Tavernier, Jaak Simm, Karl Meerbergen, Jörg K. Wegner, Hugo Ceulemans, Yves Moreau |
Pattern Recognit. | 6 |
| 2018 | ECCB 2018: The 17th European Conference on Computational BiologyabstractThis volume of Bioinformatics includes the proceedings papers of the 17th European Conference in Computational Biology (ECCB), an annual international Conference for research in computational biology and bioinformatics. The Conference is being held jointly with ISMB in the odd-numbered years and independently in the even-numbered years. This year, the 17th ECCB (ECCB 2018) will take place in Athens, at the Stavros Niarchos Foundation Cultural Center (SNFCC), from September 8 to 12, 2018. SNFCC, which was opened in June 2017, is a multifunctional arts, education and entertainment complex located on the edge of Faliro Bay, 4.5 km south of the center of Athens. Information on the ECCB 2018 can be found at eccb18.org and will be later archived at www.ebi.ac.uk/eccb/2018/. With more than 1,000 participants from academia and industry, ECCB is the leading European Conference on computational biology and bioinformatics and the second largest internationally, next to ISMB (Intelligent Systems in Molecular Biology). The work presented in ECCB is related to all domains of the field of computational biology and bioinformatics, ranging from molecular level to systems biology level. The proceedings papers present new computational methodologies and tools for addressing challenging problems in the field, including research in single cell high-throughput data, microbiome data, 3 D organization of the genome, and others. The Conference also includes Highlights presentations, which showcase important papers published over the last year. Since 2016, it also features an Applications track that presents computational biology work applied in industry, clinics, governmental organizations, and other fields beyond academia. ECCB meetings are held every year in a different European country. In odd-numbered years, ECCB is co-organized with the ISMB Conference, which is held in Europe. Previous meetings were organized in Prague, Czech Republic (Beerenwinkel and Bromberg, 2017, ISMB/ECCB 2017); The Hague, Netherlands (Heringa and Reinders, ECCB 2016); Dublin, Ireland (Moreau and Beerenwinkel, 2015, ISMB/ECCB 2015); Strasbourg, France (Devignes and Moreau, 2014, ECCB 2014); Berlin, Germany (Ben-Tal, 2013, ISMB/ECCB 2013); Basel, Switzerland (Schwede and Iber, 2012, ECCB 2012); Vienna, Austria (Gaasterland and Vingron, 2011, ISMB/ECCB 2011); Ghent, Belgium (Moreau and Heringa, 2010, ECCB 2010); Stockholm, Sweden (Gusfield and Tramontano, 2009, ISMB/ECCB 2009); Cagliari, Italy (Tramontano, 2008, ECCB 2008); Vienna, Austria (Lengauer et al., 2009, ISMB/ECCB 2007); Eilat, Israel (Wolfson and Safer, 2006, ECCB 2006); Madrid, Spain (Guigo et al., 2005, ECCB 2005); Glasgow, United Kingdom (Thornton et al., 2004, ISMB/ECCB 2004); Paris, France (Lenhof and Sagot, 2003, ECCB 2003); and Saarbrücken, Germany (Lengauer, 2002). ECCB 2018 continues the 16-year tradition and is held under the auspices of the Hellenic Society for Computational Biology and Bioinformatics (HSCBB; www.hscbb.gr), which promotes bioinformatics research and training in Greece since 2009. HSCBB organizes an annual conference in a different Greek city each time (e.g., Athens, Thessaloniki, Alexandroupolis, Patras, Lamia, Heraklion), aiming to provide exposure to new field developments to graduate students and researchers. The broad participation in the annual HSCBB conferences (more than 130 participants each year) show that the bioinformatics community in Greek universities and research centers has now matured significantly. HSCBB has also a strong international presence in the field of computational biology. It is an observer in the Greek ELIXIR consortium and is affiliated with the International Society for Computational Biology (ISCB) and Global Organization for Bioinformatics Learning, Education & Training (GOBLET). The organization of ECCB 2018, along with the European Student Council Symposium 2018 (see below), inspired the formation this year of the Greek chapter of the ISCB Student Council. ECCB 2018 received a record high number of 280 applications for proceedings talks from institutions from 48 countries. The submissions were organized in five themes, according to their topic: (1) Data (organization, integration, knowledge discovery, multi-scale modeling), (2) Genes (expression, function, editing, geno/phenotype), (3) Genome (sequence analysis, evolution, phylogeny, microbiome), (4) Proteins and Structural Biology (structure, function, alterations, drug design), and (5) Systems (molecular pathways, signaling, metabolomics). All proceedings submissions were subjected to a rigorous peer-review process (2-4 reviews per paper), organized by the members of the Programme Committees of the corresponding theme. The Programme Committees consisted of a Theme Chair (Michael Krauthammer, Roderic Guigo, Martin Vingron, Ivet Bahar, Alfonso Valencia, respectively) and two to four co-chairs (depending on the number of submissions assigned to this Theme). For the review, the EasyChair system (www.easychair.org) was used by 398 reviewers, recruited by the Programme Committees. Reviewers evaluated the impact and reproducibility of the presented research, as well as its suitability for the ECCB audience. Once the review was completed, the committee chairs and co-chairs selected 48 papers to be included in the ECCB 2018 proceedings (acceptance rate: 17%), with the number of papers accepted in each being proportional to the number of submissions initially assigned to this Theme. These papers required minor revisions and the authors had two weeks to modify them accordingly. All Proceedings Track papers and any supplementary files accompanying them are freely available in the electronic form of Oxford University Press journal Bioinformatics, as a special issue of September 2018. Together with the Proceedings track ECCB 2018 features 24 Highlights talks about scientific work already published in high impact science journals. They are coordinately presented and managed across the five themes with the proceeding presentations. We continued with newly established Application and ELIXIR tracks with 15 Applications talks, and 12 ELIXIR talks, all of which were selected after review. The aim of the Application track is to give a voice to those who apply computational biology in industry, clinics, governmental organizations, and other fields beyond academia. The aim of the ELIXIR Track is to showcase to the community the latest outputs and services from across the initiative. The presentations focus on developments relating to services and infrastructure within ELIXIR. ECCB 2018 also features a Poster Track, with 617 accepted posters, in which researchers present their latest findings in the corresponding thematic areas (Data, Genes, Genome, Proteins and Structural Biology, Systems). Seven distinguished keynote speakers will present their work: Prof. Bonnie Berger from the Massachusetts Institute of Technology (MIT), Prof. Christos Davatzikos from University of Pennsylvania, Prof. John Ioannidis from Stanford University, Prof. Manolis Kellis from MIT, Prof. Jose Onuchic from Rice University, Dr. Janet Thornton, Director Emeritus of the European Bioinformatics Institute (EMBL-EBI), and Dr. Eleftheria Zeggini from the German Research Center for Environmental Health (GmbH). Furthermore, ECCB 2018 is hosting for the first time three special invited talks before the lunch breaks and a joint keynote talk. The invited speakers will present the funding opportunities of the European Research Council (ERC) funding mechanisms (Dr. Maria Siomos) and a talk on the ISCB Student Council Internship Program (Farzana Rahman). The joint keynote will be about ELIXIR: A Common European Infrastructure for Bioinformatics Research presented by the ELIXIR director Dr. Niklas Blomberg. For the first time in ECCB 2018, ELIXIR organizes eight short workshops, namely ELIXIR TeSS Usability Study, Beacons, BioSchemas, Galaxy, OpenEbench Fnd Bio.tools, Research under the GDPR, Implementation of DMPs and Data Stewardship in practice, Secure access to your services using ELIXIR AAI. The 5th European Student Council Symposium (ESCS) is taking place before the main conference organized by the Student Council of the ISCB (chair Daniele Parisi from KU Leuven and co-chair Yvonne Saara Gladbach from Rostock University). ESCS highlights will be published in F1000Research via the ISCB Student Council channel. For the first time the ISCB Student Council will present also an invited talk as stated above. We are really thankful to the ISCB Student Council for their enthusiasm, hard work and genuine scientific interest, which has made the ESCS an inseparable part of ECCB. During the ECCB 2018, ten exhibitor booths of all sizes will be set up on the Exhibit Hall floor, next to the central Conference venue and close to the Poster session and lunch area, namely: ELIXIR, BioExcel, EMBL-EBI, ERC, ISCB and ISCB Student Council, Goblet, Oxford University Press, Cambridge University Press, Peer Community In. The booths will be grouped together in a way that provides an enhanced networking environment for our exhibitors and delegates. Exhibitors will showcase the latest trends in computational biology and bioinformatics technology, in scientific literature, as well as in modelling and simulation. We invite participants to visit the exhibition area and support these community-minded organizations, who deliver a strong message in supporting the Computational Biology and Bioinformatics scientific field. The weekend before the conference 8 workshops (including two from Special Interest Groups – SIGs) and 13 Tutorials will take place. The aim of workshops is to provide participants the opportunity to discuss different perspectives on the cutting edge of a selected research field, through presenting technical issues, exchanging research ideas and sharing practical experiences. The two SIGs and six workshops preceding the ECCB 2018 main conference were selected out of a total of 12 applications, running all but one for a single day: (SIG-1) BioExcel 2nd SIG Meeting: Advanced Simulations for Biomolecular Research organized by Rossen Apostolov (KTH Royal Institute of Technology in Stockholm, Sweden), Zoe Cournia (Biomedical Research Foundation of the Academy of Athens, Greece), Vera Matser, (EMBL-EBI, UK), Anastas Mishev (UKIM Saints Cyril and Methodius University of Skopje, Republic of Macedonia), Hrachya Astsatryan (Institute for Informatics and Automation Problems, Armenia), Adam Carter (EPCC, UK) (SIG-2): Intrinsically disordered proteins: advances and state-of-the-art of the field organized by Silvio Tosatto (University of Padua, Italy), Zsuzsanna Dosztanyi (Eötvös Loránd University, Hungary), Norman Davey (University College Dublin, Ireland), Damiano Piovesan (University of Padua, Italy) (W1) Computational Epigenomics a two-day workshop organized by Yassen Assenov (German Cancer Research Center, Germany), Guido Sanguinetti (University of Edinburgh, UK), Jordana Bell (King’s College London, UK), Jörn Walter (Saarland University, Germany), Christoph Bock (CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria) and Verena Wolf (Saarland University, Germany) (W2) BioNetVisA 2018 workshop: From biological network reconstruction to data visualization and analysis in molecular biology and medicine organized by Inna Kuperstein (Institut Curie, France), Emmanuel Barillot (Institut Curie, France), Andrei Zinovyev (Institut Curie, France), Luis Cristobal Monraz Gomez (Institut Curie, France), Hioraki Kitano (RIKEN Center for Integrative Medical Sciences, Japan), Minoru Kanehisa (Institute for Chemical Research, Kyoto University, Japan), Samik Ghosh (Systems Biology Institute, Tokyo, Japan), Nicolas Le Novère (Babraham Institute, UK), Robin Haw (Ontario Institute for Cancer Research, Canada), Alfonso Valencia (Spanish National Bioinformatics Institute, Madrid, Spain), Lodewyk Wessels (Netherlands Cancer Institute, Amsterdam, Netherlands), Patrick Kemmeren (Princess Maxima Center for Pediatric Oncology, Utrecht, Netherlands) (W3) CPW2018: Computational Pathology Workshop—second edition organized by Yves Sucaet (Vrije Universiteit Brussel, Belgium), Jeroen Van der Laak (UMC Radboud, Netherlands), Zev Leifer (New York College of Podiatric Medicine, USA), Yukako Yagi (Memorial Sloan Kettering Cancer Center, USA), Raphaël Marée (Université de Liège, Belgium), David Ameisen (IRIF, CNRS and Université Paris Diderot, France), Paul Van Diest (UMC Utrecht, Netherlands), Jeffrey Fine (Magee-Womens Hospital of UPMC) (W4) Interactive visualizations to guide health diagnostics and personalized medicine organized by Dimitrios Tzovaras, Konstantinos Votis and Kostas Stamatopoulos (all at Information Technologies Institute of the Center for Research and Technology Hellas, Greece) (W5) Logical modelling of cellular networks organized by Anna Niarakis (Univ Evry, Université Paris-Saclay, France) and Denis Thieffry (Ecole Normale Supérieure, Paris, France). (W6) Recent Computational Advances in Metagenomics organized by Valentin Loux, Mahendra Mariadassou, Pierre Peterlongo and Sophie Schbath (all at INRA, France) The purpose of the ECCB tutorial program is to provide participants with lectures and hands-on training to the most important and emerging topics in bioinformatics and computational biology research. The tutorials offer skills ranging from early and basic steps of computational analysis in recently introduced topics to advanced computational skills in important established topics. A total of thirteen ECCB 2018 tutorials running for half or a full day will be held at ECCB 2018. These were selected out of sixteen applications: (T1) Automated Machine Learning for Bioinformatics and Computational Biology organized by Ioannis Tsamardinos (University of Crete, Greece; Gnosis Data Analysis, Greece), Kleio Maria Verrou (University of Crete, Greece) and Vincenzo Lagani (Ilia State University, Georgia; Gnosis Data Analysis, Greece) (T2) Computational Mass Spectrometry with OpenMS—From Algorithms to Integrated Workflows organized by Julianus Pfeuffer (Freie Universität Berlin, Germany) and Timo Sachsenberg (Universität Tübingen, Germany) (T3)—Write your own R script to jointly analyze RNA-, ATAC-seq, DNA methylation, SNPs and find drug targets in signal transduction networks using TRANSFAC® organized by Philip Stegmaier (geneXplain GmbH, Germany), Olga Kel-Margoulis (geneXplain GmbH, Germany), Alexander Kel (geneXplain GmbH, Germany; Institute of Systems Biology, Russia) (T4) Deep learning for predicting protein–DNA and –RNA binding organized by Yaron Orenstein (Ben-Gurion University, Israel) (T5) Delving into non-coding RNA with RNAcentral and Rfam. Organized by Anton Petrov and Ioanna Kalvari (both at EMBL-EBI, UK) (T6) DIANA Tools and Databases: In silico investigation of miRNA functions organized by Artemis Hatzigeorgiou, Spyros Tastsoglou, Dimitra Karagkouni, Nikos Perdikopanis, Giorgos Skoufos, Ioannis Kavakiotis (all at University of Thessaly, Greece) (T7) Exploring programmatic access to Protein sequence, function and structure with UniProt and PDBe organized by Andrew Nightingale and Mihaly Varadi (both at EMBL-EBI, UK) (T8) Fifth International Hands-on Tutorial on Logical Modelling: Exploring the dynamics of biological systems organized by Tomas Helikar (University of Nebraska, USA) and Juilee Thakar (University of Rochester Medical Center, USA) (T9) From an application to a fully integrated workflow—Comprehensive software engineering with SeqAn organized by René Rahn and Hannes Hauswedell (both at Freie Universität Berlin, Germany) (T10) Hands-on on Protein Function Prediction with Machine Learning and Interactive Analytics organized by Rabie Saidi and Tunca Dogan (EMBL-EBI, UK) (T11) Single-cell RNA-Seq Data Analysis organized by Panagiotis Papasaikas (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) and Atul Sethi (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; University Hospital Basel, University of Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) (T12) Modern and scalable tools for efficient analysis of very large metagenomic datasets organized by Alexander Sczyrba (Bielefeld University, Germany), Christian Henke (Bielefeld University, Germany), Clovis Galiez (MPI for Biophysical Chemistry, Germany), Milot Mirdita (PI for Biophysical Chemistry, Germany) and Johannes Soeding (MPI for Biophysical Chemistry, Germany) (T13) User Experience Design for Computational Biologists organized by Nikiforos Karamanis and Xavier Watkins (both at EMBL-EBI, UK) Finally, following the tradition of ECCB, a number of travel fellowships, sponsored by ISCB, were given to young participants. We received 103 applications from scientists from 31 countries. After careful review and consideration, we awarded 14 travel fellowships, mainly to PhD students from nine countries that had an accepted proceedings paper. Volunteering is another important way for young scientists to visit and support the conference, as well as connect with other young scientists. The skills, talent and dedication of the volunteers are expected to largely contribute to the overall quality of the Conference. We made our best to accept all 91 applicants from 16 countries who were eligible to attend the Conference at a very low registration fee. Next to the volunteers, we would like to thank all the people who helped through their work to make this Conference a success. The Theme Chairs and Co-chairs as well all the reviewers who have been the heart of the Conference and helped selecting excellent papers, making thus this Conference a success. We are truly indebted to them. We are also grateful to the ECCB Steering Committee for their advice and continuous the of the organization of ECCB 2018. This conference would be at all the support from Anna Steering Committee and Chair ECCB our application and advice about to make a of was It was our to and inspired by for science with to in a will be in our and in our During the one year of the conference the support of Steering Committee Chair ECCB ECCB and Yves ECCB 2010, ECCB was Yves and for your continuous and We are also grateful to the ISCB Society for Computational for their and in ECCB 2018 at the international level and for This Conference would be the support of our and ELIXIR level and BioExcel was our level Cambridge University Press, Oxford University Press, Community and were ISCB, the Student Council of the ISCB, and who as exhibitors at ECCB 2018 also A thank to all for their and for to make ECCB 2018 an and Conference. A number of people to the organization of ECCB 2018. They beyond their of for making this conference really Maria a as was for the the the venue the and was for and Panagiotis was for the technical support of the registration Dr. organized the members of the DIANA also this Spyros Tastsoglou, Dimitra Karagkouni, Perdikopanis, Dimitrios Finally, the of a conference is its presentations, the conference for its participants. We thank all of for to ECCB 2018. We will the and in the city of Athens, the city and were of Artemis G. Hatzigeorgiou, Pantelis G. Bagos, Panayiotis V. Benos, Christoforos Nikolaou, Yves Moreau, Ioannis Kavakiotis |
Bioinform. | 5 |
| 2018 | pBRIT: gene prioritization by correlating functional and phenotypic annotations through integrative data fusionabstractMotivation: Computational gene prioritization can aid in disease gene identification. Here, we propose pBRIT (prioritization using Bayesian Ridge regression and Information Theoretic model), a novel adaptive and scalable prioritization tool, integrating Pubmed abstracts, Gene Ontology, Sequence similarities, Mammalian and Human Phenotype Ontology, Pathway, Interactions, Disease Ontology, Gene Association database and Human Genome Epidemiology database, into the prediction model. We explore and address effects of sparsity and inter-feature dependencies within annotation sources, and the impact of bias towards specific annotations. Results: pBRIT models feature dependencies and sparsity by an Information-Theoretic (data driven) approach and applies intermediate integration based data fusion. Following the hypothesis that genes underlying similar diseases will share functional and phenotype characteristics, it incorporates Bayesian Ridge regression to learn a linear mapping between functional and phenotype annotations. Genes are prioritized on phenotypic concordance to the training genes. We evaluated pBRIT against nine existing methods, and on over 2000 HPO-gene associations retrieved after construction of pBRIT data sources. We achieve maximum AUC scores ranging from 0.92 to 0.96 against benchmark datasets and of 0.80 against the time-stamped HPO entries, indicating good performance with high sensitivity and specificity. Our model shows stable performance with regard to changes in the underlying annotation data, is fast and scalable for implementation in routine pipelines. Availability and implementation: http://biomina.be/apps/pbrit/; https://bitbucket.org/medgenua/pbrit. Supplementary information: Supplementary data are available at Bioinformatics online. Ajay Anand Kumar, Lut Van Laer, Maaike Alaerts, Amin Ardeshirdavani, Yves Moreau, Kris Laukens, Bart Loeys, Geert Vandeweyer |
Bioinform. | 5 |
| 2018 | Ultra-fast global homology detection with Discrete Cosine Transform and Dynamic Time WarpingabstractMotivation: Evolutionary information is crucial for the annotation of proteins in bioinformatics. The amount of retrieved homologs often correlates with the quality of predicted protein annotations related to structure or function. With a growing amount of sequences available, fast and reliable methods for homology detection are essential, as they have a direct impact on predicted protein annotations. Results: We developed a discriminative, alignment-free algorithm for homology detection with quasi-linear complexity, enabling theoretically much faster homology searches. To reach this goal, we convert the protein sequence into numeric biophysical representations. These are shrunk to a fixed length using a novel vector quantization method which uses a Discrete Cosine Transform compression. We then compute, for each compressed representation, similarity scores between proteins with the Dynamic Time Warping algorithm and we feed them into a Random Forest. The WARP performances are comparable with state of the art methods. Availability and implementation: The method is available at http://ibsquare.be/warp. Supplementary information: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Yves Moreau, Wim F. Vranken |
Bioinform. | 3 |
| 2018 | Gene prioritization using Bayesian matrix factorization with genomic and phenotypic side informationabstractMotivation: Most gene prioritization methods model each disease or phenotype individually, but this fails to capture patterns common to several diseases or phenotypes. To overcome this limitation, we formulate the gene prioritization task as the factorization of a sparsely filled gene-phenotype matrix, where the objective is to predict the unknown matrix entries. To deliver more accurate gene-phenotype matrix completion, we extend classical Bayesian matrix factorization to work with multiple side information sources. The availability of side information allows us to make non-trivial predictions for genes for which no previous disease association is known. Results: Our gene prioritization method can innovatively not only integrate data sources describing genes, but also data sources describing Human Phenotype Ontology terms. Experimental results on our benchmarks show that our proposed model can effectively improve accuracy over the well-established gene prioritization method, Endeavour. In particular, our proposed method offers promising results on diseases of the nervous system; diseases of the eye and adnexa; endocrine, nutritional and metabolic diseases; and congenital malformations, deformations and chromosomal abnormalities, when compared to Endeavour. Availability and implementation: The Bayesian data fusion method is implemented as a Python/C++ package: https://github.com/jaak-s/macau. It is also available as a Julia package: https://github.com/jaak-s/BayesianDataFusion.jl. All data and benchmarks generated or analyzed during this study can be downloaded at https://owncloud.esat.kuleuven.be/index.php/s/UGb89WfkZwMYoTn. Supplementary information: Supplementary data are available at Bioinformatics online. Pooya Zakeri, Jaak Simm, Adam Arany, Sarah ElShal, Yves Moreau |
Bioinform. | 5 |
| 2018 | Towards practical privacy-preserving genome-wide association studyabstractThe deployment of Genome-wide association studies (GWASs) requires genomic information of a large population to produce reliable results. This raises significant privacy concerns, making people hesitate to contribute their genetic information to such studies. We propose two provably secure solutions to address this challenge: (1) a somewhat homomorphic encryption (HE) approach, and (2) a secure multiparty computation (MPC) approach. Unlike previous work, our approach does not rely on adding noise to the input data, nor does it reveal any information about the patients. Our protocols aim to prevent data breaches by calculating the χ 2 statistic in a privacy-preserving manner, without revealing any information other than whether the statistic is significant or not. Specifically, our protocols compute the χ 2 statistic, but only return a yes/no answer, indicating significance. By not revealing the statistic value itself but only the significance, our approach thwarts attacks exploiting statistic values. We significantly increased the efficiency of our HE protocols by introducing a new masking technique to perform the secure comparison that is necessary for determining significance. We show that full-scale privacy-preserving GWAS is practical, as long as the statistics can be computed by low degree polynomials. Our implementations demonstrated that both approaches are efficient. The secure multiparty computation technique completes its execution in approximately 2 ms for data contributed by one million subjects. Charlotte Bonte, Eleftheria Makri, Amin Ardeshirdavani, Jaak Simm, Yves Moreau, Frederik Vercauteren |
BMC Bioinform. | 5 |
| 2016 | Topic modeling of biomedical textabstractThe massive growth of biomedical text makes it very challenging for researchers to review all relevant work and generate all possible hypotheses in a reasonable amount of time. Many text mining methods have been developed to simplify this process and quickly present the researcher with a learned set of biomedical hypotheses that could be potentially validated. Previously, we have focused on the task of identifying genes that are linked with a given disease by text mining the PubMed abstracts. We applied a word-based concept profile similarity to learn patterns between disease and gene entities and hence identify links between them. In this work, we study an alternative approach based on topic modelling to learn different patterns between the disease and the gene entities and measure how well this affects the identified links. We investigated multiple input corpuses, word representations, topic parameters, and similarity measures. On one hand, our results show that when we (1) learn the topics from an input set of gene-clustered set of abstracts, and (2) apply the dot-product similarity measure, we succeed to improve our original methods and identify more correct disease-gene links. On the other hand, the results also show that the learned topics remain limited to the diseases existing in our vocabulary such that scaling the methodology to new disease queries becomes non trivial. Sarah ElShal, Mithila Mathad, Jaak Simm, Jesse Davis, Yves Moreau |
BIBM | 5 |
| 2016 | Highlights from the 11th ISCB Student Council Symposium 2015: Dublin, Ireland. 10 July 2015abstractTable of contents A1 Highlights from the eleventh ISCB Student Council Symposium 2015 Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman O1 Prioritizing a drug’s targets using both gene expression and structural similarity Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau O2 Organism specific protein-RNA recognition: A computational analysis of protein-RNA complex structures from different organisms Nagarajan Raju, Sonia Pankaj Chothani, C. Ramakrishnan, Masakazu Sekijima; M. Michael Gromiha O3 Detection of Heterogeneity in Single Particle Tracking Trajectories Paddy J Slator, Nigel J Burroughs O4 3D-NOME: 3D NucleOme Multiscale Engine for data-driven modeling of three-dimensional genome architecture Przemysław Szałaj, Zhonghui Tang, Paul Michalski, Oskar Luo, Xingwang Li, Yijun Ruan, Dariusz Plewczynski O5 A novel feature selection method to extract multiple adjacent solutions for viral genomic sequences classification Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici O6 A Systems Biology Compendium for Leishmania donovani Bart Cuypers, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens O7 Unravelling signal coordination from large scale phosphorylation kinetic data Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob B. Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan F. DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman, Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau, Raju Nagarajan, Sonia P. Chothani, C. Ramakrishnan, Masakazu Sekijima, M. Michael Gromiha, Paddy Slator, Nigel J. Burroughs, Przemyslaw Szalaj, Zhonghui Tang, Paul J. Michalski, Oskar Luo, Xingwang Li 0004, Yijun Ruan, Dariusz Plewczynski, Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens, Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic |
BMC Bioinform. | 14 |
| 2015 | Gene prioritization through geometric-inspired kernel data fusionabstractIn biology there is often the need to discover the most promising genes, among a large list of candidate genes, to further investigate. While a single data source might not be effective enough, integrating several complementary genomic data sources leads to more accurate prediction. We propose a kernel-based gene prioritization framework using geometric kernel fusion which we have recently developed as a powerful tool for protein fold classification [I]. It has been shown that taking more involved geometry means of their corresponding kernel matrices is less sensitive in dealing with complementary and noisy kernel matrices compared to standard multiple kernel learning methods. Since genomic kernels often encodes the complementary characteristics of biological data, this leads us to research the application of geometric kernel fusion in the gene prioritization task. We utilize an unbiased and prospective benchmark based on the OMIM [2] associations. Experimental results on our prospective benchmark show that our model can improve the accuracy of the state-of-the-art gene prioritization model. Pooya Zakeri, Sarah ElShal, Yves Moreau |
BIBM | 3 |
| 2015 | ISMB/ECCB 2015abstractThis special issue of Bioinformatics serves as the proceedings of the joint 23rd annual meeting of Intelligent Systems for Molecular Biology (ISMB) and 14th European Conference on Computational Biology (ECCB), which took place in Dublin, Ireland, July 10–14, 2015 (http://www.iscb.org/ismbeccb2015). ISMB/ECCB 2015, the official conference of the International Society for Computational Biology (ISCB, http://www.iscb.org/), was accompanied by nine Special Interest Group meetings of 1 or 2 days each, and two satellite meetings. Since its inception, ISMB/ECCB has been the largest international conference in computational biology and bioinformatics. It is the leading forum in the field for presenting new research results, disseminating methods and techniques, and facilitating discussions among leading researchers, practitioners, and students in the field. The 42 papers in this volume were selected from 241 original submissions divided into 13 research areas, collectively led by 25 Area Chairs. For each area, the Area Chairs selected an expert program committee for their subdiscipline and oversaw the reviewing process for that area. By design, the Area Chairs included a mix of experienced individuals reappointed from previous years and experts newly recruited to ensure broad technical expertise and to promote inclusivity of various elements of the research community. In total, the review process involved the 25 Area Chairs, 378 program committee members, and an additional 175 external reviewers recruited as sub-reviewers by program committee members. Table 1 provides a summary of the areas, area chairs and a review summary by area. The conference used a two-tier review system—a continuation and refinement of a process that begun with ISMB/ECCB 2013 in an effort to better ensure thorough and fair reviewing. Under the revised process, each of the 241 submissions was first reviewed by at least three expert referees, with a subset receiving between four and six reviews, as needed. Consensus on each paper was reached through online discussion among reviewers and Area Chairs. Among the 241 submissions, 27 were conditionally accepted for publication directly from the first round review. A subset of 29 papers was viewed as potentially publishable subject to revision and re-review of the manuscripts. Of the 27 papers that were resubmitted, 15 were judged to have addressed the concerns of the reviewers and were accepted for the conference proceedings, resulting in a total of 42 acceptances and an overall acceptance rate of 42/241 = 17.4%. We believe that this two-tier system, which is more reflective of typical multi-round journal review procedures, provided a means of ensuring that only the highest quality original work was accepted within the tight timing constraints imposed by the conference scheduling. We thank all authors for submitting their work. These proceedings would simply not be possible without the scientific ingenuity of the contributors of all the papers. We recognize that the process is not perfect, and some outstanding work might have been rejected despite our best efforts. Nonetheless, we are hopeful that all authors received helpful feedback on their work and that most believed their submissions were judged fairly and diligently. In total, the two-tier review process involved 896 individual reviews. We are immensely grateful to the Area Chairs, the members of the program committee and the external subreviewers for their outstanding efforts in conducting a thorough review process in just 3 months. Their contribution is at the core of the scientific quality of the conference. We also thank Steven Leard for his continuing support with the review process; the team at Oxford University Press for preparing this special proceedings volume and the Conference Chairs, Janet Kelso and Alex Bateman, the Theme Chairs and other members of the ISMB Steering Committee for their advice and supervision. We are also grateful to Russell Schwartz, Proceedings Chair of ISMB 2014, for sharing his experience and various helpful documents on the review process. ISMB/ECCB 2015 review summary by area. ISMB/ECCB 2015 review summary by area. Yves Moreau, Niko Beerenwinkel |
Bioinform. | 1 |
| 2015 | NGS-Logistics: data infrastructure for efficient analysis of NGS sequence variants across multiple centersabstractNext-Generation Sequencing (NGS) is a key tool in genomics, in particular in research and diagnostics of human Mendelian, oligogenic, and complex disorders [ 1 ]. Multiple projects now aim at mapping the human genetic variation on a large scale, such as the 1,000 Genomes Project, the UK 100k Genome Project. Meanwhile with the dramatic decrease of the price and turnaround time, large amounts of human sequencing data have been generated over the past decade [ 2 ]. As of January 2014, about 2,555 sequencers were spread over 920 centers across the world [ 3 ]. As a result, about 100,000 human exome have been sequenced so far [ 4 ]. Crucially, the speed at which NGS data is produced greatly surpasses Moore's law [ 5 ] and challenges our ability to conveniently store, exchange, and analyze this data. Data pre-processing is needed to extract reliable information from sequencing data and it can be divided into two major steps: primary analysis (image analysis and base calling) and secondary analysis. When looking for variation in the human genome, secondary analysis consists of aligning/mapping the reads against the reference genome and scanning the alignment for variation. Both raw data and mapped reads are large files occupying significant disk storage space. The collection of files resulting from the analysis of a single whole genome study can take up to 50Gb of disk space. This raises significant issues in terms of computing and data storage and transfer, with off-site data transfer currently being a key bottleneck. Moreover, the analysis of NGS data also raises the major challenge of how to reconcile federated analysis of personal genomic data and confidentiality of data to protect privacy. In many situations, the analysis of data from a single study alone will be much less powerful than if it can be correlated with other studies. In particular, when investigating a mutation of interest, it is extremely useful to obtain data about other patients or controls sharing similar mutations. However, personal genome data (whole genome, exome, transcriptome data, etc.) is sensitive personal data. Confidentiality of this data must be guaranteed at all times and only duly authorized researchers should access such personal data. To address all challenges described above, we developed a data structure NGS-Logistics , which fulfills all requirements of a successful application that can process data inclusively and comprehensively from multiple sources while guaranteeing privacy and security. NGS-Logistics is a web-based application providing a data structure to analyze NGS data in a distributed way. The data can be located in any data center, anywhere in the world. NGS-Logistics provides an environment in which researchers do not need to worry about the physical location of the data (Figure 1 ). With respect to users rights, queries will be sent to each remote server. The host will process the request and return the results back to the main server where all the privacy limitations are controlled for the data. Once the results are ready, the end user can see the desired information. Depending on the type of query, results will be divided into two parts, the first part is related to the samples to which the user has authorized access, and for which the users can see all details. The second part contains results for the whole population, for which the user has only access to some aggregate statistics without details. An example of such a query would be to review the mutations present at a single genomic position in each individual patient from a set of patients to which the user has authorized access (1st part) and to contrast these results with background frequency of mutation in the reference populations (2nd part) (Figure 2 ). NGS-Logistics components . Users pass their queries from the NGS-Logistics web interface to the clients. Request are stored and scheduled in the main database. Each center has one database, being the only way of communication between centers and the main system. Centers and their databases are connected through a secured connection, to which only valid and trusted IPs are allowed to connect. The query manager is responsible for tracking and running the request, as well as collecting and returning the results to the main system. Single Point Query result page (Statistic section) for chr9:2115841 . The query of chr9:2115841 shows that only one sample is polymorphic at this position. All samples that can be genotyped at this position from the active data set, control data set and whole are homozygous reference. The MAF of this variant in each data set is thus very low. The pilot version of NGS-Logistics has been installed and is currently being beta-tested by users at the Center for Human Genetics of the University of Leuven. Currently we have two installations of the system, the first one at the Leuven University Hospitals and the second one at the Flemish Supercomputing Center (VSC). The development of NGS-Logistics has significantly reduced the effort and time needed to evaluate the significance of mutations from full genome sequencing and exome sequencing, in a safe and confidential environment. This platform provides more opportunities for operators who are interested in expanding their queries and further analysis. Amin Ardeshirdavani, Erika L. Souche, Luc Dehaspe, Jeroen K. J. Van Houdt, Joris Robert Vermeesch, Yves Moreau |
BMC Bioinform. | 6 |
| 2015 | Problems with the nested granularity of feature domains in bioinformatics: the eXtasy caseabstractBACKGROUND: Data from biomedical domains often have an inherit hierarchical structure. As this structure is usually implicit, its existence can be overlooked by practitioners interested in constructing and evaluating predictive models from such data. Ignoring these constructs leads to potentially problematic and the routinely unrecognized bias in the models and results. In this work, we discuss this bias in detail and propose a simple, sampling-based solution for it. Next, we explore its sources and extent on synthetic data. Finally, we demonstrate how the state-of-the-art variant prioritization framework, eXtasy, benefits from using the described approach in its Random forest-based core classification model. RESULTS AND CONCLUSIONS: The conducted simulations clearly indicate that the heterogeneous granularity of feature domains poses significant problems for both the standard Random forest classifier and a modification that relies on stratified bootstrapping. Conversely, using the proposed sampling scheme when training the classifier mitigates the described bias. Furthermore, when applied to the eXtasy data under a realistic class distribution scenario, a Random forest learned using the proposed sampling scheme displays much better precision that its standard version, without degrading recall. Moreover, the largest performance gains are achieved in the most important part of the operating range: the top of prioritized gene list. Dusan Popovic, Alejandro Sifrim, Jesse Davis, Yves Moreau, Bart De Moor |
BMC Bioinform. | 4 |
| 2014 | A Self-Tuning Genetic Algorithm with Applications in Biomarker DiscoveryabstractRecent developments in the field of-omics technologies brought great potential for conducting biomedical research in very efficient manner, but also raised a plethora of new computational challenges to be addressed. Extremely high dimensionality accompanied with poor signal-to-noise ratio and small sample size of data resulting from high-throughput experiments pose previously unprecedented problem, creating an increasing demand for innovative analytical strategies. In this work we propose an island model-based genetic algorithm for multivariate feature selection in the context of-omics data, which accommodates to a particular classification scenario via dynamic tuning of its parameters. We demonstrate it on two publicly available data sets containing gene expression profiles corresponding to the two distinct biomedical questions. We show that the algorithm consistently outperforms two additional feature selection schemes across data sets, regardless to which method is used in the subsequent classification step. Dusan Popovic, Charalampos N. Moschopoulos, Ryo Sakai, Alejandro Sifrim, Jan Aerts, Yves Moreau, Bart De Moor |
CBMS | 6 |
| 2014 | ECCB 2014: The 13th European Conference on Computational BiologyabstractInternational audience Marie-Dominique Devignes, Yves Moreau |
Bioinform. | 2 |
| 2014 | Protein fold recognition using geometric kernel data fusionabstractMOTIVATION: Various approaches based on features extracted from protein sequences and often machine learning methods have been used in the prediction of protein folds. Finding an efficient technique for integrating these different protein features has received increasing attention. In particular, kernel methods are an interesting class of techniques for integrating heterogeneous data. Various methods have been proposed to fuse multiple kernels. Most techniques for multiple kernel learning focus on learning a convex linear combination of base kernels. In addition to the limitation of linear combinations, working with such approaches could cause a loss of potentially useful information. RESULTS: We design several techniques to combine kernel matrices by taking more involved, geometry inspired means of these matrices instead of convex linear combinations. We consider various sequence-based protein features including information extracted directly from position-specific scoring matrices and local sequence alignment. We evaluate our methods for classification on the SCOP PDB-40D benchmark dataset for protein fold recognition. The best overall accuracy on the protein fold recognition test set obtained by our methods is ∼ 86.7%. This is an improvement over the results of the best existing approach. Moreover, our computational model has been developed by incorporating the functional domain composition of proteins through a hybridization model. It is observed that by using our proposed hybridization model, the protein fold recognition accuracy is further improved to 89.30%. Furthermore, we investigate the performance of our approach on the protein remote homology detection problem by fusing multiple string kernels. AVAILABILITY AND IMPLEMENTATION: The MATLAB code used for our proposed geometric kernel fusion frameworks are publicly available at http://people.cs.kuleuven.be/∼raf.vandebril/homepage/software/geomean.php?menu=5/. Pooya Zakeri, Ben Jeuris, Raf Vandebril, Yves Moreau |
Bioinform. | 4 |
| 2013 | eXtasy simplified-towards opening the black boxabstractExome sequencing remarkably simplifies the search for mutations causing rare monogenic disorders. Still, due to a big number of potential candidate variants, computational methods are needed to facilitate this process. Recently, an algorithm based on genomic data fusion has been proposed in this context (eXtasy), which exhibits highly competitive performances among the state of the art methods. Nonetheless, being based on a Random Forest classifier, its core model is characterized by a prohibitive size, slow execution speed and difficulties associated with gaining insights in the decision-making process. Here we propose a simplification of the original eXtasy algorithm that retains superior ranking capability of former without suffering from the both high complexity and low interpretability. Dusan Popovic, Alejandro Sifrim, Yves Moreau, Bart De Moor |
BIBM | 3 |
| 2013 | A Genetic Algorithm for Pancreatic Cancer Diagnosis
Charalampos N. Moschopoulos, Dusan Popovic, Alejandro Sifrim, Grigorios N. Beligiannis, Bart De Moor, Yves Moreau |
EANN (2) | 6 |
| 2013 | A Hybrid Approach to Feature Ranking for Microarray Data Classification
Dusan Popovic, Alejandro Sifrim, Charalampos N. Moschopoulos, Yves Moreau, Bart De Moor |
EANN (2) | 4 |
| 2012 | Applying Kernel Methods on Protein Complexes Detection Problem
Charalampos N. Moschopoulos, Griet Laenen, George D. Kritikos, Yves Moreau |
EANN | 4 |
| 2012 | An unbiased evaluation of gene prioritization toolsabstractMOTIVATION: Gene prioritization aims at identifying the most promising candidate genes among a large pool of candidates-so as to maximize the yield and biological relevance of further downstream validation experiments and functional studies. During the past few years, several gene prioritization tools have been defined, and some of them have been implemented and made available through freely available web tools. In this study, we aim at comparing the predictive performance of eight publicly available prioritization tools on novel data. We have performed an analysis in which 42 recently reported disease-gene associations from literature are used to benchmark these tools before the underlying databases are updated. RESULTS: Cross-validation on retrospective data provides performance estimate likely to be overoptimistic because some of the data sources are contaminated with knowledge from disease-gene association. Our approach mimics a novel discovery more closely and thus provides more realistic performance estimates. There are, however, marked differences, and tools that rely on more advanced data integration schemes appear more powerful. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniela Börnigen, Léon-Charles Tranchevent, Francisco Bonachela Capdevila, Koenraad Devriendt, Bart De Moor, Patrick De Causmaecker, Yves Moreau |
Bioinform. | 7 |
| 2012 | ReLiance: a machine learning and literature-based prioritization of receptor - ligand pairingsabstractMOTIVATION: The prediction of receptor-ligand pairings is an important area of research as intercellular communications are mediated by the successful interaction of these key proteins. As the exhaustive assaying of receptor-ligand pairs is impractical, a computational approach to predict pairings is necessary. We propose a workflow to carry out this interaction prediction task, using a text mining approach in conjunction with a state of the art prediction method, as well as a widely accessible and comprehensive dataset. Among several modern classifiers, random forests have been found to be the best at this prediction task. The training of this classifier was carried out using an experimentally validated dataset of Database of Ligand-Receptor Partners (DLRP) receptor-ligand pairs. New examples, co-cited with the training receptors and ligands, are then classified using the trained classifier. After applying our method, we find that we are able to successfully predict receptor-ligand pairs within the GPCR family with a balanced accuracy of 0.96. Upon further inspection, we find several supported interactions that were not present in the Database of Interacting Proteins (DIPdatabase). We have measured the balanced accuracy of our method resulting in high quality predictions stored in the available database ReLiance. AVAILABILITY: http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB/index.php CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau |
Bioinform. | 7 |
| 2012 | A bioinformatics e-dating story: computational prediction and prioritization of receptor-ligand pairsabstractRegulation of cellular events is initiated, often, via extracellular signaling when a circulating protein ligand interacts with one or more membrane-bound protein receptors. Identification of receptor-ligand pairs is thus an important and difficult task to address as this form of interaction is transient and not well studied. In order to address this problem, we collect the most readily available data from repositories (expression, domain, pathway, sequence, and text-based), and apply a high through-put analysis to this problem. We have worked on the receptor-ligand pairing problem in three main studies. In our first study, using a LS-SVM classifier, we show that we are able to more aptly match members of the chemokine and tgfβ families than a previously published method [ 1 ]. Notably, we are able to achieve an increase in recall of 0.76 over the 0.44 for the matching of receptor-ligands in the tgfβ family. In our subsequent study, we benchmarked several machine learning techniques, and essayed several parameters, on the receptior-ligand interaction prediction task. We found that we could reach a balanced accuracy of 0.84. In our final work, we produce a publicly available database of our results with respect to a text-based in silico prediction workflow. The resulting database, contains several key findings, particularly predictions in the GPCR family with a balanced accuracy of 0.96. The receptor-ligand prediction task is an essential one, as the challenge of predicting such pairs is an important issue in wet-labs, biotech, and pharmaceutical companies. Through several studies, we have determined the most appropriate methodology to predict the receptor-ligand pairs and have made available high-quality predictions at our ReLianceDB website ( http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB ), a tool to aid in performing effective and targeted research. Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau |
BMC Bioinform. | 7 |
| 2012 | Optimized Data Fusion for Kernel k-Means ClusteringabstractThis paper presents a novel optimized kernel k-means algorithm (OKKC) to combine multiple data sources for clustering analysis. The algorithm uses an alternating minimization framework to optimize the cluster membership and kernel coefficients as a nonconvex problem. In the proposed algorithm, the problem to optimize the cluster membership and the problem to optimize the kernel coefficients are all based on the same Rayleigh quotient objective; therefore the proposed algorithm converges locally. OKKC has a simpler procedure and lower complexity than other algorithms proposed in the literature. Simulated and real-life data fusion applications are experimentally studied, and the results validate that the proposed algorithm has comparable performance, moreover, it is more efficient on large-scale data sets. (The Matlab implementation of OKKC algorithm is downloadable from http://homes.esat.kuleuven.be/~sistawww/bio/syu/okkc.html.). Léon-Charles Tranchevent, Xinhai Liu, Wolfgang Glänzel, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2011 | A guide to web tools to prioritize candidate genesabstractFinding the most promising genes among large lists of candidate genes has been defined as the gene prioritization problem. It is a recurrent problem in genetics in which genetic conditions are reported to be associated with chromosomal regions. In the last decade, several different computational approaches have been developed to tackle this challenging task. In this study, we review 19 computational solutions for human gene prioritization that are freely accessible as web tools and illustrate their differences. We summarize the various biological problems to which they have been successfully applied. Ultimately, we describe several research directions that could increase the quality and applicability of the tools. In addition we developed a website (http://www.esat.kuleuven.be/gpp) containing detailed information about these and other tools, which is regularly updated. This review and the associated website constitute together a guide to help users select a gene prioritization strategy that suits best their needs. Léon-Charles Tranchevent, Francisco Bonachela Capdevila, Daniela Nitsch, Bart De Moor, Patrick De Causmaecker, Yves Moreau |
Briefings Bioinform. | 6 |
| 2011 | Optimized data fusion for K-means Laplacian clusteringabstractMOTIVATION: We propose a novel algorithm to combine multiple kernels and Laplacians for clustering analysis. The new algorithm is formulated on a Rayleigh quotient objective function and is solved as a bi-level alternating minimization procedure. Using the proposed algorithm, the coefficients of kernels and Laplacians can be optimized automatically. RESULTS: Three variants of the algorithm are proposed. The performance is systematically validated on two real-life data fusion applications. The proposed Optimized Kernel Laplacian Clustering (OKLC) algorithms perform significantly better than other methods. Moreover, the coefficients of kernels and Laplacians optimized by OKLC show some correlation with the rank of performance of individual data source. Though in our evaluation the K values are predefined, in practical studies, the optimal cluster number can be consistently estimated from the eigenspectrum of the combined kernel Laplacian matrix. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/oklc.html. Xinhai Liu, Léon-Charles Tranchevent, Wolfgang Glänzel, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
Bioinform. | 7 |
| 2011 | Predicting Receptor-Ligand Pairs through Kernel LearningabstractBACKGROUND: Regulation of cellular events is, often, initiated via extracellular signaling. Extracellular signaling occurs when a circulating ligand interacts with one or more membrane-bound receptors. Identification of receptor-ligand pairs is thus an important and specific form of PPI prediction. RESULTS: Given a set of disparate data sources (expression data, domain content, and phylogenetic profile) we seek to predict new receptor-ligand pairs. We create a combined kernel classifier and assess its performance with respect to the Database of Ligand-Receptor Partners (DLRP) 'golden standard' as well as the method proposed by Gertz et al. Among our findings, we discover that our predictions for the tgfβ family accurately reconstruct over 76% of the supported edges (0.76 recall and 0.67 precision) of the receptor-ligand bipartite graph defined by the DLRP "golden standard". In addition, for the tgfβ family, the combined kernel classifier is able to relatively improve upon the Gertz et al. work by a factor of approximately 1.5 when considering that our method has an F-measure of 0.71 while that of Gertz et al. has a value of 0.48. CONCLUSIONS: The prediction of receptor-ligand pairings is a difficult and complex task. We have demonstrated that using kernel learning on multiple data sources provides a stronger alternative to the existing method in solving this task. Ernesto Iacucci, Fabian Ojeda, Bart De Moor, Yves Moreau |
BMC Bioinform. | 4 |
| 2010 | Towards Better Receptor-Ligand Prioritization: How Machine Learning on Protein-Protein Interaction Data Can Provide Insight Into Receptor-Ligand Pairs
Ernesto Iacucci, Yves Moreau |
ICANN (1) | 2 |
| 2010 | Biological knowledge bases using Wikis: combining the flexibility of Wikis with the structure of databasesabstractSUMMARY: In recent years, the number of knowledge bases developed using Wiki technology has exploded. Unfortunately, next to their numerous advantages, classical Wikis present a critical limitation: the invaluable knowledge they gather is represented as free text, which hinders their computational exploitation. This is in sharp contrast with the current practice for biological databases where the data is made available in a structured way. Here, we present WikiOpener an extension for the classical MediaWiki engine that augments Wiki pages by allowing on-the-fly querying and formatting resources external to the Wiki. Those resources may provide data extracted from databases or DAS tracks, or even results returned by local or remote bioinformatics analysis tools. This also implies that structured data can be edited via dedicated forms. Hence, this generic resource combines the structure of biological databases with the flexibility of collaborative Wikis. AVAILABILITY: The source code and its documentation are freely available on the MediaWiki website: http://www.mediawiki.org/wiki/Extension:WikiOpener. Sylvain Brohée, Roland Barriot, Yves Moreau |
Bioinform. | 3 |
| 2010 | Large-scale benchmark of Endeavour using MetaCore mapsabstractAbstract Summary: Endeavour is a tool that detects the most promising genes within large lists of candidates with respect to a biological process of interest and by combining several genomic data sources. We have benchmarked Endeavour using 450 pathway maps and 826 disease marker sets from MetaCoreTM of GeneGo, Inc. containing a total of 9911 and 12 432 genes, respectively. We obtained an area under the receiver operating characteristic curves of 0.97 for pathway and of 0.91 for disease gene sets. These results indicate that Endeavour can be used to efficiently prioritize candidate genes for pathways and diseases. Availability: Endeavour is available at http://www.esat.kuleuven.be/endeavour Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Sven Schuierer, Léon-Charles Tranchevent, Uwe Dengler, Yves Moreau |
Bioinform. | 4 |
| 2010 | Candidate gene prioritization by network analysis of differential expression using machine learning approachesabstractBACKGROUND: Discovering novel disease genes is still challenging for diseases for which no prior knowledge--such as known disease genes or disease-related pathways--is available. Performing genetic studies frequently results in large lists of candidate genes of which only few can be followed up for further investigation. We have recently developed a computational method for constitutional genetic disorders that identifies the most promising candidate genes by replacing prior knowledge by experimental data of differential gene expression between affected and healthy individuals.To improve the performance of our prioritization strategy, we have extended our previous work by applying different machine learning approaches that identify promising candidate genes by determining whether a gene is surrounded by highly differentially expressed genes in a functional association or protein-protein interaction network. RESULTS: We have proposed three strategies scoring disease candidate genes relying on network-based machine learning approaches, such as kernel ridge regression, heat kernel, and Arnoldi kernel approximation. For comparison purposes, a local measure based on the expression of the direct neighbors is also computed. We have benchmarked these strategies on 40 publicly available knockout experiments in mice, and performance was assessed against results obtained using a standard procedure in genetics that ranks candidate genes based solely on their differential expression levels (Simple Expression Ranking). Our results showed that our four strategies could outperform this standard procedure and that the best results were obtained using the Heat Kernel Diffusion Ranking leading to an average ranking position of 8 out of 100 genes, an AUC value of 92.3% and an error reduction of 52.8% relative to the standard procedure approach which ranked the knockout gene on average at position 17 with an AUC value of 83.7%. CONCLUSION: In this study we could identify promising candidate genes using network based machine learning approaches even if no knowledge is available about the disease or phenotype. Daniela Nitsch, Joana P. Gonçalves, Fabian Ojeda, Bart De Moor, Yves Moreau |
BMC Bioinform. | 5 |
| 2010 | Estimating the individualized HIV-1 genetic barrier to resistance using a nelfinavir fitness landscapeabstractBACKGROUND: Failure on Highly Active Anti-Retroviral Treatment is often accompanied with development of antiviral resistance to one or more drugs included in the treatment. In general, the virus is more likely to develop resistance to drugs with a lower genetic barrier. Previously, we developed a method to reverse engineer, from clinical sequence data, a fitness landscape experienced by HIV-1 under nelfinavir (NFV) treatment. By simulation of evolution over this landscape, the individualized genetic barrier to NFV resistance may be estimated for an isolate. RESULTS: We investigated the association of estimated genetic barrier with risk of development of NFV resistance at virological failure, in 201 patients that were predicted fully susceptible to NFV at baseline, and found that a higher estimated genetic barrier was indeed associated with lower odds for development of resistance at failure (OR 0.62 (0.45 - 0.94), per additional mutation needed, p = .02). CONCLUSIONS: Thus, variation in individualized genetic barrier to NFV resistance may impact effective treatment options available after treatment failure. If similar results apply for other drugs, then estimated genetic barrier may be a new clinical tool for choice of treatment regimen, which allows consideration of available treatment options after virological failure. Kristof Theys, Koen Deforche, Gertjan Beheydt, Yves Moreau, Kristel Van Laethem, Philippe Lemey, Ricardo Camacho, Soo-Yon Rhee, Robert W. Shafer, Eric Van Wijngaerden, Anne-Mieke Vandamme |
BMC Bioinform. | 4 |
| 2010 | L2-norm multiple kernel learning and its application to biomedical data fusionabstractBACKGROUND: This paper introduces the notion of optimizing different norms in the dual problem of support vector machines with multiple kernels. The selection of norms yields different extensions of multiple kernel learning (MKL) such as L(infinity), L1, and L2 MKL. In particular, L2 MKL is a novel method that leads to non-sparse optimal kernel coefficients, which is different from the sparse kernel coefficients optimized by the existing L(infinity) MKL method. In real biomedical applications, L2 MKL may have more advantages over sparse integration method for thoroughly combining complementary information in heterogeneous data sources. RESULTS: We provide a theoretical analysis of the relationship between the L2 optimization of kernels in the dual problem with the L2 coefficient regularization in the primal problem. Understanding the dual L2 problem grants a unified view on MKL and enables us to extend the L2 method to a wide range of machine learning problems. We implement L2 MKL for ranking and classification problems and compare its performance with the sparse L(infinity) and the averaging L1 MKL methods. The experiments are carried out on six real biomedical data sets and two large scale UCI data sets. L2 MKL yields better performance on most of the benchmark data sets. In particular, we propose a novel L2 MKL least squares support vector machine (LSSVM) algorithm, which is shown to be an efficient and promising classifier for large scale data sets processing. CONCLUSIONS: This paper extends the statistical framework of genomic data fusion based on MKL. Allowing non-sparse weights on the data sources is an attractive option in settings where we believe most data sources to be relevant to the problem at hand and want to avoid a "winner-takes-all" effect seen in L(infinity) MKL, which can be detrimental to the performance in prospective studies. The notion of optimizing L2 kernels can be straightforwardly extended to ranking, classification, regression, and clustering algorithms. To tackle the computational burden of MKL, this paper proposes several novel LSSVM based MKL algorithms. Systematic comparison on real data sets shows that LSSVM MKL has comparable performance as the conventional SVM MKL algorithms. Moreover, large scale numerical experiments indicate that when cast as semi-infinite programming, LSSVM MKL can be solved more efficiently than SVM MKL. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/l2lssvm.html. Tillmann Falck, Anneleen Daemen, Léon-Charles Tranchevent, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
BMC Bioinform. | 7 |
| 2010 | Gene prioritization and clustering by multi-view text miningabstractBACKGROUND: Text mining has become a useful tool for biologists trying to understand the genetics of diseases. In particular, it can help identify the most interesting candidate genes for a disease for further experimental analysis. Many text mining approaches have been introduced, but the effect of disease-gene identification varies in different text mining models. Thus, the idea of incorporating more text mining models may be beneficial to obtain more refined and accurate knowledge. However, how to effectively combine these models still remains a challenging question in machine learning. In particular, it is a non-trivial issue to guarantee that the integrated model performs better than the best individual model. RESULTS: We present a multi-view approach to retrieve biomedical knowledge using different controlled vocabularies. These controlled vocabularies are selected on the basis of nine well-known bio-ontologies and are applied to index the vast amounts of gene-based free-text information available in the MEDLINE repository. The text mining result specified by a vocabulary is considered as a view and the obtained multiple views are integrated by multi-source learning algorithms. We investigate the effect of integration in two fundamental computational disease gene identification tasks: gene prioritization and gene clustering. The performance of the proposed approach is systematically evaluated and compared on real benchmark data sets. In both tasks, the multi-view approach demonstrates significantly better performance than other comparing methods. CONCLUSIONS: In practical research, the relevance of specific vocabulary pertaining to the task is usually unknown. In such case, multi-view text mining is a superior and promising strategy for text-based disease gene identification. Léon-Charles Tranchevent, Bart De Moor, Yves Moreau |
BMC Bioinform. | 4 |
| 2010 | Weighted hybrid clustering by combining text mining and bibliometrics on a large-scale journal databaseabstractAbstract We propose a new hybrid clustering framework to incorporate text mining with bibliometrics in journal set analysis. The framework integrates two different approaches: clustering ensemble and kernel‐fusion clustering. To improve the flexibility and the efficiency of processing large‐scale data, we propose an information‐based weighting scheme to leverage the effect of multiple data sources in hybrid clustering. Three different algorithms are extended by the proposed weighting scheme and they are employed on a large journal set retrieved from the Web of Science (WoS) database. The clustering performance of the proposed algorithms is systematically evaluated using multiple evaluation methods, and they were cross‐compared with alternative methods. Experimental results demonstrate that the proposed weighted hybrid clustering strategy is superior to other methods in clustering performance and efficiency. The proposed approach also provides a more refined structural mapping of journal sets, which is useful for monitoring and detecting new trends in different scientific fields. Xinhai Liu, Frizo A. L. Janssens, Wolfgang Glänzel, Yves Moreau, Bart De Moor |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2009 | Hybrid Clustering of Text Mining and Bibliometrics Applied to Journal SetsabstractTo obtain correlated and complementary information contained in text mining and bibliometrics, hybrid clustering to incorporate textual content and citation information has become a popular strategy. In this paper, we propose a new computational framework of integrating text mining and bibliometrics to provide a mapping of journal sets. Two different approaches of hybrid clustering methods are applied in this paper. The first category is ensemble clustering, which combines different clustering results obtained from individual data into a consolidated clustering result. The second category is kernel fusion, which maps heterogeneous data sets into the kernel space and combines the kernel matrices for clustering. Kernels can be combined either averagely, or by an optimized weighted linear combination model. In this paper, we propose a novel adaptive kernel K-means clustering algorithm to combine textual content and citation information for clustering. The proposed algorithm is systematically compared with other methods on a clustering problem of 1869 journals published in 2002–2006. Based on several validation indices, the experimental results demonstrate that our hybrid clustering strategy is able to provide clustering result as well as the best individual data source. Xinhai Liu, Yves Moreau, Bart De Moor, Wolfgang Glänzel, Frizo A. L. Janssens |
SDM | 3 |
| 2009 | An experimental loop design for the detection of constitutional chromosomal aberrations by array CGHabstractBACKGROUND: Comparative genomic hybridization microarrays for the detection of constitutional chromosomal aberrations is the application of microarray technology coming fastest into routine clinical application. Through genotype-phenotype association, it is also an important technique towards the discovery of disease causing genes and genomewide functional annotation in human. When using a two-channel microarray of genomic DNA probes for array CGH, the basic setup consists in hybridizing a patient against a normal reference sample. Two major disadvantages of this setup are (1) the use of half of the resources to measure a (little informative) reference sample and (2) the possibility that deviating signals are caused by benign copy number variation in the "normal" reference instead of a patient aberration. Instead, we apply an experimental loop design that compares three patients in three hybridizations. RESULTS: We develop and compare two statistical methods (linear models of log ratios and mixed models of absolute measurements). In an analysis of 27 patients seen at our genetics center, we observed that the linear models of the log ratios are advantageous over the mixed models of the absolute intensities. CONCLUSION: The loop design and the performance of the statistical analysis contribute to the quick adoption of array CGH as a routine diagnostic tool. They lower the detection limit of mosaicisms and improve the assignment of copy number variation for genetic association studies. Joke Allemeersch, Steven Van Vooren, Femke Hannes, Bart De Moor, Joris Robert Vermeesch, Yves Moreau |
BMC Bioinform. | 6 |
| 2008 | Estimation of an in vivo fitness landscape experienced by HIV-1 under drug selective pressure useful for prediction of drug resistance evolution during treatmentabstractMOTIVATION: HIV-1 antiviral resistance is a major cause of antiviral treatment failure. The in vivo fitness landscape experienced by the virus in presence of treatment could in principle be used to determine both the susceptibility of the virus to the treatment and the genetic barrier to resistance. We propose a method to estimate this fitness landscape from cross-sectional clinical genetic sequence data of different subtypes, by reverse engineering the required selective pressure for HIV-1 sequences obtained from treatment naive patients, to evolve towards sequences obtained from treated patients. The method was evaluated for recovering 10 random fictive selective pressures in simulation experiments, and for modeling the selective pressure under treatment with the protease inhibitor nelfinavir. RESULTS: The estimated fitness function under nelfinavir treatment considered fitness contributions of 114 mutations at 48 sites. Estimated fitness correlated significantly with the in vitro resistance phenotype in 519 matched genotype-phenotype pairs (R(2) = 0.47 (0.41 - 0.54)) and variation in predicted evolution under nelfinavir selective pressure correlated significantly with observed in vivo evolution during nelfinavir treatment for 39 mutations (with FDR = 0.05). AVAILABILITY: The software is available on request from the authors, and data sets are available from http://jose.med.kuleuven.be/~kdforc0/nfv-fitness-data/. Koen Deforche, Ricardo Camacho, Kristel Van Laethem, Philippe Lemey, Andrew Rambaut, Yves Moreau, Anne-Mieke Vandamme |
Bioinform. | 6 |
| 2007 | Query-driven module discovery in microarray dataabstractAbstract Motivation: Existing (bi)clustering methods for microarray data analysis often do not answer the specific questions of interest to a biologist. Such specific questions could be derived from other information sources, including expert prior knowledge. More specifically, given a set of seed genes which are believed to have a common function, we would like to recruit genes with similar expression profiles as the seed genes in a significant subset of experimental conditions. Results: We introduce QDB, a novel Bayesian query-driven biclustering framework in which the prior distributions allow introducing knowledge from a set of seed genes (query) to guide the pattern search. In two well-known yeast compendia, we grow highly functionally enriched biclusters from small sets of seed genes using a resolution sweep approach. In addition, relevant conditions are identified and modularity of the biclusters is demonstrated, including the discovery of overlapping modules. Finally, our method deals with missing values naturally, performs well on artificial data from a recent biclustering benchmark study and has a number of conceptual advantages when compared to existing approaches for focused module search. Availability: Software is available on the Supplementary Material. Contact: [email protected] Supplementary information: Available on http://homes.esat.kuleuven.be/~tdhollan/Supplementary_Information_Dhollander_2007/index.html Thomas Dhollander, Qizheng Sheng, Karen Lemmens, Bart De Moor, Kathleen Marchal, Yves Moreau |
Bioinform. | 6 |
| 2007 | CATMA, a comprehensive genome-scale resource for silencing and transcript profiling of Arabidopsis genesabstractBACKGROUND: The Complete Arabidopsis Transcript MicroArray (CATMA) initiative combines the efforts of laboratories in eight European countries 1 to deliver gene-specific sequence tags (GSTs) for the Arabidopsis research community. The CATMA initiative offers the power and flexibility to regularly update the GST collection according to evolving knowledge about the gene repertoire. These GST amplicons can easily be reamplified and shared, subsets can be picked at will to print dedicated arrays, and the GSTs can be cloned and used for other functional studies. This ongoing initiative has already produced approximately 24,000 GSTs that have been made publicly available for spotted microarray printing and RNA interference. RESULTS: GSTs from the CATMA version 2 repertoire (CATMAv2, created in 2002) were mapped onto the gene models from two independent Arabidopsis nuclear genome annotation efforts, TIGR5 and PSB-EuGène, to consolidate a list of genes that were targeted by previously designed CATMA tags. A total of 9,027 gene models were not tagged by any amplified CATMAv2 GST, and 2,533 amplified GSTs were no longer predicted to tag an updated gene model. To validate the efficacy of GST mapping criteria and design rules, the predicted and experimentally observed hybridization characteristics associated to GST features were correlated in transcript profiling datasets obtained with the CATMAv2 microarray, confirming the reliability of this platform. To complete the CATMA repertoire, all 9,027 gene models for which no GST had yet been designed were processed with an adjusted version of the Specific Primer and Amplicon Design Software (SPADS). A total of 5,756 novel GSTs were designed and amplified by PCR from genomic DNA. Together with the pre-existing GST collection, this new addition constitutes the CATMAv3 repertoire. It comprises 30,343 unique amplified sequences that tag 24,202 and 23,009 protein-encoding nuclear gene models in the TAIR6 and EuGène genome annotations, respectively. To cover the remaining untagged genes, we identified 543 additional GSTs using less stringent design criteria and designed 990 sequence tags matching multiple members of gene families (Gene Family Tags or GFTs) to cover any remaining untagged genes. These latter 1,533 features constitute the CATMAv4 addition. CONCLUSION: To update the CATMA GST repertoire, we designed 7,289 additional sequence tags, bringing the total number of tagged TAIR6-annotated Arabidopsis nuclear protein-coding genes to 26,173. This resource is used both for the production of spotted microarrays and the large-scale cloning of hairpin RNA silencing vectors. All information about the resulting updated CATMA repertoire is available through the CATMA database http://www.catma.org. Gert Sclep, Joke Allemeersch, Robin Liechti, Björn De Meyer, Jim Beynon, Rishikesh Bhalerao, Yves Moreau, Wilfried Nietfeld, Jean-Pierre Renou, Philippe Reymond, Martin Kuiper, Pierre Hilson |
BMC Bioinform. | 7 |
| 2006 | Analysis of HIV-1 pol sequences using Bayesian Networks: implications for drug resistanceabstractHuman Immunodeficiency Virus-1 (HIV-1) antiviral resistance is a major cause of antiviral therapy failure and compromises future treatment options. As a consequence, resistance testing is the standard of care. Because of the high degree of HIV-1 natural variation and complex interactions, the role of resistance mutations is in many cases insufficiently understood. We applied a probabilistic model, Bayesian networks, to analyze direct influences between protein residues and exposure to treatment in clinical HIV-1 protease sequences from diverse subtypes. We can determine the specific role of many resistance mutations against the protease inhibitor nelfinavir, and determine relationships between resistance mutations and polymorphisms. We can show for example that in addition to the well-known major mutations 90M and 30N for nelfinavir resistance, 88S should not be treated as 88D but instead considered as a major mutation and explain the subtype-dependent prevalence of the 30N resistance pathway. Koen Deforche, Tomi Silander, Ricardo Camacho, Zehava Grossman, M. A. Soares, Kristel Van Laethem, Rami Kantor, Yves Moreau, Anne-Mieke Vandamme |
Bioinform. | 8 |
| 2005 | BioMart and Bioconductor: a powerful link between biological databases and microarray data analysisabstractbiomaRt is a new Bioconductor package that integrates BioMart data resources with data analysis software in Bioconductor. It can annotate a wide range of gene or gene product identifiers (e.g. Entrez-Gene and Affymetrix probe identifiers) with information such as gene symbol, chromosomal coordinates, Gene Ontology and OMIM annotation. Furthermore biomaRt enables retrieval of genomic sequences and single nucleotide polymorphism information, which can be used in data analysis. Fast and up-to-date data retrieval is possible as the package executes direct SQL queries to the BioMart databases (e.g. Ensembl). The biomaRt package provides a tight integration of large, public or locally installed BioMart databases with data analysis in Bioconductor creating a powerful environment for biological data mining. Steffen Durinck, Yves Moreau, Arek Kasprzyk, Sean R. Davis, Bart De Moor, Alvis Brazma, Wolfgang Huber |
Bioinform. | 2 |
| 2005 | arrayCGHbase: an analysis platform for comparative genomic hybridization microarraysabstractBACKGROUND: The availability of the human genome sequence as well as the large number of physically accessible oligonucleotides, cDNA, and BAC clones across the entire genome has triggered and accelerated the use of several platforms for analysis of DNA copy number changes, amongst others microarray comparative genomic hybridization (arrayCGH). One of the challenges inherent to this new technology is the management and analysis of large numbers of data points generated in each individual experiment. RESULTS: We have developed arrayCGHbase, a comprehensive analysis platform for arrayCGH experiments consisting of a MIAME (Minimal Information About a Microarray Experiment) supportive database using MySQL underlying a data mining web tool, to store, analyze, interpret, compare, and visualize arrayCGH results in a uniform and user-friendly format. Following its flexible design, arrayCGHbase is compatible with all existing and forthcoming arrayCGH platforms. Data can be exported in a multitude of formats, including BED files to map copy number information on the genome using the Ensembl or UCSC genome browser. CONCLUSION: ArrayCGHbase is a web based and platform independent arrayCGH data analysis tool, that allows users to access the analysis suite through the internet or a local intranet after installation on a private server. ArrayCGHbase is available at http://medgen.ugent.be/arrayCGHbase/. Björn Menten, Filip Pattyn, Katleen De Preter, Piet Robbrecht, Evi Michels, Karen Buysse, Geert Mortier, Anne De Paepe, Steven Van Vooren, Joris Robert Vermeesch, Yves Moreau, Bart De Moor, Stefan Vermeulen, Frank Speleman, Jo Vandesompele |
BMC Bioinform. | 11 |
| 2004 | Using literature and data to learn Bayesian networks as clinical models of ovarian tumors
Péter Antal, Geert Fannes, Dirk Timmerman, Yves Moreau, Bart De Moor |
Artif. Intell. Medicine | 4 |
| 2004 | A genetic algorithm for the detection of new cis-regulatory modules in sets of coregulated genesabstractSUMMARY: The implementation of a genetic algorithm is described that provides a fast method of searching for the optimal combination of transcription factor binding sites in a set of regulatory sequences. AVAILABILITY: The algorithm can be used transparently as a web service from within the Toucan software. Toucan can be accessed at http://www.esat.kuleuven.ac.be/~saerts/software/toucan.php. A standalone version of the software is available upon request. Stein Aerts, Peter Van Loo, Yves Moreau, Bart De Moor |
Bioinform. | 3 |
| 2004 | Importing MAGE-ML format microarray data into BioConductorabstractUNLABELLED: The microarray gene expression markup language (MAGE-ML) is a widely used XML (eXtensible Markup Language) standard for describing and exchanging information about microarray experiments. It can describe microarray designs, microarray experiment designs, gene expression data and data analysis results. We describe RMAGEML, a new Bioconductor package that provides a link between cDNA microarray data stored in MAGE-ML format and the Bioconductor framework for preprocessing, visualization and analysis of microarray experiments. AVAILABILITY: http://www.bioconductor.org. Open Source. Steffen Durinck, Joke Allemeersch, Vincent Carey, Yves Moreau, Bart De Moor |
Bioinform. | 4 |
| 2003 | Bayesian applications of belief networks and multilayer perceptrons for ovarian tumor classification with rejection
Péter Antal, Geert Fannes, Dirk Timmerman, Yves Moreau, Bart De Moor |
Artif. Intell. Medicine | 4 |
| 2002 | Web-based Data Collection for Uterine Adnexal Tumors: A Case StudyabstractWe have developed a World Wide Web application for the collection of EPRs (electronic patient records) from uterine adnexal masses pre-operatively examined with transvaginal ultrasonography. The application has been used intensively since November 2000 by nine of the 19 international centers that joined the International Ovarian Tumor Analysis (IOTA) consortium. The IOTA database contains 68 parameters for 1,150 masses. We report the design and implementation of the generic Web-based clinical data entry system and describe the advantages and drawbacks that we have experienced while developing, using and maintaining the system. The data model, the user interface, the help system, the constraints (mandatory/optional) and the quality checking were all based on the medical protocol created by the IOTA consortium. The data collection system has become an open and transparent implementation of the formalized protocol. It covers the complete path of the patient data from the clinical situation to the finalized database. This approach provides new types of possibilities for the data analysis, since all aspects of the data collection are documented and formally available to the data analyst. The IOTA Web site can be found at, which also serves as the entry point for the secure EPR application. Stein Aerts, Péter Antal, Dirk Timmerman, Bart De Moor, Yves Moreau |
CBMS | 5 |
| 2002 | Adaptive quality-based clustering of gene expression profilesabstractMOTIVATION: Microarray experiments generate a considerable amount of data, which analyzed properly help us gain a huge amount of biologically relevant information about the global cellular behaviour. Clustering (grouping genes with similar expression profiles) is one of the first steps in data analysis of high-throughput expression measurements. A number of clustering algorithms have proved useful to make sense of such data. These classical algorithms, though useful, suffer from several drawbacks (e.g. they require the predefinition of arbitrary parameters like the number of clusters; they force every gene into a cluster despite a low correlation with other cluster members). In the following we describe a novel adaptive quality-based clustering algorithm that tackles some of these drawbacks. RESULTS: We propose a heuristic iterative two-step algorithm: First, we find in the high-dimensional representation of the data a sphere where the "density" of expression profiles is locally maximal (based on a preliminary estimate of the radius of the cluster-quality-based approach). In a second step, we derive an optimal radius of the cluster (adaptive approach) so that only the significantly coexpressed genes are included in the cluster. This estimation is achieved by fitting a model to the data using an EM-algorithm. By inferring the radius from the data itself, the biologist is freed from finding an optimal value for this radius by trial-and-error. The computational complexity of this method is approximately linear in the number of gene expression profiles in the data set. Finally, our method is successfully validated using existing data sets. AVAILABILITY: http://www.esat.kuleuven.ac.be/~thijs/Work/Clustering.html Frank De Smet, Janick Mathys, Kathleen Marchal, Gert Thijs, Bart De Moor, Yves Moreau |
Bioinform. | 6 |
| 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif SamplingabstractAbstract Summary: INCLUSive allows automatic multistep analysis of microarray data (clustering and motif finding). The clustering algorithm (adaptive quality-based clustering) groups together genes with highly similar expression profiles. The upstream sequences of the genes belonging to a cluster are automatically retrieved from GenBank and can be fed directly into Motif Sampler, a Gibbs sampling algorithm that retrieves statistically over-represented motifs in sets of sequences, in this case upstream regions of co-expressed genes. Availability: For academic purposes at http://www.esat.kuleuven.ac.be/~dna/BioI/Software.html Contact: [email protected] * To whom correspondence should be addressed. Email: [email protected]. Gert Thijs, Yves Moreau, Frank De Smet, Janick Mathys, Magali Lescot, Stephane Rombauts, Pierre Rouzé, Bart De Moor, Kathleen Marchal |
Bioinform. | 2 |
| 2002 | Functional bioinformatics of microarray data: from expression to regulationabstractUsing microarrays is a powerful technique to monitor the expression of thousands of genes in a single experiment. From series of such experiments, it is possible to identify the mechanisms that govern the activation of genes in an organism. Short deoxyribonucleic acid patterns (called binding sites) near the genes serve as switches that control gene expression. As a result similar patterns of expression can correspond to similar binding site patterns. Here we integrate clustering of coexpressed genes with the discovery of binding motifs. We overview several important clustering techniques and present a clustering algorithm (called adaptive quality-based clustering), which we have developed to address several shortcomings of existing methods. We overview the different techniques for motif finding, in particular the technique of Gibbs sampling, and we present several extensions of this technique in our Motif Sampler. Finally, we present an integrated web tool called INCLUSive (available online at http://www.esat.kuleuven.ac.be//spl sim/dna/BioI/Software.html) that allows the easy analysis of microarray data for motif finding. Yves Moreau, Frank De Smet, Gert Thijs, Kathleen Marchal, Bart De Moor |
Proc. IEEE | 1 |
| 2001 | Extended Bayesian Regression Models: A Symbiotic Application of Belief Networks and Multilayer Perceptrons for the Classification of Ovarian Tumors
Péter Antal, Geert Fannes, Bart De Moor, Joos Vandewalle, Yves Moreau, Dirk Timmerman |
AIME | 5 |
| 2001 | A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genesabstractMicroarray experiments can reveal useful information on the transcriptional regulation. We try to find regulatory elements in the region upstream of translation start of coexpressed genes. Here we present a modification to the original Gibbs Sampling algorithm [12]. We introduce a probability distribution to estimate the number of copies of the motif in a sequence. The second modification is the incorporation of a higher-order background model. We have successfully tested our algorithm on several data sets. First we show results on two selected data set: sequences from plants containing the G-box motif and the upstream sequences from bacterial genes regulated by O2-responsive protein FNR. In both cases the motif sampler is able to find the expected motifs. Finally, the sampler is tested on 4 clusters of coexpressed genes from a wounding experiment in Arabidopsis thaliana. We find several putative motifs that are related to the pathways involved in the plant defense mechanism. Gert Thijs, Kathleen Marchal, Magali Lescot, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau |
RECOMB | 7 |
| 2001 | A higher-order background model improves the detection of promoter regulatory elements by Gibbs samplingabstractMOTIVATION: Transcriptome analysis allows detection and clustering of genes that are coexpressed under various biological circumstances. Under the assumption that coregulated genes share cis-acting regulatory elements, it is important to investigate the upstream sequences controlling the transcription of these genes. To improve the robustness of the Gibbs sampling algorithm to noisy data sets we propose an extension of this algorithm for motif finding with a higher-order background model. RESULTS: Simulated data and real biological data sets with well-described regulatory elements are used to test the influence of the different background models on the performance of the motif detection algorithm. We show that the use of a higher-order model considerably enhances the performance of our motif finding algorithm in the presence of noisy data. For Arabidopsis thaliana, a reliable background model based on a set of carefully selected intergenic sequences was constructed. AVAILABILITY: Our implementation of the Gibbs sampler called the Motif Sampler can be used through a web interface: http://www.esat.kuleuven.ac.be/~thijs/Work/MotifSampler.html. CONTACT: [email protected]; [email protected] Gert Thijs, Magali Lescot, Kathleen Marchal, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau |
Bioinform. | 7 |
| 1999 | A hybrid system for fraud detection in mobile communications
Yves Moreau, Ellen Lerouge, Herman Verrelst, Joos Vandewalle, Christof Störmann, Peter Burge |
ESANN | 1 |
| 1999 | Embedding recurrent neural networks into predator-prey models
Yves Moreau, Stéphane Louiès, Joos Vandewalle, Léon Brenig |
Neural Networks | 1 |
| 1998 | To stop learning using the evidence
Yves Moreau, Joos Vandewalle |
ESANN | 1 |
| 1997 | Composition methods for the integration of dynamical neural networks
Yves Moreau, Joos Vandewalle |
ESANN | 1 |
| 1997 | Detection of Mobile Phone Fraud Using Supervised Neural Networks: A First Prototype
Yves Moreau, Herman Verrelst, Joos Vandewalle |
ICANN | 1 |
| 1997 | Use of a Multi-Layer Perceptron to Predict Malignancy in Ovarian Tumors
Herman Verrelst, Yves Moreau, Joos Vandewalle, Dirk Timmerman |
NIPS | 2 |
| 1996 | Prediction of dynamical systems with composition networks
Yves Moreau, Joos Vandewalle |
ESANN | 1 |