EDBT 2026 Demo / reviewers in the wild / expert
Bart De Moor
dblp:66/3559 · also B. L. R. De Moor
· DBLP profile ↗
136ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-1154-5028ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 42 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 1 since 2021Human-computer interaction and ubiquitous computing · 4Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
19 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
4 papers |
Learning theory · 47% Kernel, tree and ensemble methods · 39% Probabilistic and Bayesian machine learning · 14% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 90% Information retrieval · 10% | |
| Theoretical computer science
2 papers |
Algorithms and data structures · 60% Mathematical optimization · 20% Information theory · 20% |
Topics — the 30 heaviest of 47, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.4 | 8 | 2007 | CALIB: a Bioconductor package for estimating absolute expression levels from two-color microarray data · Bioinform. 2007 A calibration method for estimating absolute expression levels from microarray data · Bioinform. 2006 M@CBETH: a microarray classification benchmarking tool · Bioinform. 2005 |
Bioinformatics and computational biology
gene expression analysis |
0.3 | 6 | 2007 | CALIB: a Bioconductor package for estimating absolute expression levels from two-color microarray data · Bioinform. 2007 Query-driven module discovery in microarray data · Bioinform. 2007 Importing MAGE-ML format microarray data into BioConductor · Bioinform. 2004 |
Data mining
clustering |
0.3 | 3 | 2013 | Multiview Partitioning via Tensor Methods · IEEE Trans. Knowl. Data Eng. 2013 Dynamic hybrid clustering of bioinformatics by incorporating text mining and citation analysis · KDD 2007 Query-driven module discovery in microarray data · Bioinform. 2007 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.2 | 1 | 2014 | EnsembleSVM: a library for ensemble learning using support vector machines · J. Mach. Learn. Res. 2014 |
Machine learning › Kernel, tree and ensemble methods
support vector machine |
0.2 | 1 | 2014 | EnsembleSVM: a library for ensemble learning using support vector machines · J. Mach. Learn. Res. 2014 |
Machine learning › Probabilistic and Bayesian machine learning
derivative estimation |
0.2 | 1 | 2013 | Derivative estimation with local polynomial fitting · J. Mach. Learn. Res. 2013 |
Machine learning › Learning theory
nonparametric regression |
0.2 | 1 | 2013 | Derivative estimation with local polynomial fitting · J. Mach. Learn. Res. 2013 |
Data mining › clustering
multi-view clustering |
0.2 | 1 | 2013 | Multiview Partitioning via Tensor Methods · IEEE Trans. Knowl. Data Eng. 2013 |
Data mining › clustering
spectral clustering |
0.2 | 1 | 2013 | Multiview Partitioning via Tensor Methods · IEEE Trans. Knowl. Data Eng. 2013 |
Bioinformatics and computational biology
biomedical text mining |
0.1 | 1 | 2012 | ReLiance: a machine learning and literature-based prioritization of receptor - ligand pairings · Bioinform. 2012 |
Bioinformatics and computational biology › genomics › computational genomics
gene prioritization |
0.1 | 1 | 2012 | An unbiased evaluation of gene prioritization tools · Bioinform. 2012 |
Algorithms and data structures
clustering |
0.1 | 1 | 2012 | Optimized Data Fusion for Kernel k-Means Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Algorithms and data structures › clustering › k-means clustering
kernel k-means |
0.1 | 1 | 2012 | Optimized Data Fusion for Kernel k-Means Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Machine learning › Learning theory › nonparametric regression
kernel regression |
0.1 | 1 | 2011 | Kernel Regression in the Presence of Correlated Errors · J. Mach. Learn. Res. 2011 |
Bioinformatics and computational biology
multi-omics data integration |
0.1 | 1 | 2011 | Optimized data fusion for K-means Laplacian clustering · Bioinform. 2011 |
Bioinformatics and computational biology
transcriptomics |
0.1 | 1 | 2009 | ViTraM: visualization of transcriptional modules · Bioinform. 2009 |
Information theory › probability theory › measure concentration
concentration inequalities |
0.1 | 1 | 2009 | Least conservative support and tolerance tubes · IEEE Trans. Inf. Theory 2009 |
Mathematical optimization
statistical estimation |
0.1 | 1 | 2009 | Least conservative support and tolerance tubes · IEEE Trans. Inf. Theory 2009 |
Machine learning › Learning theory
generalization bounds |
0.1 | 1 | 2007 | A Risk Minimization Principle for a Class of Parzen Estimators · NIPS 2007 |
Machine learning › Learning theory › generalization bounds
rademacher complexity |
0.1 | 1 | 2007 | A Risk Minimization Principle for a Class of Parzen Estimators · NIPS 2007 |
Bioinformatics and computational biology › gene expression analysis
biclustering |
0.1 | 1 | 2007 | Query-driven module discovery in microarray data · Bioinform. 2007 |
Information retrieval
citation analysis |
0.1 | 1 | 2007 | Dynamic hybrid clustering of bioinformatics by incorporating text mining and citation analysis · KDD 2007 |
Data mining › clustering
hybrid clustering |
0.1 | 1 | 2007 | Dynamic hybrid clustering of bioinformatics by incorporating text mining and citation analysis · KDD 2007 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.1 | 2 | 2013 | Derivative estimation with local polynomial fitting · J. Mach. Learn. Res. 2013 A Risk Minimization Principle for a Class of Parzen Estimators · NIPS 2007 |
Bioinformatics and computational biology › sequence analysis
motif discovery |
0.1 | 2 | 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002 A higher-order background model improves the detection of promoter regulatory elements by Gibbs sampling · Bioinform. 2001 |
Bioinformatics and computational biology › gene expression analysis
gene expression quantification |
0.1 | 1 | 2006 | A calibration method for estimating absolute expression levels from microarray data · Bioinform. 2006 |
Bioinformatics and computational biology › data integration
biological database integration |
0.1 | 1 | 2005 | BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis · Bioinform. 2005 |
Bioinformatics and computational biology › genome annotation
gene annotation |
0.1 | 1 | 2005 | BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis · Bioinform. 2005 |
Bioinformatics and computational biology › cancer genomics
cancer classification |
0.0 | 1 | 2004 | Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reduction · Bioinform. 2004 |
Bioinformatics and computational biology › gene regulation › regulatory element discovery
cis-regulatory module prediction |
0.0 | 1 | 2004 | A genetic algorithm for the detection of new cis-regulatory modules in sets of coregulated genes · Bioinform. 2004 |
Methods — techniques the papers use, named apart from their topics
alternating minimization · 0.3cross-validation · 0.2support vector machine · 0.2ensemble learning · 0.2tensor decomposition · 0.2local polynomial fitting · 0.2frobenius norm · 0.2text mining · 0.1rayleigh quotient optimization · 0.1random forest · 0.1co-citation analysis · 0.1benchmarking · 0.1gibbs sampling · 0.1rayleigh quotient · 0.1multiple kernel learning · 0.1kernel regression · 0.1module detection · 0.1fano's inequality · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calibration in Multiple Instance Learning: Evaluating Aggregation Methods for Ultrasound-Based Diagnosis
Axel Geysels, Ben Van Calster, Bart De Moor, Wouter Froyman, Dirk Timmerman |
MICCAI (15) | 3 |
| 2025 | DSS4EX: A Decision Support System framework to explore Artificial Intelligence pipelines with an application in time series forecastingabstractUnderstanding complex Artificial Intelligence (AI) pipelines for time series forecasting can be challenging for both experts and non-experts. This work introduces DSS4EX, a Decision Support System (DSS) framework designed to facilitate the exploration and comprehension of AI pipelines. The framework is demonstrated through a software demo currently in a prototype phase, specifically tailored for electricity demand forecasting that utilizes the Decomposition-Residuals Deep Neural Network (DR-DNN) pipeline. The dataset used in the software demo, spanning March 18th, 2017, to February 16th, 2021, contains seven hourly time series from an unknown location. It covers 34,360 timesteps of data, including power demand, air pressure, cloud coverage, humidity, temperature, wind direction, and wind speed. The prototype software demo was created to showcase the usability of the DSS4EX framework, allowing users to interactively configure, visualize, and understand the AI pipeline. This interactive approach enhances user engagement and comprehension, promoting informed decision-making. The DSS4EX framework makes complex models more transparent and accessible, bridging the gap between advanced AI models and user understanding. This research demonstrates the potential of DSS4EX in supporting further advancements in AI applications by making sophisticated AI techniques more understandable and usable for a broader audience. Giulia Rinaldi, Konstantinos Theodorakos, Fernando Crema Garcia, Oscar Mauricio Agudelo, Bart De Moor |
Expert Syst. Appl. | 5 |
| 2024 | Length of stay prediction for hospital management using domain adaptationabstractAnticipating inpatient length of stay (LoS) can be of crucial aid for efficient hospital management, admissions planning, resource allocation and care quality improvement. Anticipating LoS can be achieved by applying machine learning techniques on historical patient data for developing predictive models. Ethically, these models cannot serve as substitutes for medical unit heads, solely responsible for authorizing patients’ discharge, but they can be exploited by management systems for effective hospital planning. Therefore, the prediction system should be designed and adapted to work in a true hospital setting. This study focuses on early hospital LoS prediction for different admission units and employs domain adaptation to leverage information learned from a potential source domain. Time series data from 110,079 admissions into 8 intensive care units (ICUs) and 60,492 admissions into 9 ICUs were respectively extracted from eICU-CRD and MIMIC-IV and fed into a Long-Short Term Memory network and a Fully connected network to train a source domain model. Learned weights were then transferred either partially or fully from a source domain to target domains. SHapley Additive exPlanations (SHAP) algorithms were used to study the effect of weight transfer on feature importance. Compared against the benchmark, the proposed weight transfer model showed statistically significant gains in prediction accuracy (between 1% and 5%) as well as computation time (up to 2 h saved) for some target domains. The proposed method provides an adaptable clinical decision support system for hospital management that can facilitate processes of data access via the ethical committee, computation infrastructures and time. Lyse Naomi Wamba Momo, Nyalleng Moorosi, Elaine O. Nsoesie, Frank E. Rademakers, Bart De Moor |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Client Recruitment for Federated Learning in ICU Length of Stay PredictionabstractMachine and deep learning methods for medical and healthcare applications have shown significant progress and performance improvement in recent years. These methods require vast amounts of training data which are available in the medical sector, albeit decentralized. Medical institutions generate vast amounts of data for which sharing and centralizing remains a challenge as the result of data and privacy regulations. Federated Learning (FL) is well-suited to tackle these challenges. However, FL comes with a new set of open problems related to communication overhead, efficient parameter aggregation, client selection strategies and more. In this work, we address the step prior to the initiation of a federated network for model training, client recruitment. By intelligently recruiting clients, communication overhead and overall cost of training can be reduced without sacrificing predictive performance. Client recruitment aims at pre-excluding potential clients from partaking in the federation based on a set of criteria indicative of their eventual contributions to the federation. In this work, we propose a client recruitment approach using only the output distribution and sample size at the client site. We show how a subset of clients can be recruited without sacrificing model performance whilst significantly improving computation time. By applying the recruitment approach to the training of federated models for accurate patient Length of Stay prediction using data from 189 intensive care units (clients), we show how the models trained in federations made up from only recruited clients significantly outperform federated models trained with the standard procedure in terms of predictive power and training time. Vincent Scheltjens, Lyse Naomi Wamba Momo, Wouter Verbeke, Bart De Moor |
e-Science | 4 |
| 2023 | Multi-view kernel PCA for time series forecastingabstractIn this paper, we propose a kernel principal component analysis model for multi-variate time series forecasting, where the training and prediction schemes are derived from the multi-view formulation of Restricted Kernel Machines. The training problem is simply an eigenvalue decomposition of the summation of two kernel matrices corresponding to the views of the input and output data. When a linear kernel is used for the output view, it is shown that the forecasting equation takes the form of kernel ridge regression. When that kernel is non-linear, a pre-image problem has to be solved to forecast a point in the input space. We evaluate the model on several standard time series datasets, perform ablation studies, benchmark with closely related models and discuss its results. Arun Pandey, Hannes De Meulemeester, Bart De Moor, Johan A. K. Suykens |
Neurocomputing | 3 |
| 2023 | Island Transpeciation: A Co-Evolutionary Neural Architecture Search, Applied to Country-Scale Air-Quality ForecastingabstractAir pollution causes around 400 000 premature deaths per year in Europe due to Particulate Matter, nitrogen oxides, and ground-level ozone pollutants. Multiple-input multiple-output nonlinear auto-regressive exogenous deep neural networks are frequently used to predict a day before, air-quality pollution incidents, at a country scale. With complexity and data sizes increasing, finding performant models becomes harder. We propose island transpeciation to optimize hyperparameters and architectures. Unlike using a single optimizer, island transpeciation combines results from multiple optimizers, to consistently provide excellent performance. Moreover, we show that island transpeciation outperforms random model search and other previous modeling efforts. Island transpeciation is a neural architecture search that uses co-evolution (genes), to combine (transpeciation) populations of incompatible optimizers (species) organized in island formations. In island transpeciation, architecture search is parallelized and utilizes a distributed pool of hardware resources. We have successfully used these techniques to predict next-day ozone concentrations across the Belgian territory. Konstantinos Theodorakos, Oscar Mauricio Agudelo, Joachim Schreurs, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Evol. Comput. | 5 |
| 2022 | Recurrent Restricted Kernel Machines for Time-series ForecastingabstractIn this paper, we propose a novel method for time-series modeling and forecasting.It is based on the temporal formulation of Restricted Kernel Machines leading to a dynamical equation in the latent-variables.Forecasting involves finding the next latent variable and then solving a pre-image problem to predict a new-point in the input space.Further, we benchmark our model on several standard data sets against other well-known time-series models. Arun Pandey, Hannes De Meulemeester, Henri De Plaen, Bart De Moor, Johan A. K. Suykens |
ESANN | 4 |
| 2021 | The Bures Metric for Generative Adversarial Networks
Hannes De Meulemeester, Joachim Schreurs, Michaël Fanuel, Bart De Moor, Johan A. K. Suykens |
ECML/PKDD (2) | 4 |
| 2020 | Unsupervised Embeddings for Categorical VariablesabstractReal-world data sets often contain both continuous and categorical variables yet most popular machine learning methods cannot by default handle both data types. This creates the need for researchers to transform their data into a continuous format. When no prior information is available, the most widely applied methods are simple ones such as one-hot encoding. However, they ignore many possible sources of information, in particular, categorical dependencies, which could enrich the vector representations. We investigate the effect of natural language processing techniques for learning continuous word-vector representations on categorical variables. We show empirically that the learned vector representations of the categorical variables capture information about the variables themselves and their dependencies with other variables similar to how word embeddings capture semantic and syntactic information. We also show that machine learning models using unsupervised categorical embeddings are competitive with supervised embeddings, and outperform them when fine-tuned, on various classification benchmark data sets. Hannes De Meulemeester, Bart De Moor |
IJCNN | 2 |
| 2015 | Problems with the nested granularity of feature domains in bioinformatics: the eXtasy caseabstractBACKGROUND: Data from biomedical domains often have an inherit hierarchical structure. As this structure is usually implicit, its existence can be overlooked by practitioners interested in constructing and evaluating predictive models from such data. Ignoring these constructs leads to potentially problematic and the routinely unrecognized bias in the models and results. In this work, we discuss this bias in detail and propose a simple, sampling-based solution for it. Next, we explore its sources and extent on synthetic data. Finally, we demonstrate how the state-of-the-art variant prioritization framework, eXtasy, benefits from using the described approach in its Random forest-based core classification model. RESULTS AND CONCLUSIONS: The conducted simulations clearly indicate that the heterogeneous granularity of feature domains poses significant problems for both the standard Random forest classifier and a modification that relies on stratified bootstrapping. Conversely, using the proposed sampling scheme when training the classifier mitigates the described bias. Furthermore, when applied to the eXtasy data under a realistic class distribution scenario, a Random forest learned using the proposed sampling scheme displays much better precision that its standard version, without degrading recall. Moreover, the largest performance gains are achieved in the most important part of the operating range: the top of prioritized gene list. Dusan Popovic, Alejandro Sifrim, Jesse Davis, Yves Moreau, Bart De Moor |
BMC Bioinform. | 5 |
| 2015 | A robust ensemble approach to learn from positive and unlabeled data using SVM base models
Marc Claesen, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
Neurocomputing | 4 |
| 2014 | A Self-Tuning Genetic Algorithm with Applications in Biomarker DiscoveryabstractRecent developments in the field of-omics technologies brought great potential for conducting biomedical research in very efficient manner, but also raised a plethora of new computational challenges to be addressed. Extremely high dimensionality accompanied with poor signal-to-noise ratio and small sample size of data resulting from high-throughput experiments pose previously unprecedented problem, creating an increasing demand for innovative analytical strategies. In this work we propose an island model-based genetic algorithm for multivariate feature selection in the context of-omics data, which accommodates to a particular classification scenario via dynamic tuning of its parameters. We demonstrate it on two publicly available data sets containing gene expression profiles corresponding to the two distinct biomedical questions. We show that the algorithm consistently outperforms two additional feature selection schemes across data sets, regardless to which method is used in the subsequent classification step. Dusan Popovic, Charalampos N. Moschopoulos, Ryo Sakai, Alejandro Sifrim, Jan Aerts, Yves Moreau, Bart De Moor |
CBMS | 7 |
| 2014 | New Bandwidth Selection Criterion for Kernel PCA: Approach to Dimensionality Reduction and Classification ProblemsabstractBACKGROUND: DNA microarrays are potentially powerful technology for improving diagnostic classification, treatment selection, and prognostic assessment. The use of this technology to predict cancer outcome has a history of almost a decade. Disease class predictors can be designed for known disease cases and provide diagnostic confirmation or clarify abnormal cases. The main input to this class predictors are high dimensional data with many variables and few observations. Dimensionality reduction of these features set significantly speeds up the prediction task. Feature selection and feature transformation methods are well known preprocessing steps in the field of bioinformatics. Several prediction tools are available based on these techniques. RESULTS: Studies show that a well tuned Kernel PCA (KPCA) is an efficient preprocessing step for dimensionality reduction, but the available bandwidth selection method for KPCA was computationally expensive. In this paper, we propose a new data-driven bandwidth selection criterion for KPCA, which is related to least squares cross-validation for kernel density estimation. We propose a new prediction model with a well tuned KPCA and Least Squares Support Vector Machine (LS-SVM). We estimate the accuracy of the newly proposed model based on 9 case studies. Then, we compare its performances (in terms of test set Area Under the ROC Curve (AUC) and computational time) with other well known techniques such as whole data set + LS-SVM, PCA + LS-SVM, t-test + LS-SVM, Prediction Analysis of Microarrays (PAM) and Least Absolute Shrinkage and Selection Operator (Lasso). Finally, we assess the performance of the proposed strategy with an existing KPCA parameter tuning algorithm by means of two additional case studies. CONCLUSION: We propose, evaluate, and compare several mathematical/statistical techniques, which apply feature transformation/selection for subsequent classification, and consider its application in medical diagnostics. Both feature selection and feature transformation perform well on classification tasks. Due to the dynamic selection property of feature selection, it is hard to define significant features for the classifier, which predicts classes of future samples. Moreover, the proposed strategy enjoys a distinctive advantage with its relatively lesser time complexity. Minta Thomas, Kris De Brabanter, Bart De Moor |
BMC Bioinform. | 3 |
| 2014 | Predicting breast cancer using an expression values weighted clinical classifierabstractBACKGROUND: Clinical data, such as patient history, laboratory analysis, ultrasound parameters-which are the basis of day-to-day clinical decision support-are often used to guide the clinical management of cancer in the presence of microarray data. Several data fusion techniques are available to integrate genomics or proteomics data, but only a few studies have created a single prediction model using both gene expression and clinical data. These studies often remain inconclusive regarding an obtained improvement in prediction performance. To improve clinical management, these data should be fully exploited. This requires efficient algorithms to integrate these data sets and design a final classifier. LS-SVM classifiers and generalized eigenvalue/singular value decompositions are successfully used in many bioinformatics applications for prediction tasks. While bringing up the benefits of these two techniques, we propose a machine learning approach, a weighted LS-SVM classifier to integrate two data sources: microarray and clinical parameters. RESULTS: We compared and evaluated the proposed methods on five breast cancer case studies. Compared to LS-SVM classifier on individual data sets, generalized eigenvalue decomposition (GEVD) and kernel GEVD, the proposed weighted LS-SVM classifier offers good prediction performance, in terms of test area under ROC Curve (AUC), on all breast cancer case studies. CONCLUSIONS: Thus a clinical classifier weighted with microarray data set results in significantly improved diagnosis, prognosis and prediction responses to therapy. The proposed model has been shown as a promising mathematical framework in both data fusion and non-linear classification problems. Minta Thomas, Kris De Brabanter, Johan A. K. Suykens, Bart De Moor |
BMC Bioinform. | 4 |
| 2014 | Incremental kernel spectral clustering for online learning of non-stationary data
Rocco Langone, Oscar Mauricio Agudelo, Bart De Moor, Johan A. K. Suykens |
Neurocomputing | 3 |
| 2014 | EnsembleSVM: a library for ensemble learning using support vector machines
Marc Claesen, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
J. Mach. Learn. Res. | 4 |
| 2014 | Maximum Likelihood Estimation ofGEVD: Applications in BioinformaticsabstractWe propose a method, maximum likelihood estimation of generalized eigenvalue decomposition (MLGEVD) that employs a well known technique relying on the generalization of singular value decomposition (SVD). The main aim of the work is to show the tight equivalence between MLGEVD and generalized ridge regression. This relationship reveals an important mathematical property of GEVD in which the second argument act as prior information in the model. Thus we show that MLGEVD allows the incorporation of external knowledge about the quantities of interest into the estimation problem. We illustrate the importance of prior knowledge in clinical decision making/identifying differentially expressed genes with case studies for which microarray data sets with corresponding clinical/literature information are available. On all of these three case studies, MLGEVD outperformed GEVD on prediction in terms of test area under the ROC curve (test AUC). MLGEVD results in significantly improved diagnosis, prognosis and prediction of therapy response. Minta Thomas, Anneleen Daemen, Bart De Moor |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | eXtasy simplified-towards opening the black boxabstractExome sequencing remarkably simplifies the search for mutations causing rare monogenic disorders. Still, due to a big number of potential candidate variants, computational methods are needed to facilitate this process. Recently, an algorithm based on genomic data fusion has been proposed in this context (eXtasy), which exhibits highly competitive performances among the state of the art methods. Nonetheless, being based on a Random Forest classifier, its core model is characterized by a prohibitive size, slow execution speed and difficulties associated with gaining insights in the decision-making process. Here we propose a simplification of the original eXtasy algorithm that retains superior ranking capability of former without suffering from the both high complexity and low interpretability. Dusan Popovic, Alejandro Sifrim, Yves Moreau, Bart De Moor |
BIBM | 4 |
| 2013 | A Genetic Algorithm for Pancreatic Cancer Diagnosis
Charalampos N. Moschopoulos, Dusan Popovic, Alejandro Sifrim, Grigorios N. Beligiannis, Bart De Moor, Yves Moreau |
EANN (2) | 5 |
| 2013 | A Hybrid Approach to Feature Ranking for Microarray Data Classification
Dusan Popovic, Alejandro Sifrim, Charalampos N. Moschopoulos, Yves Moreau, Bart De Moor |
EANN (2) | 5 |
| 2013 | Derivative estimation with local polynomial fitting
Kris De Brabanter, Jos De Brabanter, Bart De Moor, Irène Gijbels |
J. Mach. Learn. Res. | 3 |
| 2013 | Multiview Partitioning via Tensor MethodsabstractClustering by integrating multiview representations has become a crucial issue for knowledge discovery in heterogeneous environments. However, most prior approaches assume that the multiple representations share the same dimension, limiting their applicability to homogeneous environments. In this paper, we present a novel tensor-based framework for integrating heterogeneous multiview data in the context of spectral clustering. Our framework includes two novel formulations; that is multiview clustering based on the integration of the Frobenius-norm objective function (MC-FR-OI) and that based on matrix integration in the Frobenius-norm objective function (MC-FR-MI). We show that the solutions for both formulations can be computed by tensor decompositions. We evaluated our methods on synthetic data and two real-world data sets in comparison with baseline methods. Experimental results demonstrate that the proposed formulations are effective in integrating multiview data in heterogeneous environments. Xinhai Liu, Shuiwang Ji, Wolfgang Glänzel, Bart De Moor |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | maximum likelihood estimation and polynomial system solving
Kim Batselier, Philippe Dreesen, Bart De Moor |
ESANN | 3 |
| 2012 | Deconvolution in nonparametric statistics
Kris De Brabanter, Bart De Moor |
ESANN | 2 |
| 2012 | Weighted/Structured Total Least Squares problems and polynomial system solving
Philippe Dreesen, Kim Batselier, Bart De Moor |
ESANN | 3 |
| 2012 | Joint Regression and Linear Combination of Time Series for Optimal Prediction
Dries Geebelen, Kim Batselier, Philippe Dreesen, Marco Signoretto, Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
ESANN | 6 |
| 2012 | Robustness of kernel based regression: Influence and weight functionsabstractIt has been shown that kernel based regression (KBR) with a least squares loss has some undesirable properties from robustness point of view. KBR with more robust loss functions, e.g. Huber or Logistic losses, often give rise to more complicated computations. In classical statistics, robustness is improved by reweighting the original estimate. We study the influence of reweighting the LS-KBR estimate using three well-known weight functions and one new weight function called Myriad. Our results give practical guidelines in order to choose the weights, providing robustness and fast convergence. It turns out that Logistic and Myriad weights are suitable reweighting schemes when outliers are present in the data. In fact, the Myriad shows better performance over the others in the presence of extreme outliers (e.g. Cauchy distributed errors). These findings are then illustrated on toy example as well as on a real life data sets. Finally, we establish an empirical maxbias curve to demonstrate the ability of the proposed methodology. Kris De Brabanter, Jos De Brabanter, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
IJCNN | 5 |
| 2012 | Improved modeling of clinical data with kernel methods
Anneleen Daemen, Dirk Timmerman, Thierry Van den Bosch, Cecilia Bottomley, Emma Kirk, Caroline Van Holsbeke, Lil Valentin, Tom Bourne, Bart De Moor |
Artif. Intell. Medicine | 9 |
| 2012 | An unbiased evaluation of gene prioritization toolsabstractMOTIVATION: Gene prioritization aims at identifying the most promising candidate genes among a large pool of candidates-so as to maximize the yield and biological relevance of further downstream validation experiments and functional studies. During the past few years, several gene prioritization tools have been defined, and some of them have been implemented and made available through freely available web tools. In this study, we aim at comparing the predictive performance of eight publicly available prioritization tools on novel data. We have performed an analysis in which 42 recently reported disease-gene associations from literature are used to benchmark these tools before the underlying databases are updated. RESULTS: Cross-validation on retrospective data provides performance estimate likely to be overoptimistic because some of the data sources are contaminated with knowledge from disease-gene association. Our approach mimics a novel discovery more closely and thus provides more realistic performance estimates. There are, however, marked differences, and tools that rely on more advanced data integration schemes appear more powerful. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniela Börnigen, Léon-Charles Tranchevent, Francisco Bonachela Capdevila, Koenraad Devriendt, Bart De Moor, Patrick De Causmaecker, Yves Moreau |
Bioinform. | 5 |
| 2012 | ReLiance: a machine learning and literature-based prioritization of receptor - ligand pairingsabstractMOTIVATION: The prediction of receptor-ligand pairings is an important area of research as intercellular communications are mediated by the successful interaction of these key proteins. As the exhaustive assaying of receptor-ligand pairs is impractical, a computational approach to predict pairings is necessary. We propose a workflow to carry out this interaction prediction task, using a text mining approach in conjunction with a state of the art prediction method, as well as a widely accessible and comprehensive dataset. Among several modern classifiers, random forests have been found to be the best at this prediction task. The training of this classifier was carried out using an experimentally validated dataset of Database of Ligand-Receptor Partners (DLRP) receptor-ligand pairs. New examples, co-cited with the training receptors and ligands, are then classified using the trained classifier. After applying our method, we find that we are able to successfully predict receptor-ligand pairs within the GPCR family with a balanced accuracy of 0.96. Upon further inspection, we find several supported interactions that were not present in the Database of Interacting Proteins (DIPdatabase). We have measured the balanced accuracy of our method resulting in high quality predictions stored in the available database ReLiance. AVAILABILITY: http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB/index.php CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau |
Bioinform. | 5 |
| 2012 | A bioinformatics e-dating story: computational prediction and prioritization of receptor-ligand pairsabstractRegulation of cellular events is initiated, often, via extracellular signaling when a circulating protein ligand interacts with one or more membrane-bound protein receptors. Identification of receptor-ligand pairs is thus an important and difficult task to address as this form of interaction is transient and not well studied. In order to address this problem, we collect the most readily available data from repositories (expression, domain, pathway, sequence, and text-based), and apply a high through-put analysis to this problem. We have worked on the receptor-ligand pairing problem in three main studies. In our first study, using a LS-SVM classifier, we show that we are able to more aptly match members of the chemokine and tgfβ families than a previously published method [ 1 ]. Notably, we are able to achieve an increase in recall of 0.76 over the 0.44 for the matching of receptor-ligands in the tgfβ family. In our subsequent study, we benchmarked several machine learning techniques, and essayed several parameters, on the receptior-ligand interaction prediction task. We found that we could reach a balanced accuracy of 0.84. In our final work, we produce a publicly available database of our results with respect to a text-based in silico prediction workflow. The resulting database, contains several key findings, particularly predictions in the GPCR family with a balanced accuracy of 0.96. The receptor-ligand prediction task is an essential one, as the challenge of predicting such pairs is an important issue in wet-labs, biotech, and pharmaceutical companies. Through several studies, we have determined the most appropriate methodology to predict the receptor-ligand pairs and have made available high-quality predictions at our ReLianceDB website ( http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB ), a tool to aid in performing effective and targeted research. Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau |
BMC Bioinform. | 5 |
| 2012 | Optimized Data Fusion for Kernel k-Means ClusteringabstractThis paper presents a novel optimized kernel k-means algorithm (OKKC) to combine multiple data sources for clustering analysis. The algorithm uses an alternating minimization framework to optimize the cluster membership and kernel coefficients as a nonconvex problem. In the proposed algorithm, the problem to optimize the cluster membership and the problem to optimize the kernel coefficients are all based on the same Rayleigh quotient objective; therefore the proposed algorithm converges locally. OKKC has a simpler procedure and lower complexity than other algorithms proposed in the literature. Simulated and real-life data fusion applications are experimentally studied, and the results validate that the proposed algorithm has comparable performance, moreover, it is more efficient on large-scale data sets. (The Matlab implementation of OKKC algorithm is downloadable from http://homes.esat.kuleuven.be/~sistawww/bio/syu/okkc.html.). Léon-Charles Tranchevent, Xinhai Liu, Wolfgang Glänzel, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2012 | Confidence bands for least squares support vector machine classifiers: A regression approach
Kris De Brabanter, Peter Karsmakers, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Pattern Recognit. | 5 |
| 2011 | A guide to web tools to prioritize candidate genesabstractFinding the most promising genes among large lists of candidate genes has been defined as the gene prioritization problem. It is a recurrent problem in genetics in which genetic conditions are reported to be associated with chromosomal regions. In the last decade, several different computational approaches have been developed to tackle this challenging task. In this study, we review 19 computational solutions for human gene prioritization that are freely accessible as web tools and illustrate their differences. We summarize the various biological problems to which they have been successfully applied. Ultimately, we describe several research directions that could increase the quality and applicability of the tools. In addition we developed a website (http://www.esat.kuleuven.be/gpp) containing detailed information about these and other tools, which is regularly updated. This review and the associated website constitute together a guide to help users select a gene prioritization strategy that suits best their needs. Léon-Charles Tranchevent, Francisco Bonachela Capdevila, Daniela Nitsch, Bart De Moor, Patrick De Causmaecker, Yves Moreau |
Briefings Bioinform. | 4 |
| 2011 | Optimized data fusion for K-means Laplacian clusteringabstractMOTIVATION: We propose a novel algorithm to combine multiple kernels and Laplacians for clustering analysis. The new algorithm is formulated on a Rayleigh quotient objective function and is solved as a bi-level alternating minimization procedure. Using the proposed algorithm, the coefficients of kernels and Laplacians can be optimized automatically. RESULTS: Three variants of the algorithm are proposed. The performance is systematically validated on two real-life data fusion applications. The proposed Optimized Kernel Laplacian Clustering (OKLC) algorithms perform significantly better than other methods. Moreover, the coefficients of kernels and Laplacians optimized by OKLC show some correlation with the rank of performance of individual data source. Though in our evaluation the K values are predefined, in practical studies, the optimal cluster number can be consistently estimated from the eigenspectrum of the combined kernel Laplacian matrix. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/oklc.html. Xinhai Liu, Léon-Charles Tranchevent, Wolfgang Glänzel, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
Bioinform. | 6 |
| 2011 | Predicting Receptor-Ligand Pairs through Kernel LearningabstractBACKGROUND: Regulation of cellular events is, often, initiated via extracellular signaling. Extracellular signaling occurs when a circulating ligand interacts with one or more membrane-bound receptors. Identification of receptor-ligand pairs is thus an important and specific form of PPI prediction. RESULTS: Given a set of disparate data sources (expression data, domain content, and phylogenetic profile) we seek to predict new receptor-ligand pairs. We create a combined kernel classifier and assess its performance with respect to the Database of Ligand-Receptor Partners (DLRP) 'golden standard' as well as the method proposed by Gertz et al. Among our findings, we discover that our predictions for the tgfβ family accurately reconstruct over 76% of the supported edges (0.76 recall and 0.67 precision) of the receptor-ligand bipartite graph defined by the DLRP "golden standard". In addition, for the tgfβ family, the combined kernel classifier is able to relatively improve upon the Gertz et al. work by a factor of approximately 1.5 when considering that our method has an F-measure of 0.71 while that of Gertz et al. has a value of 0.48. CONCLUSIONS: The prediction of receptor-ligand pairings is a difficult and complex task. We have demonstrated that using kernel learning on multiple data sources provides a stronger alternative to the existing method in solving this task. Ernesto Iacucci, Fabian Ojeda, Bart De Moor, Yves Moreau |
BMC Bioinform. | 3 |
| 2011 | Kernel Regression in the Presence of Correlated Errors
Kris De Brabanter, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
J. Mach. Learn. Res. | 4 |
| 2011 | Tensor Versus Matrix Completion: A Comparison With Application to Spectral DataabstractTensor completion recently emerged as a generalization of matrix completion for higher order arrays. This problem formulation allows one to exploit the structure of data that intrinsically have multiple dimensions. In this work, we recall a convex formulation for minimum (multilinear) ranks completion of arrays of arbitrary order. Successively we focus on completion of partially observed spectral images; the latter can be naturally represented as third order tensors and typically exhibit intraband correlations. We compare different convex formulations and assess them through case studies. Marco Signoretto, Raf Van de Plas, Bart De Moor, Johan A. K. Suykens |
IEEE Signal Process. Lett. | 3 |
| 2011 | Approximate Confidence and Prediction Intervals for Least Squares Support Vector RegressionabstractBias-corrected approximate 100(1-α)% pointwise and simultaneous confidence and prediction intervals for least squares support vector machines are proposed. A simple way of determining the bias without estimating higher order derivatives is formulated. A variance estimator is developed that works well in the homoscedastic and heteroscedastic case. In order to produce simultaneous confidence intervals, a simple Šidák correction and a more involved correction (based on upcrossing theory) are used. The obtained confidence intervals are compared to a state-of-the-art bootstrap-based method. Simulations show that the proposed method obtains similar intervals compared to the bootstrap at a lower computational cost. Kris De Brabanter, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Neural Networks | 4 |
| 2010 | Polynomial componentwise LS-SVM: Fast variable selection using low rank updatesabstractThis paper describes a Least Squares Support Vector Machines (LS-SVM) approach to estimate additive models as a sum of non-linear components. In particular, this work discusses the low rank matrix modifications for componentwise polynomial kernels, which allow the factors of the modified kernel-matrix to be directly updated. The main concept refers to the use of a valid explicit feature map for polynomial kernels in an additive setting. By exploiting the structure of such feature map the model parameters of the classification/regression problem can be easily modified and updated when new variables are added. Therefore, the low rank updates constitute an algorithmic tool to efficiently obtain the model parameters once the system has been altered in some minimal sense. Such strategy allows, for instance, the development of algorithms for sequential variable ranking in high dimensional settings, while non-linearity is provided by the polynomial feature map. Moreover relevant variables can be robustly ranked using the closed form of the leave-one-out (LOO) error estimator, obtained as a by-product of the low rank modifications. Fabian Ojeda, Tillmann Falck, Bart De Moor, Johan A. K. Suykens |
IJCNN | 3 |
| 2010 | Hybrid Clustering of Multiple Information Sources via HOSVD
Xinhai Liu, Lieven De Lathauwer, Frizo A. L. Janssens, Bart De Moor |
ISNN (2) | 4 |
| 2010 | Candidate gene prioritization by network analysis of differential expression using machine learning approachesabstractBACKGROUND: Discovering novel disease genes is still challenging for diseases for which no prior knowledge--such as known disease genes or disease-related pathways--is available. Performing genetic studies frequently results in large lists of candidate genes of which only few can be followed up for further investigation. We have recently developed a computational method for constitutional genetic disorders that identifies the most promising candidate genes by replacing prior knowledge by experimental data of differential gene expression between affected and healthy individuals.To improve the performance of our prioritization strategy, we have extended our previous work by applying different machine learning approaches that identify promising candidate genes by determining whether a gene is surrounded by highly differentially expressed genes in a functional association or protein-protein interaction network. RESULTS: We have proposed three strategies scoring disease candidate genes relying on network-based machine learning approaches, such as kernel ridge regression, heat kernel, and Arnoldi kernel approximation. For comparison purposes, a local measure based on the expression of the direct neighbors is also computed. We have benchmarked these strategies on 40 publicly available knockout experiments in mice, and performance was assessed against results obtained using a standard procedure in genetics that ranks candidate genes based solely on their differential expression levels (Simple Expression Ranking). Our results showed that our four strategies could outperform this standard procedure and that the best results were obtained using the Heat Kernel Diffusion Ranking leading to an average ranking position of 8 out of 100 genes, an AUC value of 92.3% and an error reduction of 52.8% relative to the standard procedure approach which ranked the knockout gene on average at position 17 with an AUC value of 83.7%. CONCLUSION: In this study we could identify promising candidate genes using network based machine learning approaches even if no knowledge is available about the disease or phenotype. Daniela Nitsch, Joana P. Gonçalves, Fabian Ojeda, Bart De Moor, Yves Moreau |
BMC Bioinform. | 4 |
| 2010 | L2-norm multiple kernel learning and its application to biomedical data fusionabstractBACKGROUND: This paper introduces the notion of optimizing different norms in the dual problem of support vector machines with multiple kernels. The selection of norms yields different extensions of multiple kernel learning (MKL) such as L(infinity), L1, and L2 MKL. In particular, L2 MKL is a novel method that leads to non-sparse optimal kernel coefficients, which is different from the sparse kernel coefficients optimized by the existing L(infinity) MKL method. In real biomedical applications, L2 MKL may have more advantages over sparse integration method for thoroughly combining complementary information in heterogeneous data sources. RESULTS: We provide a theoretical analysis of the relationship between the L2 optimization of kernels in the dual problem with the L2 coefficient regularization in the primal problem. Understanding the dual L2 problem grants a unified view on MKL and enables us to extend the L2 method to a wide range of machine learning problems. We implement L2 MKL for ranking and classification problems and compare its performance with the sparse L(infinity) and the averaging L1 MKL methods. The experiments are carried out on six real biomedical data sets and two large scale UCI data sets. L2 MKL yields better performance on most of the benchmark data sets. In particular, we propose a novel L2 MKL least squares support vector machine (LSSVM) algorithm, which is shown to be an efficient and promising classifier for large scale data sets processing. CONCLUSIONS: This paper extends the statistical framework of genomic data fusion based on MKL. Allowing non-sparse weights on the data sources is an attractive option in settings where we believe most data sources to be relevant to the problem at hand and want to avoid a "winner-takes-all" effect seen in L(infinity) MKL, which can be detrimental to the performance in prospective studies. The notion of optimizing L2 kernels can be straightforwardly extended to ranking, classification, regression, and clustering algorithms. To tackle the computational burden of MKL, this paper proposes several novel LSSVM based MKL algorithms. Systematic comparison on real data sets shows that LSSVM MKL has comparable performance as the conventional SVM MKL algorithms. Moreover, large scale numerical experiments indicate that when cast as semi-infinite programming, LSSVM MKL can be solved more efficiently than SVM MKL. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/l2lssvm.html. Tillmann Falck, Anneleen Daemen, Léon-Charles Tranchevent, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
BMC Bioinform. | 6 |
| 2010 | Gene prioritization and clustering by multi-view text miningabstractBACKGROUND: Text mining has become a useful tool for biologists trying to understand the genetics of diseases. In particular, it can help identify the most interesting candidate genes for a disease for further experimental analysis. Many text mining approaches have been introduced, but the effect of disease-gene identification varies in different text mining models. Thus, the idea of incorporating more text mining models may be beneficial to obtain more refined and accurate knowledge. However, how to effectively combine these models still remains a challenging question in machine learning. In particular, it is a non-trivial issue to guarantee that the integrated model performs better than the best individual model. RESULTS: We present a multi-view approach to retrieve biomedical knowledge using different controlled vocabularies. These controlled vocabularies are selected on the basis of nine well-known bio-ontologies and are applied to index the vast amounts of gene-based free-text information available in the MEDLINE repository. The text mining result specified by a vocabulary is considered as a view and the obtained multiple views are integrated by multi-source learning algorithms. We investigate the effect of integration in two fundamental computational disease gene identification tasks: gene prioritization and gene clustering. The performance of the proposed approach is systematically evaluated and compared on real benchmark data sets. In both tasks, the multi-view approach demonstrates significantly better performance than other comparing methods. CONCLUSIONS: In practical research, the relevance of specific vocabulary pertaining to the task is usually unknown. In such case, multi-view text mining is a superior and promising strategy for text-based disease gene identification. Léon-Charles Tranchevent, Bart De Moor, Yves Moreau |
BMC Bioinform. | 3 |
| 2010 | Weighted hybrid clustering by combining text mining and bibliometrics on a large-scale journal databaseabstractAbstract We propose a new hybrid clustering framework to incorporate text mining with bibliometrics in journal set analysis. The framework integrates two different approaches: clustering ensemble and kernel‐fusion clustering. To improve the flexibility and the efficiency of processing large‐scale data, we propose an information‐based weighting scheme to leverage the effect of multiple data sources in hybrid clustering. Three different algorithms are extended by the proposed weighting scheme and they are employed on a large journal set retrieved from the Web of Science (WoS) database. The clustering performance of the proposed algorithms is systematically evaluated using multiple evaluation methods, and they were cross‐compared with alternative methods. Experimental results demonstrate that the proposed weighted hybrid clustering strategy is superior to other methods in clustering performance and efficiency. The proposed approach also provides a more refined structural mapping of journal sets, which is useful for monitoring and detecting new trends in different scientific fields. Xinhai Liu, Frizo A. L. Janssens, Wolfgang Glänzel, Yves Moreau, Bart De Moor |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2009 | Identifying Customer Profiles in Power Load Time Series Using Spectral Clustering
Carlos Alzate, Marcelo Espinoza, Bart De Moor, Johan A. K. Suykens |
ICANN (2) | 3 |
| 2009 | Robustness of Kernel Based Regression: A Comparison of Iterative Weighting Schemes
Kris De Brabanter, Kristiaan Pelckmans, Jos De Brabanter, Michiel Debruyne, Johan A. K. Suykens, Mia Hubert, Bart De Moor |
ICANN (1) | 7 |
| 2009 | Hybrid Clustering of Text Mining and Bibliometrics Applied to Journal SetsabstractTo obtain correlated and complementary information contained in text mining and bibliometrics, hybrid clustering to incorporate textual content and citation information has become a popular strategy. In this paper, we propose a new computational framework of integrating text mining and bibliometrics to provide a mapping of journal sets. Two different approaches of hybrid clustering methods are applied in this paper. The first category is ensemble clustering, which combines different clustering results obtained from individual data into a consolidated clustering result. The second category is kernel fusion, which maps heterogeneous data sets into the kernel space and combines the kernel matrices for clustering. Kernels can be combined either averagely, or by an optimized weighted linear combination model. In this paper, we propose a novel adaptive kernel K-means clustering algorithm to combine textual content and citation information for clustering. The proposed algorithm is systematically compared with other methods on a clustering problem of 1869 journals published in 2002–2006. Based on several validation indices, the experimental results demonstrate that our hybrid clustering strategy is able to provide clustering result as well as the best individual data source. Xinhai Liu, Yves Moreau, Bart De Moor, Wolfgang Glänzel, Frizo A. L. Janssens |
SDM | 4 |
| 2009 | ViTraM: visualization of transcriptional modulesabstractMOTIVATION: We developed ViTraM, a tool that allows visualizing overlapping transcriptional modules in an intuitive way. By visualizing not only the genes and the experiments in which the genes are co-expressed, but also additional properties of the modules such as the regulators and regulatory motifs that are responsible for the observed co-expression, ViTraM can assist in the biological analysis and interpretation of the output of module detection tools. AVAILABILITY: The ViTraM software is platform-independent. The software and supplementary material are available at: http://homes.esat.kuleuven.be/~kmarchal/ViTraM/Index.html Hong Sun 0002, Karen Lemmens, Tim Van den Bulcke, Kristof Engelen, Bart De Moor, Kathleen Marchal |
Bioinform. | 5 |
| 2009 | An experimental loop design for the detection of constitutional chromosomal aberrations by array CGHabstractBACKGROUND: Comparative genomic hybridization microarrays for the detection of constitutional chromosomal aberrations is the application of microarray technology coming fastest into routine clinical application. Through genotype-phenotype association, it is also an important technique towards the discovery of disease causing genes and genomewide functional annotation in human. When using a two-channel microarray of genomic DNA probes for array CGH, the basic setup consists in hybridizing a patient against a normal reference sample. Two major disadvantages of this setup are (1) the use of half of the resources to measure a (little informative) reference sample and (2) the possibility that deviating signals are caused by benign copy number variation in the "normal" reference instead of a patient aberration. Instead, we apply an experimental loop design that compares three patients in three hybridizations. RESULTS: We develop and compare two statistical methods (linear models of log ratios and mixed models of absolute measurements). In an analysis of 27 patients seen at our genetics center, we observed that the linear models of the log ratios are advantageous over the mixed models of the absolute intensities. CONCLUSION: The loop design and the performance of the statistical analysis contribute to the quick adoption of array CGH as a routine diagnostic tool. They lower the detection limit of mosaicisms and improve the assignment of copy number variation for genetic association studies. Joke Allemeersch, Steven Van Vooren, Femke Hannes, Bart De Moor, Joris Robert Vermeesch, Yves Moreau |
BMC Bioinform. | 4 |
| 2009 | ModuleDigger: an itemset mining framework for the detection of cis-regulatory modulesabstractBACKGROUND: The detection of cis-regulatory modules (CRMs) that mediate transcriptional responses in eukaryotes remains a key challenge in the postgenomic era. A CRM is characterized by a set of co-occurring transcription factor binding sites (TFBS). In silico methods have been developed to search for CRMs by determining the combination of TFBS that are statistically overrepresented in a certain geneset. Most of these methods solve this combinatorial problem by relying on computational intensive optimization methods. As a result their usage is limited to finding CRMs in small datasets (containing a few genes only) and using binding sites for a restricted number of transcription factors (TFs) out of which the optimal module will be selected. RESULTS: We present an itemset mining based strategy for computationally detecting cis-regulatory modules (CRMs) in a set of genes. We tested our method by applying it on a large benchmark data set, derived from a ChIP-Chip analysis and compared its performance with other well known cis-regulatory module detection tools. CONCLUSION: We show that by exploiting the computational efficiency of an itemset mining approach and combining it with a well-designed statistical scoring scheme, we were able to prioritize the biologically valid CRMs in a large set of coregulated genes using binding sites for a large number of potential TFs as input. Hong Sun 0002, Tijl De Bie, Valerie Storms, Qiang Fu 0009, Thomas Dhollander, Karen Lemmens, Annemieke Verstuyf, Bart De Moor, Kathleen Marchal |
BMC Bioinform. | 8 |
| 2009 | Hybrid clustering for validation and improvement of subject-classification schemes
Frizo A. L. Janssens, Lin Zhang 0004, Bart De Moor, Wolfgang Glänzel |
Inf. Process. Manag. | 3 |
| 2009 | Least conservative support and tolerance tubesabstractThis paper studies a distribution-free estimator of the conditional support and tolerance intervals of a distributions underlying a set of paired independent and identically distributed (i.i.d.) observations. The key ingredients are (a) an appropriate notion of risk which measures what probability mass is not captured by the estimate, (b) a uniform concentration inequality for the empirical risk based on a compression argument, and (c) the derivation of a lower bound to the mutual information, dictating how to maximize the informativeness of the estimator. For this result we extend Fano's inequality to the bivariate case. Kristiaan Pelckmans, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Inf. Theory | 4 |
| 2008 | Classification of Sporadic and BRCA1 Ovarian Cancer Based on a Genome-Wide Study of Copy Number Variations
Anneleen Daemen, Olivier Gevaert, Karin Leunen, Vanessa Vanspauwen, Geneviève Michils, Eric Legius, Ignace Vergote, Bart De Moor |
KES (2) | 8 |
| 2008 | Exploring the Operational Characteristics of Inference Algorithms for Transcriptional Networks by Means of Synthetic DataabstractThe development of structure-learning algorithms for gene regulatory networks depends heavily on the availability of synthetic data sets that contain both the original network and associated expression data. This article reports the application of SynTReN, an existing network generator that samples topologies from existing biological networks and uses Michaelis-Menten and Hill enzyme kinetics to simulate gene interactions. We illustrate the effects of different aspects of the expression data on the quality of the inferred network. The tested expression data parameters are network size, network topology, type and degree of noise, quantity of expression data, and interaction types between genes. This is done by applying three well-known inference algorithms to SynTReN data sets. The results show the power of synthetic data in revealing operational characteristics of inference algorithms that are unlikely to be discovered by means of biological microarray data only. Koenraad Van Leemput, Tim Van den Bulcke, Thomas Dhollander, Bart De Moor, Kathleen Marchal, Piet van Remortel |
Artif. Life | 4 |
| 2008 | Multiple-vector user profiles in support of knowledge sharing
Joris Vertommen, Frizo A. L. Janssens, Bart De Moor, Joost R. Duflou |
Inf. Sci. | 3 |
| 2008 | Low rank updated LS-SVM classifiers for fast variable selection
Fabian Ojeda, Johan A. K. Suykens, Bart De Moor |
Neural Networks | 3 |
| 2007 | Convex optimization for the design of learning machines
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ESANN | 3 |
| 2007 | State-of-the-Art and Evolution in Public Data Sets and Competitions for System Identification, Time Series Prediction and Pattern RecognitionabstractIt is the aim of reproducible research to provide mechanisms for objective comparison of methods, algorithms, software and procedures in various research topics. In this paper, we discuss the role of data sets, benchmarks and competitions in the fields of system identification, time series prediction, classification, and pattern recognition in view of creating an environment of reproducible research. Important elements are the data sets, their origin, and the comparison measures that will be used to rank the performance of the methods. The issues are discussed, a comparison is made and recommendations are given. Joos Vandewalle, Johan A. K. Suykens, Bart De Moor, Amaury Lendasse |
ICASSP (4) | 3 |
| 2007 | Variable selection by rank-one updates for least squares support vector machinesabstractLeast squares support vector machines (LS-SVM) classifiers are a class of simple, yet powerful, kernel methods whose solution follows from a set of linear equations. Here, forward and backward algorithms, based on this technique, are proposed for fast and efficient variable selection. By exploiting the structure of the LS-SVM solution a closed form expression for the leave-one-out (LOO) estimator, useful for selecting variables, is obtained. For inclusion or removal of a new variable, rank-one adjustments in the kernel matrix (linear kernel) allow for updating, rather than recomputing, the LS-SVM solution. The proposed approach is applied to microarray data for gene selection. Simulations clearly show lower computational complexity along with good stability on the generalization performance when compared to other related algorithms. Fabian Ojeda, Johan A. K. Suykens, Bart De Moor |
IJCNN | 3 |
| 2007 | Dynamic hybrid clustering of bioinformatics by incorporating text mining and citation analysisabstractTo unravel the concept structure and dynamics of the bioinformatics field, we analyze a set of 7401 publications from the Web of Science and MEDLINE databases, publication years 1981–2004. For delineating this complex, interdisciplinary field, a novel bibliometric retrieval strategy is used. Given that the performance of unsupervised clustering and classification of scientific publications is significantly improved by deeply merging textual contents with the structure of the citation graph, we proceed with a hybrid clustering method based on Fisher’s inverse chi-square. The optimal number of clusters is determined by a compound semi-automatic strategy comprising a combination of distance-based and stability-based methods. We also investigate the relationship between number of Latent Semantic Indexing factors, number of clusters, and clustering performance. The HITS and PageRank algorithms are used to determine representative publications in each cluster. Next, we develop a methodology for dynamic hybrid clustering of evolving bibliographic data sets. The same clustering methodology is applied to consecutive periods defined by time windows on the set, and in a subsequent phase chains are formed by matching and tracking clusters through time. Term networks for the eleven resulting cluster chains present the cognitive structure of the field. Finally, we provide a view on how much attention the bioinformatics community has devoted to the different subfields through time. Frizo A. L. Janssens, Wolfgang Glänzel, Bart De Moor |
KDD | 3 |
| 2007 | A Risk Minimization Principle for a Class of Parzen EstimatorsabstractThis paper explores the use of a Maximal Average Margin (MAM) optimality principle for the design of learning algorithms. It is shown that the application of this risk minimization principle results in a class of (computationally) simple learning machines similar to the classical Parzen window classifier. A direct relation with the Rademacher complexities is established, as such facilitating analysis and providing a notion of certainty of prediction. This analysis is related to Support Vector Machines by means of a margin transformation. The power of the MAM principle is illustrated further by application to ordinal regression tasks, resulting in an $O(n)$ algorithm able to process large datasets in reasonable time. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
NIPS | 3 |
| 2007 | Query-driven module discovery in microarray dataabstractAbstract Motivation: Existing (bi)clustering methods for microarray data analysis often do not answer the specific questions of interest to a biologist. Such specific questions could be derived from other information sources, including expert prior knowledge. More specifically, given a set of seed genes which are believed to have a common function, we would like to recruit genes with similar expression profiles as the seed genes in a significant subset of experimental conditions. Results: We introduce QDB, a novel Bayesian query-driven biclustering framework in which the prior distributions allow introducing knowledge from a set of seed genes (query) to guide the pattern search. In two well-known yeast compendia, we grow highly functionally enriched biclusters from small sets of seed genes using a resolution sweep approach. In addition, relevant conditions are identified and modularity of the biclusters is demonstrated, including the discovery of overlapping modules. Finally, our method deals with missing values naturally, performs well on artificial data from a recent biclustering benchmark study and has a number of conceptual advantages when compared to existing approaches for focused module search. Availability: Software is available on the Supplementary Material. Contact: [email protected] Supplementary information: Available on http://homes.esat.kuleuven.be/~tdhollan/Supplementary_Information_Dhollander_2007/index.html Thomas Dhollander, Qizheng Sheng, Karen Lemmens, Bart De Moor, Kathleen Marchal, Yves Moreau |
Bioinform. | 4 |
| 2007 | CALIB: a Bioconductor package for estimating absolute expression levels from two-color microarray dataabstractUNLABELLED: In this article we describe a new Bioconductor package 'CALIB' for normalization of two-color microarray data. This approach is based on the measurements of external controls and estimates an absolute target level for each gene and condition pair, as opposed to working with log-ratios as a relative measure of expression. Moreover, this method makes no assumptions regarding the distribution of gene expression divergence. AVAILABILITY: http://bioconductor.org/packages/2.0/bioc Open Source. Kristof Engelen, Bart De Moor, Kathleen Marchal |
Bioinform. | 3 |
| 2007 | Efficiently updating and tracking the dominant kernel principal components
Luc Hoegaerts, Lieven De Lathauwer, Ivan Goethals, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
Neural Networks | 6 |
| 2007 | A Convex Approach to Validation-Based Learning of the Regularization ConstantabstractThis letter investigates a tight convex relaxation to the problem of tuning the regularization constant with respect to a validation based criterion. A number of algorithms is covered including ridge regression, regularization networks, smoothing splines, and least squares support vector machines (LS-SVMs) for regression. This convex approach allows the application of reliable and efficient tools, thereby improving computational cost and automatization of the learning method. It is shown that all solutions of the relaxation allow an interpretation in terms of a solution to a weighted LS-SVM. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Neural Networks | 3 |
| 2006 | Region-Based Statistical Background Modeling for Foreground Object SegmentationabstractThis paper proposes a novel region-based scheme for dynamically modeling time-evolving statistics of video background, leading to an effective segmentation of foreground moving objects for a video surveillance system. In (L. Li et al., 2004) statistical-based video surveillance systems employ a Bayes decision rule for classifying foreground and background changes in individual pixels. Although principal feature representations significantly reduce the size of tables of statistics, pixel-wise maintenance remains a challenge due to the computations and memory requirement. The proposed region-based scheme, which is an extension of the above method, replaces pixel-based statistics by region-based statistics through introducing dynamic background region (or pixel) merging and splitting. Simulations have been performed to several outdoor and indoor image sequences, and results have shown a significant reduction of memory requirements for tables of statistics while maintaining relatively good quality in foreground segmented video objects. Kristof Op De Beeck, Irene Y. H. Gu, Liyuan Li, Mats Viberg, Bart De Moor |
ICIP | 5 |
| 2006 | Interpreting Gene Profiles from Biomedical Literature Mining with Self Organizing Maps
Steven Van Vooren, Bert Coessens, Bart De Moor |
ISNN (2) | 4 |
| 2006 | A calibration method for estimating absolute expression levels from microarray dataabstractMOTIVATION: We describe an approach to normalize spotted microarray data, based on a physically motivated calibration model. This model consists of two major components, describing the hybridization of target transcripts to their corresponding probes on the one hand, and the measurement of fluorescence from the hybridized, labeled target on the other hand. The model parameters and error distributions are estimated from external control spikes. RESULTS: Using a publicly available dataset, we show that our procedure is capable of adequately removing the typical non-linearities of the data, without making any assumptions on the distribution of differences in gene expression from one biological sample to the next. Since our model links target concentration to measured intensity, we show how absolute expression values of target transcripts in the hybridization solution can be estimated up to a certain degree. Kristof Engelen, Bart Naudts, Bart De Moor, Kathleen Marchal |
Bioinform. | 3 |
| 2006 | SynTReN: a generator of synthetic gene expression data for design and analysis of structure learning algorithmsabstractBACKGROUND: The development of algorithms to infer the structure of gene regulatory networks based on expression data is an important subject in bioinformatics research. Validation of these algorithms requires benchmark data sets for which the underlying network is known. Since experimental data sets of the appropriate size and design are usually not available, there is a clear need to generate well-characterized synthetic data sets that allow thorough testing of learning algorithms in a fast and reproducible manner. RESULTS: In this paper we describe a network generator that creates synthetic transcriptional regulatory networks and produces simulated gene expression data that approximates experimental data. Network topologies are generated by selecting subnetworks from previously described regulatory networks. Interaction kinetics are modeled by equations based on Michaelis-Menten and Hill kinetics. Our results show that the statistical properties of these topologies more closely approximate those of genuine biological networks than do those of different types of random graph models. Several user-definable parameters adjust the complexity of the resulting data set with respect to the structure learning algorithms. CONCLUSION: This network generation technique offers a valid alternative to existing methods. The topological characteristics of the generated networks more closely resemble the characteristics of real transcriptional networks. Simulation of the network scales well to large networks. The generator models different types of biological interactions and produces biologically plausible synthetic gene expression data. Tim Van den Bulcke, Koenraad Van Leemput, Bart Naudts, Piet van Remortel, Hongwu Ma, Alain Verschoren, Bart De Moor, Kathleen Marchal |
BMC Bioinform. | 7 |
| 2006 | More robust detection of motifs in coexpressed genes by using phylogenetic informationabstractBACKGROUND: Several motif detection algorithms have been developed to discover overrepresented motifs in sets of coexpressed genes. However, in a noisy gene list, the number of genes containing the motif versus the number lacking the motif might not be sufficiently high to allow detection by classical motif detection tools. To still recover motifs which are not significantly enriched but still present, we developed a procedure in which we use phylogenetic footprinting to first delineate all potential motifs in each gene. Then we mutually compare all detected motifs and identify the ones that are shared by at least a few genes in the data set as potential candidates. RESULTS: We applied our methodology to a compiled test data set containing known regulatory motifs and to two biological data sets derived from genome wide expression studies. By executing four consecutive steps of 1) identifying conserved regions in orthologous intergenic regions, 2) aligning these conserved regions, 3) clustering the conserved regions containing similar regulatory regions followed by extraction of the regulatory motifs and 4) screening the input intergenic sequences with detected regulatory motif models, our methodology proves to be a powerful tool for detecting regulatory motifs when a low signal to noise ratio is present in the input data set. Comparing our results with two other motif detection algorithms points out the robustness of our algorithm. CONCLUSION: We developed an approach that can reliably identify multiple regulatory motifs lacking a high degree of overrepresentation in a set of coexpressed genes (motifs belonging to sparsely connected hubs in the regulatory network) by exploiting the advantages of using both coexpression and phylogenetic information. Pieter Monsieurs, Gert Thijs, Abeer A. Fadda, Sigrid C. J. De Keersmaecker, Jozef Vanderleyden, Bart De Moor, Kathleen Marchal |
BMC Bioinform. | 6 |
| 2006 | Towards mapping library and information science
Frizo A. L. Janssens, Jacqueline Leta, Wolfgang Glänzel, Bart De Moor |
Inf. Process. Manag. | 4 |
| 2006 | Additive Regularization Trade-Off: Fusion of Training and Validation Levels in Kernel Methods
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
Mach. Learn. | 3 |
| 2005 | Componentwise Support Vector Machines for Structure Detection
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ICANN (2) | 3 |
| 2005 | Maximal variation and missing values for componentwise support vector machinesabstractThis paper proposes primal-dual kernel machine classifiers based on worst-case analysis of a finite set of observations including missing values of the inputs. Key ingredients are the use of a componentwise support vector machine (cSVM) and an empirical measure of maximal variation of the components to bind the influence of the component which cannot be evaluated due to missing values. A regularization term based on the L/sub 1/ norm of the maximal variation is used to obtain a mechanism for structure detection in that context. An efficient implementation using the hierarchical kernel machines framework is elaborated. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor, Jos De Brabanter |
IJCNN | 3 |
| 2005 | BioMart and Bioconductor: a powerful link between biological databases and microarray data analysisabstractbiomaRt is a new Bioconductor package that integrates BioMart data resources with data analysis software in Bioconductor. It can annotate a wide range of gene or gene product identifiers (e.g. Entrez-Gene and Affymetrix probe identifiers) with information such as gene symbol, chromosomal coordinates, Gene Ontology and OMIM annotation. Furthermore biomaRt enables retrieval of genomic sequences and single nucleotide polymorphism information, which can be used in data analysis. Fast and up-to-date data retrieval is possible as the package executes direct SQL queries to the BioMart databases (e.g. Ensembl). The biomaRt package provides a tight integration of large, public or locally installed BioMart databases with data analysis in Bioconductor creating a powerful environment for biological data mining. Steffen Durinck, Yves Moreau, Arek Kasprzyk, Sean R. Davis, Bart De Moor, Alvis Brazma, Wolfgang Huber |
Bioinform. | 5 |
| 2005 | M@CBETH: a microarray classification benchmarking toolabstractMicroarray classification can be useful to support clinical management decisions for individual patients in, for example, oncology. However, comparing classifiers and selecting the best for each microarray dataset can be a tedious and non-straightforward task. The M@CBETH (a MicroArray Classification BEnchmarking Tool on a Host server) web service offers the microarray community a simple tool for making optimal two-class predictions. M@CBETH aims at finding the best prediction among different classification methods by using randomizations of the benchmarking dataset. The M@CBETH web service intends to introduce an optimal use of clinical microarray data classification. Nathalie Pochet, Frizo A. L. Janssens, Frank De Smet, Kathleen Marchal, Johan A. K. Suykens, Bart De Moor |
Bioinform. | 6 |
| 2005 | arrayCGHbase: an analysis platform for comparative genomic hybridization microarraysabstractBACKGROUND: The availability of the human genome sequence as well as the large number of physically accessible oligonucleotides, cDNA, and BAC clones across the entire genome has triggered and accelerated the use of several platforms for analysis of DNA copy number changes, amongst others microarray comparative genomic hybridization (arrayCGH). One of the challenges inherent to this new technology is the management and analysis of large numbers of data points generated in each individual experiment. RESULTS: We have developed arrayCGHbase, a comprehensive analysis platform for arrayCGH experiments consisting of a MIAME (Minimal Information About a Microarray Experiment) supportive database using MySQL underlying a data mining web tool, to store, analyze, interpret, compare, and visualize arrayCGH results in a uniform and user-friendly format. Following its flexible design, arrayCGHbase is compatible with all existing and forthcoming arrayCGH platforms. Data can be exported in a multitude of formats, including BED files to map copy number information on the genome using the Ensembl or UCSC genome browser. CONCLUSION: ArrayCGHbase is a web based and platform independent arrayCGH data analysis tool, that allows users to access the analysis suite through the internet or a local intranet after installation on a private server. ArrayCGHbase is available at http://medgen.ugent.be/arrayCGHbase/. Björn Menten, Filip Pattyn, Katleen De Preter, Piet Robbrecht, Evi Michels, Karen Buysse, Geert Mortier, Anne De Paepe, Steven Van Vooren, Joris Robert Vermeesch, Yves Moreau, Bart De Moor, Stefan Vermeulen, Frank Speleman, Jo Vandesompele |
BMC Bioinform. | 12 |
| 2005 | Subset based least squares subspace regression in RKHS
Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
Neurocomputing | 4 |
| 2005 | The differogram: Non-parametric noise variance estimation and its use for model selection
Kristiaan Pelckmans, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Neurocomputing | 4 |
| 2005 | Building sparse representations and structure determination on LS-SVM substrates
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
Neurocomputing | 3 |
| 2005 | Combining full text and bibliometric information in mapping scientific disciplines
Patrick Glenisson, Wolfgang Glänzel, Frizo A. L. Janssens, Bart De Moor |
Inf. Process. Manag. | 4 |
| 2005 | Handling missing values in support vector machine classifiers
Kristiaan Pelckmans, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Neural Networks | 4 |
| 2005 | Primal-Dual Monotone Kernel Regression
Kristiaan Pelckmans, Marcelo Espinoza, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Neural Process. Lett. | 5 |
| 2004 | Sparse LS-SVMs using additive regularization with a penalized validation criterion
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ESANN | 3 |
| 2004 | Nonparametric regularized time delay estimationabstractA novel nonparametric estimator, which we call the regularized time delay estimator (RTDE), is introduced for the time delay estimation problem in MISO (multiple input-single output) linear systems. This estimator decouples the input-output signals in the frequency domain using a regularized Wiener-Hopf filter. An RLS (regularized least squares) problem is formulated based on the coherence spectrum in order to find the optimal filter that can decouple the signals. Then, the corresponding delayed impulse responses of the system are computed. As a result, the time delay between the several input-output signals can be estimated. Oscar Barrero, Bart De Moor |
ICASSP (2) | 2 |
| 2004 | A Comparison of Pruning Algorithms for Sparse Least Squares Support Vector Machines
Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
ICONIP | 4 |
| 2004 | Morozov, Ivanov and Tikhonov Regularization Based LS-SVMs
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ICONIP | 3 |
| 2004 | Primal space sparse kernel partial least squares regression for large scale problemsabstractKernel based methods suffer from exceeding time and memory requirements when applied on large datasets since the involved optimization problems typically scale polynomially in the number of data samples. As a remedy we propose both working on a reduced set (for fast evaluation) and at the same time keeping the number of model parameters small (for fast training). Departing from the Nystrom based feature approximation we describe fixed-size least squares support vector machine in the context of primal space least squares regression, to extend it with a supervised counterpart, sparse kernel partial least squares. The model is illustrated on a large scale example. Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
IJCNN | 4 |
| 2004 | Regularization constants in LS-SVMs: a fast estimate via convex optimizationabstractThe tuning of the regularization constant in applications of least squares support vector machines (LS-SVMs) for regression and classification is considered. The formulation of the LS-SVM training and regularization constant tuning problem (w.r.t. the validation performance) is considered as a single constrained optimization problem. In the formulation with Tikhonov regularization the problem of estimation the weights, validation errors and the regularization constants is a non-convex problem. The main result of This work is a conversion of the nonlinear constraints into a set of linear constraints, which turns the problem into a convex one. This is done based upon a simple Nadaraya-Watson kernel estimator via approximating the LS-SVM smoother matrix by the Nadaraya-Watson smoother. The paper further illustrates how to use this initial estimate towards grid search or local search methods. Numerical examples show considerable speed-ups by the proposed method. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
IJCNN | 3 |
| 2004 | Using literature and data to learn Bayesian networks as clinical models of ovarian tumors
Péter Antal, Geert Fannes, Dirk Timmerman, Yves Moreau, Bart De Moor |
Artif. Intell. Medicine | 5 |
| 2004 | A genetic algorithm for the detection of new cis-regulatory modules in sets of coregulated genesabstractSUMMARY: The implementation of a genetic algorithm is described that provides a fast method of searching for the optimal combination of transcription factor binding sites in a set of regulatory sequences. AVAILABILITY: The algorithm can be used transparently as a web service from within the Toucan software. Toucan can be accessed at http://www.esat.kuleuven.ac.be/~saerts/software/toucan.php. A standalone version of the software is available upon request. Stein Aerts, Peter Van Loo, Yves Moreau, Bart De Moor |
Bioinform. | 4 |
| 2004 | Importing MAGE-ML format microarray data into BioConductorabstractUNLABELLED: The microarray gene expression markup language (MAGE-ML) is a widely used XML (eXtensible Markup Language) standard for describing and exchanging information about microarray experiments. It can describe microarray designs, microarray experiment designs, gene expression data and data analysis results. We describe RMAGEML, a new Bioconductor package that provides a link between cDNA microarray data stored in MAGE-ML format and the Bioconductor framework for preprocessing, visualization and analysis of microarray experiments. AVAILABILITY: http://www.bioconductor.org. Open Source. Steffen Durinck, Joke Allemeersch, Vincent Carey, Yves Moreau, Bart De Moor |
Bioinform. | 5 |
| 2004 | Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reductionabstractMOTIVATION: Microarrays are capable of determining the expression levels of thousands of genes simultaneously. In combination with classification methods, this technology can be useful to support clinical management decisions for individual patients, e.g. in oncology. The aim of this paper is to systematically benchmark the role of non-linear versus linear techniques and dimensionality reduction methods. RESULTS: A systematic benchmarking study is performed by comparing linear versions of standard classification and dimensionality reduction techniques with their non-linear versions based on non-linear kernel functions with a radial basis function (RBF) kernel. A total of 9 binary cancer classification problems, derived from 7 publicly available microarray datasets, and 20 randomizations of each problem are examined. CONCLUSIONS: Three main conclusions can be formulated based on the performances on independent test sets. (1) When performing classification with least squares support vector machines (LS-SVMs) (without dimensionality reduction), RBF kernels can be used without risking too much overfitting. The results obtained with well-tuned RBF kernels are never worse and sometimes even statistically significantly better compared to results obtained with a linear kernel in terms of test set receiver operating characteristic and test set accuracy performances. (2) Even for classification with linear classifiers like LS-SVM with linear kernel, using regularization is very important. (3) When performing kernel principal component analysis (kernel PCA) before classification, using an RBF kernel for kernel PCA tends to result in overfitting, especially when using supervised feature selection. It has been observed that an optimal selection of a large number of features is often an indication for overfitting. Kernel PCA with linear kernel gives better results. Nathalie Pochet, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
Bioinform. | 4 |
| 2004 | Benchmarking Least Squares Support Vector Machine ClassifiersabstractIn Support Vector Machines (SVMs), the solution of the classification problem is characterized by a (convex) quadratic programming (QP) problem. In a modified version of SVMs, called Least Squares SVM classifiers (LS-SVMs), a least squares cost function is proposed so as to obtain a linear set of equations in the dual space. While the SVM classifier has a large margin interpretation, the LS-SVM formulation is related in this paper to a ridge regression approach for classification with binary targets and to Fisher's linear discriminant analysis in the feature space. Multiclass categorization problems are represented by a set of binary classifiers using different output coding schemes. While regularization is used to control the effective number of parameters of the LS-SVM classifier, the sparseness property of SVMs is lost due to the choice of the 2-norm. Sparseness can be imposed in a second stage by gradually pruning the support value spectrum and optimizing the hyperparameters during the sparse approximation procedure. In this paper, twenty public domain benchmark datasets are used to evaluate the test set performance of LS-SVM classifiers with linear, polynomial and radial basis function (RBF) kernels. Both the SVM and LS-SVM classifier with RBF kernel in combination with standard cross-validation procedures for hyperparameter selection achieve comparable test set performances. These SVM and LS-SVM performances are consistently very good when compared to a variety of methods described in the literature including decision tree based algorithms, statistical algorithms and instance based learning methods. We show on ten UCI datasets that the LS-SVM sparse approximation procedure can be successfully applied. Tony Van Gestel, Johan A. K. Suykens, Bart Baesens, Stijn Viaene, Jan Vanthienen, Guido Dedene, Bart De Moor, Joos Vandewalle |
Mach. Learn. | 7 |
| 2003 | Bankruptcy prediction with least squares support vector machine classifiersabstractClassification algorithms like linear discriminant analysis and logistic regression are popular linear techniques for modelling and predicting corporate distress. These techniques aim at finding an optimal linear combination of explanatory input variables, such as, e.g., solvency and liquidity ratios, in order to analyse, model and predict corporate default risk. Recently, performant kernel based nonlinear classification techniques, like support vector machines, least squares support vector machines and kernel fisher discriminant analysis, have been developed. Basically, these methods map the inputs first in a nonlinear way to a high dimensional kernel-induced feature space, in which a linear classifier is constructed in the second step. Practical expressions are obtained in the so-called dual space by application of Mercer's theorem. In this paper, we explain the relations between linear and nonlinear kernel based classification and illustrate their performance on predicting bankruptcy of mid-cap firms in Belgium and the Netherlands. Tony Van Gestel, Bart Baesens, Johan A. K. Suykens, Marcelo Espinoza, Dirk-Emma Baestaens, Jan Vanthienen, Bart De Moor |
CIFEr | 7 |
| 2003 | Kernel PLS variants for regression
Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
ESANN | 4 |
| 2003 | Bayesian applications of belief networks and multilayer perceptrons for ovarian tumor classification with rejection
Péter Antal, Geert Fannes, Dirk Timmerman, Yves Moreau, Bart De Moor |
Artif. Intell. Medicine | 5 |
| 2003 | MARAN: Normalizing Micro-array DataabstractAbstract Summary: MARAN is a web-based application for normalizing microarray data. MARAN comprises a generic ANOVA model, an option for Loess fitting prior to ANOVA analysis, and a module for selecting genes with significantly changing expression. Availability: http://www.esat.kuleuven.ac.be/maran/ Contact: [email protected] * To whom correspondence should be addressed. Kristof Engelen, Bert Coessens, Kathleen Marchal, Bart De Moor |
Bioinform. | 4 |
| 2003 | On a cepstral norm for an ARMA model and the polar plot of the logarithm of its transfer function
Katrien De Cock, Bernard Hanzon, Bart De Moor |
Signal Process. | 3 |
| 2003 | Constraints in channel shortening equalizer design for DMT-based systems
Geert Ysebaert, Katleen Van Acker, Marc Moonen, Bart De Moor |
Signal Process. | 4 |
| 2003 | A formant filtered physical model for wind instrumentsabstractWe report on our research concerning the calibration of physical models for sound synthesis. We combine waveguide physical modeling synthesis with formant filtering, by dividing the nonlinear description of the reed mechanism into a nonlinear part and an input-dependent linear filter. We elaborate on the calibration of the model and assess its performance by comparing it to a single-reed, cylindrical bore instrument, the clarinet. Axel Nackaerts, Bart De Moor, Rudy Lauwereins |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | A support vector machine formulation to PCA analysis and its kernel versionabstractIn this paper, we present a simple and straightforward primal-dual support vector machine formulation to the problem of principal component analysis (PCA) in dual variables. By considering a mapping to a high-dimensional feature space and application of the kernel trick (Mercer theorem), kernel PCA is obtained as introduced by Scholkopf et al. (2002). While least squares support vector machine classifiers have a natural link with the kernel Fisher discriminant analysis (minimizing the within class scatter around targets +1 and -1), for PCA analysis one can take the interpretation of a one-class modeling problem with zero target value around which one maximizes the variance. The score variables are interpreted as error variables within the problem formulation. In this way primal-dual constrained optimization problem interpretations to the linear and kernel PCA analysis are obtained in a similar style as for least square-support vector machine classifiers. Johan A. K. Suykens, Tony Van Gestel, Joos Vandewalle, Bart De Moor |
IEEE Trans. Neural Networks | 4 |
| 2002 | Web-based Data Collection for Uterine Adnexal Tumors: A Case StudyabstractWe have developed a World Wide Web application for the collection of EPRs (electronic patient records) from uterine adnexal masses pre-operatively examined with transvaginal ultrasonography. The application has been used intensively since November 2000 by nine of the 19 international centers that joined the International Ovarian Tumor Analysis (IOTA) consortium. The IOTA database contains 68 parameters for 1,150 masses. We report the design and implementation of the generic Web-based clinical data entry system and describe the advantages and drawbacks that we have experienced while developing, using and maintaining the system. The data model, the user interface, the help system, the constraints (mandatory/optional) and the quality checking were all based on the medical protocol created by the IOTA consortium. The data collection system has become an open and transparent implementation of the formalized protocol. It covers the complete path of the patient data from the clinical situation to the finalized database. This approach provides new types of possibilities for the data analysis, since all aspects of the data collection are documented and formally available to the data analyst. The IOTA Web site can be found at, which also serves as the entry point for the secure EPR application. Stein Aerts, Péter Antal, Dirk Timmerman, Bart De Moor, Yves Moreau |
CBMS | 4 |
| 2002 | Compactly Supported RBF Kernels for Sparsifying the Gram Matrix in LS-SVM Regression Models
Bart Hamers, Johan A. K. Suykens, Bart De Moor |
ICANN | 3 |
| 2002 | Adaptive quality-based clustering of gene expression profilesabstractMOTIVATION: Microarray experiments generate a considerable amount of data, which analyzed properly help us gain a huge amount of biologically relevant information about the global cellular behaviour. Clustering (grouping genes with similar expression profiles) is one of the first steps in data analysis of high-throughput expression measurements. A number of clustering algorithms have proved useful to make sense of such data. These classical algorithms, though useful, suffer from several drawbacks (e.g. they require the predefinition of arbitrary parameters like the number of clusters; they force every gene into a cluster despite a low correlation with other cluster members). In the following we describe a novel adaptive quality-based clustering algorithm that tackles some of these drawbacks. RESULTS: We propose a heuristic iterative two-step algorithm: First, we find in the high-dimensional representation of the data a sphere where the "density" of expression profiles is locally maximal (based on a preliminary estimate of the radius of the cluster-quality-based approach). In a second step, we derive an optimal radius of the cluster (adaptive approach) so that only the significantly coexpressed genes are included in the cluster. This estimation is achieved by fitting a model to the data using an EM-algorithm. By inferring the radius from the data itself, the biologist is freed from finding an optimal value for this radius by trial-and-error. The computational complexity of this method is approximately linear in the number of gene expression profiles in the data set. Finally, our method is successfully validated using existing data sets. AVAILABILITY: http://www.esat.kuleuven.ac.be/~thijs/Work/Clustering.html Frank De Smet, Janick Mathys, Kathleen Marchal, Gert Thijs, Bart De Moor, Yves Moreau |
Bioinform. | 5 |
| 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif SamplingabstractAbstract Summary: INCLUSive allows automatic multistep analysis of microarray data (clustering and motif finding). The clustering algorithm (adaptive quality-based clustering) groups together genes with highly similar expression profiles. The upstream sequences of the genes belonging to a cluster are automatically retrieved from GenBank and can be fed directly into Motif Sampler, a Gibbs sampling algorithm that retrieves statistically over-represented motifs in sets of sequences, in this case upstream regions of co-expressed genes. Availability: For academic purposes at http://www.esat.kuleuven.ac.be/~dna/BioI/Software.html Contact: [email protected] * To whom correspondence should be addressed. Email: [email protected]. Gert Thijs, Yves Moreau, Frank De Smet, Janick Mathys, Magali Lescot, Stephane Rombauts, Pierre Rouzé, Bart De Moor, Kathleen Marchal |
Bioinform. | 8 |
| 2002 | Bayesian Framework for Least-Squares Support Vector Machine Classifiers, Gaussian Processes, and Kernel Fisher Discriminant AnalysisabstractThe Bayesian evidence framework has been successfully applied to the design of multilayer perceptrons (MLPs) in the work of MacKay. Nevertheless, the training of MLPs suffers from drawbacks like the nonconvex optimization problem and the choice of the number of hidden units. In support vector machines (SVMs) for classification, as introduced by Vapnik, a nonlinear decision boundary is obtained by mapping the input vector first in a nonlinear way to a high-dimensional kernel-induced feature space in which a linear large margin classifier is constructed. Practical expressions are formulated in the dual space in terms of the related kernel function, and the solution follows from a (convex) quadratic programming (QP) problem. In least-squares SVMs (LS-SVMs), the SVM problem formulation is modified by introducing a least-squares cost function and equality instead of inequality constraints, and the solution follows from a linear system in the dual space. Implicitly, the least-squares formulation corresponds to a regression formulation and is also related to kernel Fisher discriminant analysis. The least-squares regression formulation has advantages for deriving analytic expressions in a Bayesian evidence framework, in contrast to the classification formulations used, for example, in gaussian processes (GPs). The LS-SVM formulation has clear primal-dual interpretations, and without the bias term, one explicitly constructs a model that yields the same expressions as have been obtained with GPs for regression. In this article, the Bayesian evidence framework is combined with the LS-SVM classifier formulation. Starting from the feature space formulation, analytic expressions are obtained in the dual space on the different levels of Bayesian inference, while posterior class probabilities are obtained by marginalizing over the model parameters. Empirical results obtained on 10 public domain data sets show that the LS-SVM classifier designed within the Bayesian evidence framework consistently yields good generalization performances. Tony Van Gestel, Johan A. K. Suykens, Gert R. G. Lanckriet, Annemie Lambrechts, Bart De Moor, Joos Vandewalle |
Neural Comput. | 5 |
| 2002 | Multiclass LS SVMs Moderated Outputs and Coding Decoding Schemes
Tony Van Gestel, Johan A. K. Suykens, Gert R. G. Lanckriet, Annemie Lambrechts, Bart De Moor, Joos Vandewalle |
Neural Process. Lett. | 5 |
| 2002 | Functional bioinformatics of microarray data: from expression to regulationabstractUsing microarrays is a powerful technique to monitor the expression of thousands of genes in a single experiment. From series of such experiments, it is possible to identify the mechanisms that govern the activation of genes in an organism. Short deoxyribonucleic acid patterns (called binding sites) near the genes serve as switches that control gene expression. As a result similar patterns of expression can correspond to similar binding site patterns. Here we integrate clustering of coexpressed genes with the discovery of binding motifs. We overview several important clustering techniques and present a clustering algorithm (called adaptive quality-based clustering), which we have developed to address several shortcomings of existing methods. We overview the different techniques for motif finding, in particular the technique of Gibbs sampling, and we present several extensions of this technique in our Motif Sampler. Finally, we present an integrated web tool called INCLUSive (available online at http://www.esat.kuleuven.ac.be//spl sim/dna/BioI/Software.html) that allows the easy analysis of microarray data for motif finding. Yves Moreau, Frank De Smet, Gert Thijs, Kathleen Marchal, Bart De Moor |
Proc. IEEE | 5 |
| 2001 | Extended Bayesian Regression Models: A Symbiotic Application of Belief Networks and Multilayer Perceptrons for the Classification of Ovarian Tumors
Péter Antal, Geert Fannes, Bart De Moor, Joos Vandewalle, Yves Moreau, Dirk Timmerman |
AIME | 3 |
| 2001 | Annotated Bayesian Networks: A Tool to Integrate Textual and Probabilistic Medical KnowledgeabstractWe have previously (2000) reported on the development of Bayesian network models for the pre-operative discrimination between malignant and benign ovarian masses. The models incorporated both medical background knowledge and patient data, which required the traceability of the incorporated prior medical knowledge. For this purpose, we followed a particular annotation method for Bayesian networks using a dedicated representation. In this paper, we present the resulting annotated Bayesian network (ABN) representation that consists of a regular Bayesian network, with standard probabilistic semantics, and a corresponding semantic network, to which textual information sources are attached. We demonstrate the applicability of such a dual model to represent both the rigorous probabilistic and the unconstrained textual medical knowledge. We describe methods on how these ABN models can be used: (1) as a domain model to arrange the personal textual information of a clinician according to the semantics of the domain, (2) in decision support to provide detailed (and even personalized) explanation, and (3) to enhance the information retrieval to find new textual information more efficiently. Péter Antal, Bart De Moor, Tamás Mészáros 0003, Tadeusz P. Dobrowiecki |
CBMS | 2 |
| 2001 | Automatic relevance determination for Least Squares Support Vector Machines classifiers
Tony Van Gestel, Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
ESANN | 3 |
| 2001 | Kernel Canonical Correlation Analysis and Least Squares Support Vector Machines
Tony Van Gestel, Johan A. K. Suykens, Jos De Brabanter, Bart De Moor, Joos Vandewalle |
ICANN | 4 |
| 2001 | A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genesabstractMicroarray experiments can reveal useful information on the transcriptional regulation. We try to find regulatory elements in the region upstream of translation start of coexpressed genes. Here we present a modification to the original Gibbs Sampling algorithm [12]. We introduce a probability distribution to estimate the number of copies of the motif in a sequence. The second modification is the incorporation of a higher-order background model. We have successfully tested our algorithm on several data sets. First we show results on two selected data set: sequences from plants containing the G-box motif and the upstream sequences from bacterial genes regulated by O2-responsive protein FNR. In both cases the motif sampler is able to find the expected motifs. Finally, the sampler is tested on 4 clusters of coexpressed genes from a wounding experiment in Arabidopsis thaliana. We find several putative motifs that are related to the pathways involved in the plant defense mechanism. Gert Thijs, Kathleen Marchal, Magali Lescot, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau |
RECOMB | 5 |
| 2001 | A higher-order background model improves the detection of promoter regulatory elements by Gibbs samplingabstractMOTIVATION: Transcriptome analysis allows detection and clustering of genes that are coexpressed under various biological circumstances. Under the assumption that coregulated genes share cis-acting regulatory elements, it is important to investigate the upstream sequences controlling the transcription of these genes. To improve the robustness of the Gibbs sampling algorithm to noisy data sets we propose an extension of this algorithm for motif finding with a higher-order background model. RESULTS: Simulated data and real biological data sets with well-described regulatory elements are used to test the influence of the different background models on the performance of the motif detection algorithm. We show that the use of a higher-order model considerably enhances the performance of our motif finding algorithm in the presence of noisy data. For Arabidopsis thaliana, a reliable background model based on a set of carefully selected intergenic sequences was constructed. AVAILABILITY: Our implementation of the Gibbs sampler called the Motif Sampler can be used through a web interface: http://www.esat.kuleuven.ac.be/~thijs/Work/MotifSampler.html. CONTACT: [email protected]; [email protected] Gert Thijs, Magali Lescot, Kathleen Marchal, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau |
Bioinform. | 5 |
| 2001 | Knowledge discovery in a direct marketing case using least squares support vector machinesabstractWe study the problem of repeat-purchase modeling in a direct marketing setting using Belgian data. More specifically, we investigate the detection and qualification of the most relevant explanatory variables for predicting purchase incidence. The analysis is based on a wrapped form of input selection using a sensitivity based pruning heuristic to guide a greedy, stepwise, and backward traversal of the input space. For this purpose, we make use of a powerful and promising least squares support vector machine (LS-SVM) classifier formulation. This study extends beyond the standard recency frequency monetary (RFM) modeling semantics in two ways: (1) by including alternative operationalizations of the RFM variables, and (2) by adding several other (non-RFM) predictors. Results indicate that elimination of redundant/irrelevant inputs allows significant reduction of model complexity. The empirical findings also highlight the importance of frequency and monetary variables, while the recency variable category seems to be of somewhat lesser importance to the case at hand. Results also point to the added value of including non-RFM variables for improving customer profiling. More specifically, customer/company interaction, measured using indicators of information requests and complaints, and merchandise returns provide additional predictive power to purchase incidence modeling for database marketing. © 2001 John Wiley & Sons, Inc. Stijn Viaene, Bart Baesens, Tony Van Gestel, Johan A. K. Suykens, Dirk Van den Poel, Jan Vanthienen, Bart De Moor, Guido Dedene |
Int. J. Intell. Syst. | 7 |
| 2001 | Improved Long-Term Temperature Prediction by Chaining of Neural NetworksabstractWhen an artificial neural network (ANN) is trained to predict signals p steps ahead, the quality of the prediction typically decreases for large values of p. In this paper, we compare two methods for prediction with ANNs: the classical recursion of one-step ahead predictors and a new kind of chain structure. When applying both techniques to the prediction of the temperature at the end of a blast furnace, we conclude that the chaining approach leads to an improved prediction of the temperature and avoidance of instabilities, since the chained networks gradually take the prediction of their predecessors in the chain as an extra input. It is observed that instabilities might occur in the iterative case, which does not happen with the chaining approach. To select relevant inputs and decrease the number of weights in this approach, Automatic Relevance Determination (ARD) for multilayer perceptrons is applied. Michel Duhoux, Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
Int. J. Neural Syst. | 3 |
| 2001 | Optimal control by least squares support vector machines
Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
Neural Networks | 3 |
| 2001 | IQML-like algorithms for solving structured total least squares problems: a unified view
Philippe Lemmerling, Leentje Vanhamme, Sabine Van Huffel, Bart De Moor |
Signal Process. | 4 |
| 2001 | Financial time series prediction using least squares support vector machines within the evidence frameworkabstractThe Bayesian evidence framework is applied in this paper to least squares support vector machine (LS-SVM) regression in order to infer nonlinear models for predicting a financial time series and the related volatility. On the first level of inference, a statistical framework is related to the LS-SVM formulation which allows one to include the time-varying volatility of the market by an appropriate choice of several hyper-parameters. The hyper-parameters of the model are inferred on the second level of inference. The inferred hyper-parameters, related to the volatility, are used to construct a volatility model within the evidence framework. Model comparison is performed on the third level of inference in order to automatically tune the parameters of the kernel function and to select the relevant inputs. The LS-SVM formulation allows one to derive analytic expressions in the feature space and practical expressions are obtained in the dual space replacing the inner product by the related kernel function using Mercer's theorem. The one step ahead prediction performances obtained on the prediction of the weekly 90-day T-bill rate and the daily DAX30 closing prices show that significant out of sample sign predictions can be made with respect to the Pesaran-Timmerman test statistic. Tony Van Gestel, Johan A. K. Suykens, Dirk-Emma Baestaens, Annemie Lambrechts, Gert R. G. Lanckriet, Bruno Vandaele, Bart De Moor, Joos Vandewalle |
IEEE Trans. Neural Networks | 7 |
| 2000 | Bayesian Networks in Ovarian Cancer Diagnosis: Potentials and LimitationsabstractThe pre-operative discrimination between malignant and benign masses is a crucial issue in gynaecology. Next to the large amount of background knowledge, there is a growing amount of collected patient data that can be used in inductive techniques. These two sources of information result in two different modelling strategies. Based on the background knowledge, various discrimination models have been constructed by leading experts in the field, tuned and tested by observations. Based on the patient observations, various statistical models have been developed, such as logistic regression models and artificial neural network models. For the efficient combination of prior background knowledge and observations, Bayesian network models are suggested. We summarize the applicability of this technique, report the performance of such models in ovarian cancer diagnosis and outline a possible hybrid usage of this technique. Péter Antal, Herman Verrelst, Dirk Timmerman, Sabine Van Huffel, Bart De Moor, Ignace Vergote |
CBMS | 5 |
| 2000 | SVD-based methodologies for fetal electrocardiogram extractionabstractThis paper deals with the extraction of the antepartum foetal electrocardiogram (ECG) from multilead cutaneous potential recordings, and the modelling of the transfer to the electrodes. We give an overview of a class of algebraic approaches, based on variants of the singular value decomposition (SVD): the concept, pros and cons of techniques relying on the ordinary SVD, quotient SVD and multilinear SVD are discussed. Lieven De Lathauwer, Bart De Moor, Joos Vandewalle |
ICASSP | 2 |
| 2000 | An empirical assessment of kernel type performance for least squares support vector machine classifiersabstractRecently, a modified version of support vector machines (SVMs), least-squares SVM (LS-SVM) classifiers, has been introduced, which is closely related to a form of ridge regression-type SVMs. In LS-SVMs, the classifier is obtained as the solution to a linear system instead of a quadratic programming problem. In this paper, UCI (University of California at Irvine) benchmark data sets are used to evaluate the performance of LS-SVM classifiers with linear, polynomial and radial basis function (RBF) kernels. The hyperparameters of the LS-SVM problem formulation are tuned using a 10-fold cross-validation procedure and a grid search mechanism. When comparing the performance of a nonlinear (RBF or polynomial) LS-SVM classifier with that of a linear LS-SVM, additional insight can be gained into the degree of nonlinearity of the classification problem at hand. Using a statistical motivation, it is concluded that RBF LS-SVM classifiers consistently yield among the best results for each data set. Bart Baesens, Stijn Viaene, Tony Van Gestel, Johan A. K. Suykens, Guido Dedene, Bart De Moor, Jan Vanthienen |
KES | 6 |
| 2000 | Knowledge Discovery Using Least Squares Support Vector Machine Classifiers: A Direct Marketing Case
Stijn Viaene, Bart Baesens, Tony Van Gestel, Johan A. K. Suykens, Dirk Van den Poel, Jan Vanthienen, Bart De Moor, Guido Dedene |
PKDD | 7 |
| 2000 | An algebraic approach to the blind identification of paraunitary filtersabstractThis paper deals with the blind identification of multiple-input multiple-output finite impulse response filters. We limit ourselves to the case of 2 outputs and 2 inputs. After a classical prewhitening, the remaining problem is the blind identification of a paraunitary filter. For this task, we derive a multilinear algebraic algorithm. This procedure is a generalization of the algorithm for independent component analysis described in Comon (1994). The performance is illustrated by means of some numerical experiments. Lieven De Lathauwer, Bart De Moor, Joos Vandewalle |
WCNC | 2 |
| 2000 | Robust local stability of multilayer recurrent neural networksabstractIn this paper we derive a condition for robust local stability of multilayer recurrent neural networks with two hidden layers. The stability condition follows from linking theory about linearization, robustness analysis of linear systems under nonlinear perturbation and matrix inequalities. A characterization of the basin of attraction of the origin is given in terms of the level set of a quadratic Lyapunov function. In a similar way like for NL theory, local stability is imposed around the origin and the apparent basin of attraction is made large by applying the criterion, while the proven basin of attraction is relatively small due to conservatism of the criterion. Modifying dynamic backpropagation by the new stability condition is discussed and illustrated by simulation examples. Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 1999 | Information measure based stochastic system identification of ATM network trafficabstractFor ATM network traffic, a new approach based on the Kullback-Leibler (1959) information measure is proposed for stochastic system identification of packet traffic. This approach, equivalent to the maximum marginal likelihood estimate, can overcome the over-modeling problem discussed by De Cock and De Moor. (see Proceedings of ELISIPCO-98, 1998) such that much more parsimonious model order N can be obtained, and then can lead to a significant reduction in the latter queueing analysis involving in O(N/sup 3/) computational complexity. A practical case study is provided for a set of Internet traffic data. Baibing Li, Bart De Moor |
ICASSP | 2 |
| 1997 | NLq Theory: A Neural Control Framework with Global Asymptotic Stability Criteria
Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
Neural Networks | 2 |
| 1996 | Modelling the Belgian Gas Consumption Using Neural Networks
Johan A. K. Suykens, Philippe Lemmerling, Wouter Favoreel, Bart De Moor, M. Crepel, P. Briol |
Neural Process. Lett. | 4 |
| 1996 | Continuous-time frequency domain subspace system identification
Peter Van Overschee, Bart De Moor |
Signal Process. | 2 |
| 1995 | NLq theory: unifications in the theory of neural networks, systems and control
Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
ESANN | 2 |
| 1994 | Static and dynamic stabilizing neural controllers, applicable to transition between equilibrium points
Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
Neural Networks | 2 |
| 1991 | Generalizations of the singular value and QR decompositions
Bart De Moor |
Signal Process. | 1 |
| 1988 | A geometrical approach for the identification of state space models with singular value decompositionabstractSome geometrically inspired concepts are studied for the identification of models for multivariable linear time-invariant systems from noisy input-output observations. Starting from a fundamental highly structured input-output matrix equation, it is shown how the singular value decomposition allows the order of the observable part of the system and its state-space model matrices to be estimated. Moreover, conditions for persistency of excitation of the inputs and the behavior of the algorithm when the data are perturbed by noise can easily be studied from a geometrical point of view. The singular values allow these concepts to be quantified. An example with an industrial plant identification is presented.> Bart De Moor, Marc Moonen, Lieven Vandenberghe, Joos Vandewalle |
ICASSP | 1 |
| 1988 | Computing all invariant states of a neural network
Bart De Moor, Lieven Vandenberghe, Joos Vandewalle |
Neural Networks | 1 |