EDBT 2026 Demo / reviewers in the wild / expert
Johan A. K. Suykens
dblp:61/3224
· DBLP profile ↗
263ranked-venue papers
21as first author
38since 2021 · last 2025
0000-0002-8846-6352ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 234 · 19 first-author · 36 since 2021Databases, data management, data science and information retrieval · 19 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 since 2021Systems, architecture and hardware · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generative Kernel Spectral ClusteringabstractModern clustering approaches often trade interpretability for performance, particularly in deep learning-based methods.We present Generative Kernel Spectral Clustering (GenKSC), a novel model combining kernel spectral clustering with generative modeling to produce both well-defined clusters and interpretable representations.By augmenting weighted variance maximization with reconstruction and clustering losses, our model creates an explorable latent space where cluster characteristics can be visualized through traversals along cluster directions.Results on MNIST and FashionMNIST datasets demonstrate the model's ability to learn meaningful cluster representations. Sonny Achten, David Winant, Johan A. K. Suykens |
ESANN | 3 |
| 2025 | Accelerating Spectral Clustering under Fairness ConstraintsabstractFairness of decision-making algorithms is an increasingly important issue. In this paper, we focus on spectral clustering with group fairness constraints, where every demographic group is represented in each cluster proportionally as in the general population. We present a new efficient method for fair spectral clustering (Fair SC) by casting the Fair SC problem within the difference of convex functions (DC) framework. To this end, we introduce a novel variable augmentation strategy and employ an alternating direction method of multipliers type of algorithm adapted to DC problems. We show that each associated subproblem can be solved efficiently, resulting in higher computational efficiency compared to prior work, which required a computationally expensive eigendecomposition. Numerical experiments demonstrate the effectiveness of our approach on both synthetic and real-world benchmarks, showing significant speedups in computation time over prior art, especially as the problem size grows. This work thus represents a considerable step forward towards the adoption of fair clustering in real-world applications. Francesco Tonin, Alex Lambert, Johan A. K. Suykens, Volkan Cevher |
ICML | 3 |
| 2025 | Rethinking PCA Through DualityabstractMotivated by the recently shown connection between self-attention and (kernel) principal component analysis (PCA), we revisit the fundamentals of PCA. Using the difference-of-convex (DC) framework, we present several novel formulations and provide new theoretical insights. In particular, we show the kernelizability and out-of-sample applicability for a PCA-like family of problems. Moreover, we uncover that simultaneous iteration, which is connected to the classical QR algorithm, is an instance of the difference-of-convex algorithm (DCA), offering an optimization perspective on this longstanding method. Further, we describe new algorithms for PCA and empirically compare them with state-of-the-art methods. Lastly, we introduce a kernelizable dual formulation for a robust variant of PCA that minimizes the $l_1$-deviation of the reconstruction errors. Jan Quan, Johan A. K. Suykens, Panagiotis Patrinos |
NeurIPS | 2 |
| 2025 | Nonlinear functional regression by functional deep neural network with kernel embeddingabstractRecently, deep learning has been widely applied in functional data analysis (FDA) with notable empirical success. However, the infinite dimensionality of functional data necessitates an effective dimension reduction approach for functional learning tasks, particularly in nonlinear functional regression. In this paper, we introduce a functional deep neural network with an adaptive and discretization-invariant dimension reduction method. Our functional network architecture consists of three parts: first, a kernel embedding step that features an integral transformation with an adaptive smooth kernel; next, a projection step that uses eigenfunction bases based on a projection Mercer kernel for the dimension reduction; and finally, a deep ReLU neural network is employed for the prediction. Explicit rates of approximating nonlinear smooth functionals across various input function spaces by our proposed functional network are derived. Additionally, we conduct a generalization analysis for the empirical risk minimization (ERM) algorithm applied to our functional net, by employing a novel two-stage oracle inequality and the established functional approximation results. Ultimately, we conduct numerical experiments on both simulated and real datasets to demonstrate the effectiveness and benefits of our functional net. Zhongjie Shi, Linhao Song, Ding-Xuan Zhou, Johan A. K. Suykens |
J. Mach. Learn. Res. | 5 |
| 2024 | Unsupervised Neighborhood Propagation Kernel Layers for Semi-supervised Node ClassificationabstractWe present a deep Graph Convolutional Kernel Machine (GCKM) for semi-supervised node classification in graphs. The method is built of two main types of blocks: (i) We introduce unsupervised kernel machine layers propagating the node features in a one-hop neighborhood, using implicit node feature mappings. (ii) We specify a semi-supervised classification kernel machine through the lens of the Fenchel-Young inequality. We derive an effective initialization scheme and efficient end-to-end training algorithm in the dual variables for the full architecture. The main idea underlying GCKM is that, because of the unsupervised core, the final model can achieve higher performance in semi-supervised node classification when few labels are available for training. Experimental results demonstrate the effectiveness of the proposed framework. Sonny Achten, Francesco Tonin, Panagiotis Patrinos, Johan A. K. Suykens |
AAAI | 4 |
| 2024 | Feature Learning using Multi-view Kernel Partial Least SquaresabstractThe multi-view learning deals with data of multiple views, aiming to explore the underlying relations between different views and use them for various tasks.In this paper, we derive a multi-view extension of kernel partial least squares for unsupervised feature learning.We establish the optimization objective in the primal as the pairwise covariance between the projection scores and derive that this model can be trained in the dual form by solving an eigenvalue problem.Experiments are also conducted to verify the effectiveness of the method with real-life multi-view datasets, where the proposed method is adopted as a feature extractor and then the clustering task is conducted for performance comparisons. Xinjie Zeng, Qinghua Tao, Johan A. K. Suykens |
ESANN | 3 |
| 2024 | Self-Attention through Kernel-Eigen Pair Sparse Variational Gaussian ProcessesabstractWhile the great capability of Transformers significantly boosts prediction accuracy, it could also yield overconfident predictions and require calibrated uncertainty estimation, which can be commonly tackled by Gaussian processes (GPs). Existing works apply GPs with symmetric kernels under variational inference to the attention kernel; however, omitting the fact that attention kernels are in essence asymmetric. Moreover, the complexity of deriving the GP posteriors remains high for large-scale data. In this work, we propose Kernel-Eigen Pair Sparse Variational Gaussian Processes (KEP-SVGP) for building uncertainty-aware self-attention where the asymmetry of attention kernels is tackled by Kernel SVD (KSVD) and a reduced complexity is acquired. Through KEP-SVGP, i) the SVGP pair induced by the two sets of singular vectors from KSVD w.r.t. the attention kernel fully characterizes the asymmetry; ii) using only a small set of adjoint eigenfunctions from KSVD, the derivation of SVGP posteriors can be based on the inversion of a diagonal matrix containing singular values, contributing to a reduction in time complexity; iii) an evidence lower bound is derived so that variational parameters and network weights can be optimized with it. Experiments verify our excellent performances and efficiency on in-distribution, distribution-shift and out-of-distribution benchmarks. Yingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. Suykens |
ICML | 4 |
| 2024 | Learning in Feature Spaces via Coupled Covariances: Asymmetric Kernel SVD and Nyström methodabstractIn contrast with Mercer kernel-based approaches as used e.g. in Kernel Principal Component Analysis (KPCA), it was previously shown that Singular Value Decomposition (SVD) inherently relates to asymmetric kernels and Asymmetric Kernel Singular Value Decomposition (KSVD) has been proposed. However, the existing formulation to KSVD cannot work with infinite-dimensional feature mappings, the variational objective can be unbounded, and needs further numerical evaluation and exploration towards machine learning. In this work, i) we introduce a new asymmetric learning paradigm based on coupled covariance eigenproblem (CCE) through covariance operators, allowing infinite-dimensional feature maps. The solution to CCE is ultimately obtained from the SVD of the induced asymmetric kernel matrix, providing links to KSVD. ii) Starting from the integral equations corresponding to a pair of coupled adjoint eigenfunctions, we formalize the asymmetric Nyström method through a finite sample approximation to speed up training. iii) We provide the first empirical evaluations verifying the practical utility and benefits of KSVD and compare with methods resorting to symmetrization or linear SVD across multiple tasks. Qinghua Tao, Francesco Tonin, Alex Lambert, Yingyi Chen, Panagiotis Patrinos, Johan A. K. Suykens |
ICML | 6 |
| 2024 | Explaining the model and feature dependencies by decomposition of the Shapley value
Joran Michiels, Johan A. K. Suykens, Maarten De Vos |
Decis. Support Syst. | 2 |
| 2024 | Deep Kernel Principal Component Analysis for multi-level feature learningabstractPrincipal Component Analysis (PCA) and its nonlinear extension Kernel PCA (KPCA) are widely used across science and industry for data analysis and dimensionality reduction. Modern deep learning tools have achieved great empirical success, but a framework for deep principal component analysis is still lacking. Here we develop a deep kernel PCA methodology (DKPCA) to extract multiple levels of the most informative components of the data. Our scheme can effectively identify new hierarchical variables, called deep principal components, capturing the main characteristics of high-dimensional data through a simple and interpretable numerical optimization. We couple the principal components of multiple KPCA levels, theoretically showing that DKPCA creates both forward and backward dependency across levels, which has not been explored in kernel methods and yet is crucial to extract more informative features. Various experimental evaluations on multiple data types show that DKPCA finds more efficient and disentangled representations with higher explained variance in fewer principal components, compared to the shallow KPCA. We demonstrate that our method allows for effective hierarchical data exploration, with the ability to separate the key generative factors of the input data both for large datasets and when few training samples are available. Overall, DKPCA can facilitate the extraction of useful patterns from high-dimensional data by learning more informative features organized in different levels, giving diversified aspects to explore the variation factors in the data, while maintaining a simple mathematical formulation. Francesco Tonin, Qinghua Tao, Panagiotis Patrinos, Johan A. K. Suykens |
Neural Networks | 4 |
| 2024 | Compressing Features for Learning With Noisy LabelsabstractSupervised learning can be viewed as distilling relevant information from input data into feature representations. This process becomes difficult when supervision is noisy as the distilled information might not be relevant. In fact, recent research shows that networks can easily overfit all labels including those that are corrupted, and hence can hardly generalize to clean datasets. In this article, we focus on the problem of learning with noisy labels and introduce compression inductive bias to network architectures to alleviate this overfitting problem. More precisely, we revisit one classical regularization named Dropout and its variant Nested Dropout. Dropout can serve as a compression constraint for its feature dropping mechanism, while Nested Dropout further learns ordered feature representations with respect to feature importance. Moreover, the trained models with compression regularization are further combined with co-teaching for performance boost. Theoretically, we conduct bias variance decomposition of the objective function under compression regularization. We analyze it for both single model and co-teaching. This decomposition provides three insights: 1) it shows that overfitting is indeed an issue in learning with noisy labels; 2) through an information bottleneck formulation, it explains why the proposed feature compression helps in combating label noise; and 3) it gives explanations on the performance boost brought by incorporating compression regularization into co-teaching. Experiments show that our simple approach can have comparable or even better performance than the state-of-the-art methods on benchmarks with real-world label noise including Clothing1M and ANIMAL-10N. Our implementation is available at https://yingyichen-cyy.github.io/CompressFeatNoisyLabels/. Yingyi Chen, Shell Xu Hu, Xi Shen 0001, Chunrong Ai, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Unbalanced Optimal Transport: A Unified Framework for Object DetectionabstractDuring training, supervised object detection tries to correctly match the predicted bounding boxes and associated classification scores to the ground truth. This is essential to determine which predictions are to be pushed towards which solutions, or to be discarded. Popular matching strategies include matching to the closest ground truth box (mostly used in combination with anchors), or matching via the Hungarian algorithm (mostly used in anchor free methods). Each of these strategies comes with its own properties, underlying losses, and heuristics. We show how Unbalanced Optimal Transport unifies these different approaches and opens a whole continuum of methods in between. This allows for a finer selection of the desired properties. Experimentally, we show that training an object detection model with Unbalanced Optimal Transport is able to reach the state-of-the-art both in terms of Average Precision and Average Recall as well as to provide a faster initial convergence. The approach is well suited for GPU implementation, which proves to be an advantage for large-scale models. Henri De Plaen, Pierre-François De Plaen, Johan A. K. Suykens, Marc Proesmans, Tinne Tuytelaars, Luc Van Gool |
CVPR | 3 |
| 2023 | Tensorized LSSVMS For Multitask RegressionabstractMultitask learning (MTL) can utilize the relatedness between multiple tasks for performance improvement. The advent of multimodal data allows tasks to be referenced by multiple indices. High-order tensors are capable of providing efficient representations for such tasks, while preserving structural task-relations. In this paper, a new MTL method is proposed by leveraging low-rank tensor analysis and constructing tensorized Least Squares Support Vector Machines, namely the tLSSVM-MTL, where multilinear modelling and its nonlinear extensions can be flexibly exerted. We employ a high-order tensor for all the weights with each mode relating to an index and factorize it with CP decomposition, assigning a shared factor for all tasks and retaining task-specific latent factors along each index. Then an alternating algorithm is derived for the nonconvex optimization, where each resulting subproblem is solved by a linear system. Experimental results demonstrate promising performances of our tLSSVM-MTL. Jiani Liu 0002, Qinghua Tao, Ce Zhu, Yipeng Liu 0001, Johan A. K. Suykens |
ICASSP | 5 |
| 2023 | Extending Kernel PCA through Dualization: Sparsity, Robustness and Fast AlgorithmsabstractThe goal of this paper is to revisit Kernel Principal Component Analysis (KPCA) through dualization of a difference of convex functions. This allows to naturally extend KPCA to multiple objective functions and leads to efficient gradient-based algorithms avoiding the expensive SVD of the Gram matrix. Particularly, we consider objective functions that can be written as Moreau envelopes, demonstrating how to promote robustness and sparsity within the same framework. The proposed method is evaluated on synthetic and realworld benchmarks, showing significant speedup in KPCA training time as well as highlighting the benefits in terms of robustness and sparsity. Francesco Tonin, Alex Lambert, Panagiotis Patrinos, Johan A. K. Suykens |
ICML | 4 |
| 2023 | Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal RepresentationabstractRecently, a new line of works has emerged to understand and improve self-attention in Transformers by treating it as a kernel machine. However, existing works apply the methods for symmetric kernels to the asymmetric self-attention, resulting in a nontrivial gap between the analytical understanding and numerical implementation. In this paper, we provide a new perspective to represent and optimize self-attention through asymmetric Kernel Singular Value Decomposition (KSVD), which is also motivated by the low-rank property of self-attention normally observed in deep layers. Through asymmetric KSVD, i) a primal-dual representation of self-attention is formulated, where the optimization objective is cast to maximize the projection variances in the attention outputs; ii) a novel attention mechanism, i.e., Primal-Attention, is proposed via the primal representation of KSVD, avoiding explicit computation of the kernel matrix in the dual; iii) with KKT conditions, we prove that the stationary solution to the KSVD optimization in Primal-Attention yields a zero-value objective. In this manner, KSVD optimization can be implemented by simply minimizing a regularization loss, so that low-rank property is promoted without extra decomposition. Numerical experiments show state-of-the-art performance of our Primal-Attention with improved efficiency. Moreover, we demonstrate that the deployed KSVD optimization regularizes Primal-Attention with a sharper singular value decay than that of the canonical self-attention, further verifying the great potential of our method. To the best of our knowledge, this is the first work that provides a primal-dual representation for the asymmetric kernel in self-attention and successfully applies it to modelling and optimization. Yingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. Suykens |
NeurIPS | 4 |
| 2023 | Multi-view kernel PCA for time series forecastingabstractIn this paper, we propose a kernel principal component analysis model for multi-variate time series forecasting, where the training and prediction schemes are derived from the multi-view formulation of Restricted Kernel Machines. The training problem is simply an eigenvalue decomposition of the summation of two kernel matrices corresponding to the views of the input and output data. When a linear kernel is used for the output view, it is shown that the forecasting equation takes the form of kernel ridge regression. When that kernel is non-linear, a pre-image problem has to be solved to forecast a point in the input space. We evaluate the model on several standard time series datasets, perform ablation studies, benchmark with closely related models and discuss its results. Arun Pandey, Hannes De Meulemeester, Bart De Moor, Johan A. K. Suykens |
Neurocomputing | 4 |
| 2023 | Learning With Asymmetric Kernels: Least Squares and Feature InterpretationabstractAsymmetric kernels naturally exist in real life, e.g., for conditional probability and directed graphs. However, most of the existing kernel-based learning methods require kernels to be symmetric, which prevents the use of asymmetric kernels. This paper addresses the asymmetric kernel-based learning in the framework of the least squares support vector machine named AsK-LS, resulting in the first classification method that can utilize asymmetric kernels directly. We will show that AsK-LS can learn with asymmetric features, namely source and target features, while the kernel trick remains applicable, i.e., the source and target features exist but are not necessarily known. Besides, the computational burden of AsK-LS is as cheap as dealing with symmetric kernels. Experimental results on various tasks, including Corel, PASCAL VOC, Satellite, directed graphs, and UCI database, all show that in the case asymmetric information is crucial, the proposed AsK-LS can learn with asymmetric kernels and performs much better than the existing kernel methods that rely on symmetrization to accommodate asymmetric kernels. Mingzhen He, Lei Shi 0010, Xiaolin Huang, Johan A. K. Suykens |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Jigsaw-ViT: Learning jigsaw puzzles in vision transformerabstractThe success of Vision Transformer (ViT) in various computer vision tasks has promoted the ever-increasing prevalence of this convolution-free network. The fact that ViT works on image patches makes it potentially relevant to the problem of jigsaw puzzle solving, which is a classical self-supervised task aiming at reordering shuffled sequential image patches back to their original form. Solving jigsaw puzzle has been demonstrated to be helpful for diverse tasks using Convolutional Neural Networks (CNNs), such as feature representation learning, domain generalization and fine-grained classification. In this paper, we explore solving jigsaw puzzle as a self-supervised auxiliary loss in ViT for image classification, named Jigsaw-ViT. We show two modifications that can make Jigsaw-ViT superior to standard ViT: discarding positional embeddings and masking patches randomly. Yet simple, we find that the proposed Jigsaw-ViT is able to improve on both generalization and robustness over the standard ViT, which is usually rather a trade-off. Numerical experiments verify that adding the jigsaw puzzle branch provides better generalization to ViT on large-scale image classification on ImageNet. Moreover, such auxiliary loss also improves robustness against noisy labels on Animal-10N, Food-101N, and Clothing1M, as well as adversarial examples. Our implementation is available at https://yingyichen-cyy.github.io/Jigsaw-ViT. Yingyi Chen, Xi Shen 0001, Qinghua Tao, Johan A. K. Suykens |
Pattern Recognit. Lett. | 5 |
| 2023 | Island Transpeciation: A Co-Evolutionary Neural Architecture Search, Applied to Country-Scale Air-Quality ForecastingabstractAir pollution causes around 400 000 premature deaths per year in Europe due to Particulate Matter, nitrogen oxides, and ground-level ozone pollutants. Multiple-input multiple-output nonlinear auto-regressive exogenous deep neural networks are frequently used to predict a day before, air-quality pollution incidents, at a country scale. With complexity and data sizes increasing, finding performant models becomes harder. We propose island transpeciation to optimize hyperparameters and architectures. Unlike using a single optimizer, island transpeciation combines results from multiple optimizers, to consistently provide excellent performance. Moreover, we show that island transpeciation outperforms random model search and other previous modeling efforts. Island transpeciation is a neural architecture search that uses co-evolution (genes), to combine (transpeciation) populations of incompatible optimizers (species) organized in island formations. In island transpeciation, architecture search is parallelized and utilizes a distributed pool of hardware resources. We have successfully used these techniques to predict next-day ozone concentrations across the Belgian territory. Konstantinos Theodorakos, Oscar Mauricio Agudelo, Joachim Schreurs, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Evol. Comput. | 4 |
| 2022 | Recurrent Restricted Kernel Machines for Time-series ForecastingabstractIn this paper, we propose a novel method for time-series modeling and forecasting.It is based on the temporal formulation of Restricted Kernel Machines leading to a dynamical equation in the latent-variables.Forecasting involves finding the next latent variable and then solving a pre-image problem to predict a new-point in the input space.Further, we benchmark our model on several standard data sets against other well-known time-series models. Arun Pandey, Hannes De Meulemeester, Henri De Plaen, Bart De Moor, Johan A. K. Suykens |
ESANN | 5 |
| 2022 | On the Double Descent of Random Features Models Trained with SGDabstractWe study generalization properties of random features (RF) regression in high dimensions optimized by stochastic gradient descent (SGD) in under-/over-parameterized regime. In this work, we derive precise non-asymptotic error bounds of RF regression under both constant and polynomial-decay step-size SGD setting, and observe the double descent phenomenon both theoretically and empirically. Our analysis shows how to cope with multiple randomness sources of initialization, label noise, and data sampling (as well as stochastic gradients) with no closed-form solution, and also goes beyond the commonly-used Gaussian/spherical data assumption. Our theoretical results demonstrate that, with SGD training, RF regression still generalizes well for interpolation learning, and is able to characterize the double descent behavior by the unimodality of variance and monotonic decrease of bias. Besides, we also prove that the constant step-size SGD setting incurs no loss in convergence rate when compared to the exact minimum-norm interpolator, as a theoretical justification of using SGD in practice. Fanghui Liu 0001, Johan A. K. Suykens, Volkan Cevher |
NeurIPS | 2 |
| 2022 | Nyström landmark sampling and regularized Christoffel functions
Michaël Fanuel, Joachim Schreurs, Johan A. K. Suykens |
Mach. Learn. | 3 |
| 2022 | Disentangled Representation Learning and Generation With Manifold OptimizationabstractDisentanglement is a useful property in representation learning, which increases the interpretability of generative models such as variational autoencoders (VAE), generative adversarial models, and their many variants. Typically in such models, an increase in disentanglement performance is traded off with generation quality. In the context of latent space models, this work presents a representation learning framework that explicitly promotes disentanglement by encouraging orthogonal directions of variations. The proposed objective is the sum of an autoencoder error term along with a principal component analysis reconstruction error in the feature space. This has an interpretation of a restricted kernel machine with the eigenvector matrix valued on the Stiefel manifold. Our analysis shows that such a construction promotes disentanglement by matching the principal directions in the latent space with the directions of orthogonal variation in data space. In an alternating minimization scheme, we use the Cayley ADAM algorithm, a stochastic optimization method on the Stiefel manifold along with the Adam optimizer. Our theoretical discussion and various experiments show that the proposed model is an improvement over many VAE variants in terms of both generation quality and disentangled representation learning. Arun Pandey, Michaël Fanuel, Joachim Schreurs, Johan A. K. Suykens |
Neural Comput. | 4 |
| 2022 | Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and BeyondabstractThe class of random features is one of the most popular techniques to speed up kernel methods in large-scale problems. Related works have been recognized by the NeurIPS Test-of-Time award in 2017 and the ICML Best Paper Finalist in 2019. The body of work on random features has grown rapidly, and hence it is desirable to have a comprehensive overview on this topic explaining the connections among various algorithms and theoretical results. In this survey, we systematically review the work on random features from the past ten years. First, the motivations, characteristics and contributions of representative random features based algorithms are summarized according to their sampling schemes, learning procedures, variance reduction properties and how they exploit training data. Second, we review theoretical results that center around the following key question: how many random features are needed to ensure a high approximation quality or no loss in the empirical/expected risks of the learned estimator. Third, we provide a comprehensive evaluation of popular random features based algorithms on several large-scale benchmark datasets and discuss their approximation quality and prediction performance for classification. Last, we discuss the relationship between random features and modern over-parameterized deep neural networks (DNNs), including the use of high dimensional random features in the analysis of DNNs as well as the gaps between current theoretical and empirical results. This survey may serve as a gentle introduction to this topic, and as a users' guide for practitioners interested in applying the representative algorithms and understanding theoretical results under various technical assumptions. We hope that this survey will facilitate discussion on the open problems in this topic, and more importantly, shed light on future research directions. Due to the page limit, we suggest the readers refer to the full version of this survey https://arxiv.org/abs/2004.11154. Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Johan A. K. Suykens |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Towards a Unified Quadrature Framework for Large-Scale Kernel MachinesabstractIn this paper, we develop a quadrature framework for large-scale kernel machines via a numerical integration representation. Considering that the integration domain and measure of typical kernels, e.g., Gaussian kernels, arc-cosine kernels, are fully symmetric, we leverage a numerical integration technique, deterministic fully symmetric interpolatory rules, to efficiently compute quadrature nodes and associated weights for kernel approximation. Thanks to the full symmetric property, the applied interpolatory rules are able to reduce the number of needed nodes while retaining a high approximation accuracy. Further, we randomize the above deterministic rules by the classical Monte-Carlo sampling and control variates techniques with two merits: 1) The proposed stochastic rules make the dimension of the feature mapping flexibly varying, such that we can control the discrepancy between the original and approximate kernels by tuning the dimnension. 2) Our stochastic rules have nice statistical properties of unbiasedness and variance reduction. In addition, we elucidate the relationship between our deterministic/stochastic interpolatory rules and current typical quadrature based rules for kernel approximation, thereby unifying these methods under our framework. Experimental results on several benchmark datasets show that our methods compare favorably with other representative kernel approximation based methods. Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Johan A. K. Suykens |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Short-Term Traffic Flow Prediction Based on the Efficient Hinging Hyperplanes Neural NetworkabstractTraffic flow (TF) prediction is an important and yet a challenging task in transportation systems, since the TF involves high nonlinearities and is affected by many elements. Recently, neural networks have attracted much attention for TF prediction, but they are commonly black boxes with complex architectures and difficult to be interpreted, e.g., the contributions of specific traffic elements are not explicit, hardly providing informative guidance. In this paper, we aim at addressing more interpretable short-term TF prediction with joint consideration to high accuracy, and thus introduces a pragmatic method by applying the efficient hinging hyperplanes neural network (EHHNN) simply built upon sparse neuron connections. In the proposed method, different traffic factors are incorporated into the inputs, including their spatial-temporal information. Besides the pursuit of accuracy, we further extend the ANOVA decomposition of EHHNNs to the interpretation analysis with specifications to traffic data, in which the contributions concerning specific traffic variables are detected quantitatively. As such, the proposed method firstly applies the EHHNN to filter out more important traffic variables for dimensionality reduction while maintaining accurate prediction. Then, variable interpretation analysis is performed from different perspectives, e.g. to quantitatively investigate the influence of traffic factors and also their spatial-temporal impacts. Therefore, a predictor and an analyzing tool can both be attained for the TF by exerting the flexibility and extending the interpretability of EHHNNs, which is promising to provide informative guidance to future traffic control. Numerical experiments verify the effectiveness and potential of the proposed method in TF prediction and analysis. Qinghua Tao, Zhen Li 0032, Jun Xu 0008, Shu Lin 0002, Bart De Schutter, Johan A. K. Suykens |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Toward Deep Adaptive Hinging HyperplanesabstractThe adaptive hinging hyperplane (AHH) model is a popular piecewise linear representation with a generalized tree structure and has been successfully applied in dynamic system identification. In this article, we aim to construct the deep AHH (DAHH) model to extend and generalize the networking of AHH model for high-dimensional problems. The network structure of DAHH is determined through a forward growth, in which the activity ratio is introduced to select effective neurons and no connecting weights are involved between the layers. Then, all neurons in the DAHH network can be flexibly connected to the output in a skip-layer format, and only the corresponding weights are the parameters to optimize. With such a network framework, the backpropagation algorithm can be implemented in DAHH to efficiently tackle large-scale problems and the gradient vanishing problem is not encountered in the training of DAHH. In fact, the optimization problem of DAHH can maintain convexity with convex loss in the output layer, which brings natural advantages in optimization. Different from the existing neural networks, DAHH is easier to interpret, where neurons are connected sparsely and analysis of variance (ANOVA) decomposition can be applied, facilitating to revealing the interactions between variables. A theoretical analysis toward universal approximation ability and explicit domain partitions are also derived. Numerical experiments verify the effectiveness of the proposed DAHH. Qinghua Tao, Jun Xu 0008, Zhen Li 0032, Na Xie, Shuning Wang, Xiaoli Li 0011, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2021 | Fast Learning in Reproducing Kernel Krein Spaces via Signed MeasuresabstractIn this paper, we attempt to solve a long-lasting open question for non-positive definite (non-PD) kernels in machine learning community: can a given non-PD kernel be decomposed into the difference of two PD kernels (termed as positive decomposition)? We cast this question as a distribution view by introducing the signed measure, which transforms positive decomposition to measure decomposition: a series of non-PD kernels can be associated with the linear combination of specific finite Borel measures. In this manner, our distribution-based framework provides a sufficient and necessary condition to answer this open question. Specifically, this solution is also computationally implementable in practice to scale non-PD kernels in large sample cases, which allows us to devise the first random features algorithm to obtain an unbiased estimator. Experimental results on several benchmark datasets verify the effectiveness of our algorithm over the existing methods. Fanghui Liu 0001, Xiaolin Huang, Yingyi Chen, Johan A. K. Suykens |
AISTATS | 4 |
| 2021 | Kernel regression in high dimensions: Refined analysis beyond double descentabstractIn this paper, we provide a precise characterization of generalization properties of high dimensional kernel ridge regression across the under- and over-parameterized regimes, depending on whether the number of training data n exceeds the feature dimension d. By establishing a bias-variance decomposition of the expected excess risk, we show that, while the bias is (almost) independent of d and monotonically decreases with n, the variance depends on n,d and can be unimodal or monotonically decreasing under different regularization schemes. Our refined analysis goes beyond the double descent theory by showing that, depending on the data eigen-profile and the level of regularization, the kernel regression risk curve can be a double-descent-like, bell-shaped, or monotonic function of n. Experiments on synthetic and real data are conducted to support our theoretical findings. Fanghui Liu 0001, Zhenyu Liao 0001, Johan A. K. Suykens |
AISTATS | 3 |
| 2021 | Unsupervised Energy-based Out-of-distribution Detection using Stiefel-Restricted Kernel MachineabstractDetecting out-of-distribution (OOD) samples is an essential requirement for the deployment of machine learning systems in the real world. Until now, research on energy-based OOD detectors has focused on the softmax confidence score from a pre-trained neural network classifier with access to class labels. In contrast, we propose an unsupervised energy-based OOD detector leveraging the Stiefel-Restricted Kernel Machine (St-RKM). Training requires minimizing an objective function with an autoencoder loss term and the RKM energy where the interconnection matrix lies on the Stiefel manifold. Further, we outline multiple energy function definitions based on the RKM framework and discuss their utility. In the experiments on standard datasets, the proposed method improves over the existing energy-based OOD detectors and deep generative models. Through several ablation studies, we further illustrate the merit of each proposed energy function on the OOD detection performance. Francesco Tonin, Arun Pandey, Panagiotis Patrinos, Johan A. K. Suykens |
IJCNN | 4 |
| 2021 | The Bures Metric for Generative Adversarial Networks
Hannes De Meulemeester, Joachim Schreurs, Michaël Fanuel, Bart De Moor, Johan A. K. Suykens |
ECML/PKDD (2) | 5 |
| 2021 | Kernel Machines in Time (Invited Talk)abstractKernel machines is a powerful class of models in machine learning with solid foundations and many existing application fields. The scope of this talk is kernel machines in time with a main focus on least squares support vector machines, and other related methods such as kernel principal component analysis and kernel spectral clustering. For dynamical systems modelling different possible input-output and state space model structures will be discussed. Applications will be shown on electricity load forecasting and temperature prediction in weather forecasting. Approximate closed-form solutions can be given to ordinary and partial differential equations. Kernel spectral clustering applications to identifying customer profiles, pollution modelling and detecting topological changes in time-series of bridges will be shown. Finally, new synergies between kernel machines and deep learning will be presented, leading for example to generative kernel machines, with new insights on disentangled representations, explainability and latent space exploration. Application of these models will be illustrated on out-of-distribution detection of time-series data. Johan A. K. Suykens |
TIME | 1 |
| 2021 | Learning with continuous piecewise linear decision trees
Qinghua Tao, Zhen Li 0032, Jun Xu 0008, Na Xie, Shuning Wang, Johan A. K. Suykens |
Expert Syst. Appl. | 6 |
| 2021 | A novel neural grey system model with Bayesian regularization and its applications
Xin Ma 0004, Mei Xie, Johan A. K. Suykens |
Neurocomputing | 3 |
| 2021 | Generalization Properties of hyper-RKHS and its ApplicationsabstractThis paper generalizes regularized regression problems in a hyper-reproducing kernel Hilbert space (hyper-RKHS), illustrates its utility for kernel learning and out-of-sample extensions, and proves asymptotic convergence results for the introduced regression models in an approximation theory view. Algorithmically, we consider two regularized regression models with bivariate forms in this space, including kernel ridge regression (KRR) and support vector regression (SVR) endowed with hyper-RKHS, and further combine divide-and-conquer with Nyström approximation for scalability in large sample cases. This framework is general: the underlying kernel is learned from a broad class, and can be positive definite or not, which adapts to various requirements in kernel learning. Theoretically, we study the convergence behavior of regularized regression algorithms in hyper-RKHS and derive the learning rates, which goes beyond the classical analysis on RKHS due to the non-trivial independence of pairwise samples and the characterisation of hyper-RKHS. Experimentally, results on several benchmarks suggest that the employed framework is able to learn a general kernel function form an arbitrary similarity matrix, and thus achieves a satisfactory performance on classification tasks. Fanghui Liu 0001, Lei Shi 0010, Xiaolin Huang, Jie Yang 0002, Johan A. K. Suykens |
J. Mach. Learn. Res. | 5 |
| 2021 | Analysis of regularized least-squares in reproducing kernel Kreĭn spaces
Fanghui Liu 0001, Lei Shi 0010, Xiaolin Huang, Jie Yang 0002, Johan A. K. Suykens |
Mach. Learn. | 5 |
| 2021 | Generative Restricted Kernel Machines: A framework for multi-view generation and disentangled feature learning
Arun Pandey, Joachim Schreurs, Johan A. K. Suykens |
Neural Networks | 3 |
| 2021 | Unsupervised learning of disentangled representations in deep restricted kernel machines with orthogonality constraintsabstractWe introduce Constr-DRKM, a deep kernel method for the unsupervised learning of disentangled data representations. We propose augmenting the original deep restricted kernel machine formulation for kernel PCA by orthogonality constraints on the latent variables to promote disentanglement and to make it possible to carry out optimization without first defining a stabilized objective. After discussing a number of algorithms for end-to-end training, we quantitatively evaluate the proposed method's effectiveness in disentangled feature learning. We demonstrate on four benchmark datasets that this approach performs similarly overall to β-VAE on several disentanglement metrics when few training points are available while being less sensitive to randomness and hyperparameter selection than β-VAE. We also present a deterministic initialization of Constr-DRKM's training algorithm that significantly improves the reproducibility of the results. Finally, we empirically evaluate and discuss the role of the number of layers in the proposed methodology, examining the influence of each principal component in every layer and showing that components in lower layers act as local feature detectors capturing the broad trends of the data distribution, while components in deeper layers use the representation learned by previous layers and more accurately reproduce higher-level features. Francesco Tonin, Panagiotis Patrinos, Johan A. K. Suykens |
Neural Networks | 3 |
| 2020 | Random Fourier Features via Fast Surrogate Leverage Weighted SamplingabstractIn this paper, we propose a fast surrogate leverage weighted sampling strategy to generate refined random Fourier features for kernel approximation. Compared to the current state-of-the-art method that uses the leverage weighted scheme (Li et al. 2019), our new strategy is simpler and more effective. It uses kernel alignment to guide the sampling process and it can avoid the matrix inversion operator when we compute the leverage function. Given n observations and s random features, our strategy can reduce the time complexity for sampling from O(ns2+s3) to O(ns2), while achieving comparable (or even slightly better) prediction performance when applied to kernel ridge regression (KRR). In addition, we provide theoretical guarantees on the generalization performance of our approach, and in particular characterize the number of random features required to achieve statistical guarantees in KRR. Experiments on several benchmark datasets demonstrate that our algorithm achieves comparable prediction performance and takes less time cost when compared to (Li et al. 2019). Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Jie Yang 0002, Johan A. K. Suykens |
AAAI | 5 |
| 2020 | Learning from partially labeled data
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens |
ESANN | 3 |
| 2020 | Wasserstein Exponential KernelsabstractIn the context of kernel methods, the similarity between data points is encoded by the kernel function which is often defined thanks to the Euclidean distance; the squared exponential kernel is a common example. Recently, other distances relying on optimal transport theory - such as the Wasserstein distance between probability distributions - have shown their practical relevance for different machine learning techniques. In this paper, we study the use of exponential kernels defined thanks to the regularized Wasserstein distance and discuss their positive definiteness. More specifically, we define Wasserstein feature maps and illustrate their interest for supervised learning problems involving shapes and images. Empirically, Wasserstein squared exponential kernels are shown to yield smaller classification errors on small training sets of shapes, compared to analogous classifiers using Euclidean distances. Henri De Plaen, Michaël Fanuel, Johan A. K. Suykens |
IJCNN | 3 |
| 2020 | A Theoretical Framework for Target PropagationabstractThe success of deep learning, a brain-inspired form of AI, has sparked interest in understanding how the brain could similarly learn across multiple layers of neurons. However, the majority of biologically-plausible learning algorithms have not yet reached the performance of backpropagation (BP), nor are they built on strong theoretical foundations. Here, we analyze target propagation (TP), a popular but not yet fully understood alternative to BP, from the standpoint of mathematical optimization. Our theory shows that TP is closely related to Gauss-Newton optimization and thus substantially differs from BP. Furthermore, our analysis reveals a fundamental limitation of difference target propagation (DTP), a well-known variant of TP, in the realistic scenario of non-invertible neural networks. We provide a first solution to this problem through a novel reconstruction loss that improves feedback weight training, while simultaneously introducing architectural flexibility by allowing for direct feedback connections from the output to each hidden layer. Our theory is corroborated by experimental results that show significant improvements in performance and in the alignment of forward weight updates with loss gradients, compared to DTP. Alexander Meulemans, Francesco S. Carzaniga, Johan A. K. Suykens, João Sacramento, Benjamin F. Grewe |
NeurIPS | 3 |
| 2020 | A Statistical Learning Approach to Modal RegressionabstractThis paper studies the nonparametric modal regression problem systematically from a statistical learning viewpoint. Originally motivated by pursuing a theoretical understanding of the maximum correntropy criterion based regression (MCCR), our study reveals that MCCR with a tending-to-zero scale parameter is essentially modal regression. We show that the nonparametric modal regression problem can be approached via the classical empirical risk minimization. Some efforts are then made to develop a framework for analyzing and implementing modal regression. For instance, the modal regression function is described, the modal regression risk is defined explicitly and its Bayes rule is characterized; for the sake of computational tractability, the surrogate modal regression risk, which is termed as the generalization risk in our study, is introduced. On the theoretical side, the excess modal regression risk, the excess generalization risk, the function estimation error, and the relations among the above three quantities are studied rigorously. It turns out that under mild conditions, function estimation consistency and convergence may be pursued in modal regression as in vanilla regression protocols such as mean regression, median regression, and quantile regression. On the practical side, the implementation issues of modal regression including the computational algorithm and the selection of the tuning parameters are discussed. Numerical validations on modal regression are also conducted to verify our findings. Yunlong Feng, Johan A. K. Suykens |
J. Mach. Learn. Res. | 3 |
| 2020 | Transductive LSTM for time-series prediction: An application to weather forecasting
Zahra Karevan, Johan A. K. Suykens |
Neural Networks | 2 |
| 2020 | A Double-Variational Bayesian Framework in Random Fourier Features for Indefinite KernelsabstractRandom Fourier features (RFFs) have been successfully employed to kernel approximation in large-scale situations. The rationale behind RFF relies on Bochner's theorem, but the condition is too strict and excludes many widely used kernels, e.g., dot-product kernels (violates the shift-invariant condition) and indefinite kernels [violates the positive definite (PD) condition]. In this article, we present a unified RFF framework for indefinite kernel approximation in the reproducing kernel Kreĭn spaces (RKKSs). Besides, our model is also suited to approximate a dot-product kernel on the unit sphere, as it can be transformed into a shift-invariant but indefinite kernel. By the Kolmogorov decomposition scheme, an indefinite kernel in RKKS can be decomposed into the difference of two unknown PD kernels. The spectral distribution of each underlying PD kernel can be formulated as a nonparametric Bayesian Gaussian mixtures model. Based on this, we propose a double-infinite Gaussian mixture model in RFF by placing the Dirichlet process prior. It takes full advantage of high flexibility on the number of components and has the capability of approximating indefinite kernels on a wide scale. In model inference, we develop a non-conjugate variational algorithm with a sub-sampling scheme for the posterior inference. It allows for the non-conjugate case in our model and is quite efficient due to the sub-sampling strategy. Experimental results on several large classification data sets demonstrate the effectiveness of our nonparametric Bayesian model for indefinite kernel approximation when compared to other representative random feature-based methods. Fanghui Liu 0001, Xiaolin Huang, Lei Shi 0010, Jie Yang 0002, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Deep convolutional learning for general early design stage prediction models
Sundaravelpandian Singaravel, Johan A. K. Suykens, Philipp Geyer |
Adv. Eng. Informatics | 2 |
| 2019 | Robust classification of graph-based data
Carlos M. Alaíz, Michaël Fanuel, Johan A. K. Suykens |
Data Min. Knowl. Discov. | 3 |
| 2019 | Sparse Kernel Regression with Coefficient-based $\ell_q-$regularizationabstractIn this paper, we consider the $\ell_q-$regularized kernel regression with $0 < q \leq 1$. In form, the algorithm minimizes a least-square loss functional adding a coefficient-based $\ell_q-$penalty term over a linear span of features generated by a kernel function. We study the asymptotic behavior of the algorithm under the framework of learning theory. The contribution of this paper is two-fold. First, we derive a tight bound on the $\ell_2-$empirical covering numbers of the related function space involved in the error analysis. Based on this result, we obtain the convergence rates for the $\ell_1-$regularized kernel regression which is the best so far. Second, for the case $0 < q < 1$, we show that the regularization parameter plays a role as a trade-off between sparsity and convergence rates. Under some mild conditions, the fraction of non-zero coefficients in a local minimizer of the algorithm will tend to $0$ at a polynomial decay rate when the sample size $m$ becomes large. As the concerned algorithm is non-convex, we also discuss how to generate a minimizing sequence iteratively, which can help us to search a local minimizer around any initial point. Lei Shi 0010, Xiaolin Huang, Yunlong Feng, Johan A. K. Suykens |
J. Mach. Learn. Res. | 4 |
| 2019 | Indefinite Kernel Logistic Regression With Concave-Inexact-Convex ProcedureabstractIn kernel methods, the kernels are often required to be positive definitethat restricts the use of many indefinite kernels. To consider those nonpositive definite kernels, in this paper, we aim to build an indefinite kernel learning framework for kernel logistic regression (KLR). The proposed indefinite KLR (IKLR) model is analyzed in the reproducing kernel Kreĭn spaces and then becomes nonconvex. Using the positive decomposition of a nonpositive definite kernel, the derived IKLR model can be decomposed into the difference of two convex functions. Accordingly, a concave-convex procedure (CCCP) is introduced to solve the nonconvex optimization problem. Since the CCCP has to solve a subproblem in each iteration, we propose a concave-inexact-convex procedure (CCICP) algorithm with an inexact solving scheme to accelerate the solving process. Besides, we propose a stochastic variant of CCICP to efficiently obtain a proximal solution, which achieves the similar purpose with the inexact solving scheme in CCICP. The convergence analyses of the above-mentioned two variants of CCCP are conducted. By doing so, our method works effectively not only in a deterministic setting but also in a stochastic setting. Experimental results on several benchmarks suggest that the proposed IKLR model performs favorably against the standard (positive definite) KLR and other competitive indefinite learning-based algorithms. Fanghui Liu 0001, Xiaolin Huang, Chen Gong 0002, Jie Yang 0002, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Solving lp-norm regularization with tensor kernelsabstractIn this paper, we discuss how a suitable family of tensor kernels can be used to efficiently solve nonparametric extensions of lp regularized learning methods. Our main contribution is proposing a fast dual algorithm, and showing that it allows to solve the problem efficiently. Our results contrast recent findings suggesting kernel methods cannot be extended beyond Hilbert setting. Numerical experiments confirm the effectiveness of the method. Saverio Salzo, Lorenzo Rosasco, Johan A. K. Suykens |
AISTATS | 3 |
| 2018 | Shallow and Deep Models for Domain Adaptation problems
Siamak Mehrkanoon, Matthew B. Blaschko, Johan A. K. Suykens |
ESANN | 3 |
| 2018 | Generative Kernel PCA
Joachim Schreurs, Johan A. K. Suykens |
ESANN | 2 |
| 2018 | Tensor Learning in Multi-view Kernel PCA
Lynn Houthuys, Johan A. K. Suykens |
ICANN (2) | 2 |
| 2018 | Weighted Multi-view Deep Neural Networks for Weather Forecasting
Zahra Karevan, Lynn Houthuys, Johan A. K. Suykens |
ICANN (3) | 3 |
| 2018 | Deep-learning neural-network architectures and methods: Using component-based models in building-design energy predictionabstractIncreasing sustainability requirements make evaluating different design options for identifying energy-efficient design ever more important. These requirements demand simulation models that are not only accurate but also fast. Machine Learning (ML) enables effective mimicry of Building Performance Simulation (BPS) while generating results much faster than BPS. Component-Based Machine Learning (CBML) enhances the capabilities of the monolithic ML model. Extending monolithic ML approach, the paper presents deep-learning architectures, component development methods and evaluates their suitability for space exploration in building design. Results indicate that deep learning increases the performance of models over simple artificial neural network models. Methods such as transfer learning and Multi-Task Learning make the component development process more efficient. Testing the deep-learning model on 201 new design cases indicates that its cooling energy prediction (R2: 0.983) is similar to BPS, while errors for heating energy predictions (R2: 0.848) are higher than BPS. Higher heating energy prediction error can be resolved by collecting heating data using better design space sampling methods that cover the heating demand distribution effectively. Given that the accuracy of the deep-learning model for heating predictions can be increased, the major advantage of deep-learning models over BPS is their high computation speed. BPS required 1145 s to simulate 201 design cases. Using the deep-learning model, similar results can be obtained in 0.9 s. High computation speed makes deep-learning models suitable for design space exploration. Sundaravelpandian Singaravel, Johan A. K. Suykens, Philipp Geyer |
Adv. Eng. Informatics | 2 |
| 2018 | Modified Frank-Wolfe algorithm for enhanced sparsity in support vector machine classifiers
Carlos M. Alaíz, Johan A. K. Suykens |
Neurocomputing | 2 |
| 2018 | Multi-View Least Squares Support Vector Machines Classification
Lynn Houthuys, Rocco Langone, Johan A. K. Suykens |
Neurocomputing | 3 |
| 2018 | Pinball loss minimization for one-bit compressive sensing: Convex models and algorithms
Xiaolin Huang, Lei Shi 0010, Ming Yan 0006, Johan A. K. Suykens |
Neurocomputing | 4 |
| 2018 | Deep hybrid neural-kernel networks using random Fourier features
Siamak Mehrkanoon, Johan A. K. Suykens |
Neurocomputing | 2 |
| 2018 | Kernel Density Estimation for Dynamical SystemsabstractWe study the density estimation problem with observations generated by certain dynamical systems that admit a unique underlying invariant Lebesgue density. Observations drawn from dynamical systems are not independent and moreover, usual mixing concepts may not be appropriate for measuring the dependence among these observations. By employing the $\mathcal{C}$-mixing concept to measure the dependence, we conduct statistical analysis on the consistency and convergence of the kernel density estimator. Our main results are as follows: First, we show that with properly chosen bandwidth, the kernel density estimator is universally consistent under $L_1$-norm; Second, we establish convergence rates for the estimator with respect to several classes of dynamical systems under $L_1$-norm. In the analysis, the density function $f$ is only assumed to be Hölder continuous or pointwise Hölder controllable which is a weak assumption in the literature of nonparametric density estimation and also more realistic in the dynamical system context. Last but not least, we prove that the same convergence rates of the estimator under $L_\infty$-norm and $L_1$-norm can be achieved when the density function is Hölder continuous, compactly supported, and bounded. The bandwidth selection problem of the kernel density estimator for dynamical system is also discussed in our study via numerical simulations. Hanyuan Hang, Ingo Steinwart, Yunlong Feng, Johan A. K. Suykens |
J. Mach. Learn. Res. | 4 |
| 2018 | Indefinite kernel spectral learning
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens |
Pattern Recognit. | 3 |
| 2018 | Convex Formulation for Kernel PCA and Its Use in Semisupervised LearningabstractIn this brief, kernel principal component analysis (KPCA) is reinterpreted as the solution to a convex optimization problem. Actually, there is a constrained convex problem for each principal component, so that the constraints guarantee that the principal component is indeed a solution, and not a mere saddle point. Although these insights do not imply any algorithmic improvement, they can be used to further understand the method, formulate possible extensions, and properly address them. As an example, a new convex optimization problem for semisupervised classification is proposed, which seems particularly well suited whenever the number of known labels is small. Our formulation resembles a least squares support vector machine problem with a regularization parameter multiplied by a negative sign, combined with a variational principle for KPCA. Our primal optimization principle for semisupervised learning is solved in terms of the Lagrange multipliers. Numerical experiments in several classification tasks illustrate the performance of the proposed model in problems with only a few labeled data. Carlos M. Alaíz, Michaël Fanuel, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Parallelized Tensor Train Learning of Polynomial ClassifiersabstractIn pattern classification, polynomial classifiers are well-studied methods as they are capable of generating complex decision surfaces. Unfortunately, the use of multivariate polynomials is limited to kernels as in support-vector machines, because polynomials quickly become impractical for high-dimensional problems. In this paper, we effectively overcome the curse of dimensionality by employing the tensor train (TT) format to represent a polynomial classifier. Based on the structure of TTs, two learning algorithms are proposed, which involve solving different optimization problems of low computational complexity. Furthermore, we show how both regularization to prevent overfitting and parallelization, which enables the use of large training sets, are incorporated into these methods. The efficiency and efficacy of our tensor-based polynomial classifier are then demonstrated on the two popular data sets U.S. Postal Service and Modified NIST. Zhongming Chen, Kim Batselier, Johan A. K. Suykens, Ngai Wong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Classification With Truncated $\ell _{1}$ Distance KernelabstractThis brief proposes a truncated distance (TL1) kernel, which results in a classifier that is nonlinear in the global region but is linear in each subregion. With this kernel, the subregion structure can be trained using all the training data and local linear classifiers can be established simultaneously. The TL1 kernel has good adaptiveness to nonlinearity and is suitable for problems which require different nonlinearities in different areas. Though the TL1 kernel is not positive semidefinite, some classical kernel learning methods are still applicable which means that the TL1 kernel can be directly used in standard toolboxes by replacing the kernel evaluation. In numerical experiments, the TL1 kernel with a pregiven parameter achieves similar or better performance than the radial basis function kernel with the parameter tuned by cross validation, implying the TL1 kernel a promising nonlinear kernel for classification tasks. Xiaolin Huang, Johan A. K. Suykens, Shuning Wang, Joachim Hornegger, Andreas K. Maier |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Regularized Semipaired Kernel CCA for Domain AdaptationabstractDomain adaptation learning is one of the fundamental research topics in pattern recognition and machine learning. This paper introduces a regularized semipaired kernel canonical correlation analysis formulation for learning a latent space for the domain adaptation problem. The optimization problem is formulated in the primal-dual least squares support vector machine setting where side information can be readily incorporated through regularization terms. The proposed model learns a joint representation of the data set across different domains by solving a generalized eigenvalue problem or linear system of equations in the dual. The approach is naturally equipped with out-of-sample extension property, which plays an important role for model selection. Furthermore, the Nyström approximation technique is used to make the computational issues due to the large size of the matrices involved in the eigendecomposition feasible. The learned latent space of the source domain is fed to a multiclass semisupervised kernel spectral clustering model that can learn from both labeled and unlabeled data points of the source domain in order to classify the data instances of the target domain. Experimental results are given to illustrate the effectiveness of the proposed approaches on synthetic and real-life data sets. Siamak Mehrkanoon, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Moving Least Squares Support Vector Machines for weather temperature prediction
Zahra Karevan, Yunlong Feng, Johan A. K. Suykens |
ESANN | 3 |
| 2017 | Scalable Hybrid Deep Neural Kernel Networks
Siamak Mehrkanoon, Andreas Zell, Johan A. K. Suykens |
ESANN | 3 |
| 2017 | Probabilistic matrix factorization from quantized measurementsabstractWe consider the problem of factorizing a matrix with discrete-valued entries as a product of two low-rank matrices. Under a probabilistic framework, we seek for the minimum mean-square error estimates of these matrices, using full Bayes and empirical Bayes approaches. In the first case, we devise an integration scheme based on the Gibbs sampler that accounts also for hyperparameter and noise variance estimation. A similar technique is used also for the latter case, where we combine Gibbs sampling with the expectation-maximization (EM) algorithm to estimate the model parameters via marginal likelihood maximization. Extension to the case of missing values is also discussed. The proposed methods are evaluated on simulated data, and on a real data set for recommender systems. Giulio Bottegal, Johan A. K. Suykens |
IJCNN | 2 |
| 2017 | Multi-view LS-SVM regression for black-box temperature prediction in weather forecastingabstractIn multi-view regression, we have a regression problem where the input data can be represented in multiple ways. These different representations are called views. The aim of multi-view regression is to increase the performance of using only one view by taking into account the information available from all views. In this paper, we introduce a novel multi-view regression model called Multi-View Least Squares Support Vector Machines (MV LS-SVM) regression. This model is formulated in the primal-dual setting typical to Least Squares Support Vector Machines (LS-SVM) where a coupling term is introduced in the primal objective. This form of coupling allows for some degree of freedom to model the different representations while being able to incorporate the information from all views in the training phase. This work was motivated by the challenge of predicting temperature in weather forecasting. Black-box weather forecasting deals with a large number of observations and features and is one of the most challenging learning task around. In order to predict the temperature in a city, the historical data from that city as well as from the neighboring cities are taking into account. In the past, the data for different cities were usually simply concatenated. In this work, we use MV LS-SVM to do temperature prediction by regarding each city as a different view. Experimental results on the minimum and maximum temperature prediction in Brussels, show the improvement of the multi-view method with regard to previous work and that this technique is competitive to the existing state-of-the-art methods in weather prediction. Lynn Houthuys, Zahra Karevan, Johan A. K. Suykens |
IJCNN | 3 |
| 2017 | Fast kernel spectral clustering
Rocco Langone, Johan A. K. Suykens |
Neurocomputing | 2 |
| 2017 | Supervised aggregated feature learning for multiple instance classification
Rocco Langone, Johan A. K. Suykens |
Inf. Sci. | 2 |
| 2017 | Deep Restricted Kernel Machines Using Conjugate Feature DualityabstractThe aim of this letter is to propose a theory of deep restricted kernel machines offering new foundations for deep learning with kernel machines. From the viewpoint of deep learning, it is partially related to restricted Boltzmann machines, which are characterized by visible and hidden units in a bipartite graph without hidden-to-hidden connections and deep learning extensions as deep belief networks and deep Boltzmann machines. From the viewpoint of kernel machines, it includes least squares support vector machines for classification and regression, kernel principal component analysis (PCA), matrix singular value decomposition, and Parzen-type models. A key element is to first characterize these kernel machines in terms of so-called conjugate feature duality, yielding a representation with visible and hidden units. It is shown how this is related to the energy form in restricted Boltzmann machines, with continuous variables in a nonprobabilistic setting. In this new framework of so-called restricted kernel machine (RKM) representations, the dual variables correspond to hidden features. Deep RKM are obtained by coupling the RKMs. The method is illustrated for deep RKM, consisting of three levels with a least squares support vector machine regression level and two kernel PCA levels. In its primal form also deep feedforward neural networks can be trained within this framework. Johan A. K. Suykens |
Neural Comput. | 1 |
| 2017 | Editorial: A Successful Year and Looking Forward to 2017 and BeyondabstractThis issue marks the first anniversary issue since I was honored to serve as the Editor-in-Chief (EiC) of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS). I am happy to report that we had a very successful year and here are a few highlights that I would like to share with the community.•The latest impact factor of TNNLS is 4.854 according to the Journal Citation Reports. This marks a record high impact factor for our journal and places TNNLS as the number one scholarly publication in Computer Science (Hardware & Architecture), number three in Computer Science (Theory & Methods), and number ten in Electrical and Electronic Engineering journals. Haibo He, Barbara Hammer, Daniel W. C. Ho, Fakhri Karray, Dhireesha Kudithipudi, José Antonio Lozano 0001, Teresa Bernarda Ludermir, Jacek Mandziuk, Stefano Melacci, Antonio Paiva, Hong Qiao, Alain Rakotomamonjy, Shiliang Sun, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 16 |
| 2017 | Solution Path for Pin-SVM Classifiers With Positive and Negative τ ValuesabstractApplying the pinball loss in a support vector machine (SVM) classifier results in pin-SVM. The pinball loss is characterized by a parameter τ . Its value is related to the quantile level and different τ values are suitable for different problems. In this paper, we establish an algorithm to find the entire solution path for pin-SVM with different τ values. This algorithm is based on the fact that the optimal solution to pin-SVM is continuous and piecewise linear with respect to τ . We also show that the nonnegativity constraint on τ is not necessary, i.e., τ can be extended to negative values. First, in some applications, a negative τ leads to better accuracy. Second, τ = -1 corresponds to a simple solution that links SVM and the classical kernel rule. The solution for τ = -1 can be obtained directly and then be used as a starting point of the solution path. The proposed method efficiently traverses τ values through the solution path, and then achieves good performance by a suitable τ . In particular, τ = 0 corresponds to C-SVM, meaning that the traversal algorithm can output a result at least as good as C-SVM with respect to validation error. Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Efficient multiple scale kernel classifiersabstractWhile kernel methods using a single Gaussian kernel have proven to be very successful for nonlinear classification, in case of learning problems with a more complex underlying structure it is often desirable to use a linear combination of kernels with different widths. To address this issue, this paper presents a classification algorithm based on a jointly convex constrained optimization formulation. The primal problem is defined as jointly learning a combination of kernel classification models formulated in different feature spaces, which account for various representations or scales. The solution can be found by either solving a system of linear equations in case of equal combination weights or by means of a block coordinate descent scheme. The dual model is represented by a classifier using multiple kernels in the decision function. Furthermore, time and space complexity are reduced by adopting a divide and conquer strategy and through the use of the Nyström approximation of the eigenfunctions. Several experiments show the effectiveness of the proposed algorithms in dealing with datasets containing up to millions of instances. Rocco Langone, Johan A. K. Suykens |
IEEE BigData | 2 |
| 2016 | Efficient Sparse Approximation of Support Vector Machines Solving a Kernel Lasso
Marcelo Aliquintuy, Emanuele Frandi, Ricardo Ñanculef, Johan A. K. Suykens |
CIARP | 4 |
| 2016 | Clustering from two data sources using a kernel-based approach with weight coupling
Lynn Houthuys, Rocco Langone, Johan A. K. Suykens |
ESANN | 3 |
| 2016 | Spatio-temporal feature selection for black-box weather forecasting
Zahra Karevan, Johan A. K. Suykens |
ESANN | 2 |
| 2016 | Fast in-memory spectral clustering using a fixed-size approach
Rocco Langone, Raghvendra Mall, Vilen Jumutc, Johan A. K. Suykens |
ESANN | 4 |
| 2016 | Clustering-based feature selection for black-box weather temperature predictionabstractReliable weather forecasting is one of the challenging tasks that deals with a large number of observations and features. In this paper, a data-driven modeling technique is proposed for temperature prediction. To investigate local learning, Soft Kernel Spectral Clustering (SKSC) is used to find similar samples to the test point to be used for training. Due to the high dimensionality, Elastic net is employed as a feature selection approach. Features are selected in each cluster independently and then, Least Squares Support Vector Machines (LS-SVM) regression is used to learn the data. Finally, the predicted values by LS-SVMs are averaged based on the membership of the test point to each cluster. In the experimental results, the performance of the proposed method and “Weather underground” are compared and it is shown that the data-driven technique is competitive with the existing weather temperature prediction sites. For the case study, the prediction of the temperature in Brussels is considered. Zahra Karevan, Johan A. K. Suykens |
IJCNN | 2 |
| 2016 | Denoised Kernel Spectral data ClusteringabstractKernel Spectral Clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. It builds an unsupervised model on a small subset of data using the dual solution of the optimization problem. This allows KSC to have a powerful out-of-sample extension property leading to good cluster generalization w.r.t. unseen data points. However, in the presence of noise that causes overlapping data, the technique often fails to provide good generalization capability. In this paper, we propose a two-step process for clustering noisy data. We first denoise the data using kernel principal component analysis (KPCA) with a recently proposed Model selection criterion based on point-wise Distance Distributions (MDD) to obtain the underlying information in the data. We then use the KSC technique on this denoised data to obtain good quality clusters. One advantage of model based techniques is that we can use the same training and validation set for denoising and for clustering. We discovered that using the same kernel bandwidth parameter obtained from MDD for KPCA works efficiently with KSC in combination with the optimal number of clusters k to produce good quality clusters. We compare the proposed approach with normal KSC and KSC with KPCA using a heuristic method based on reconstruction error for several synthetic and real-world datasets to showcase the effectiveness of the proposed approach. Raghvendra Mall, Halima Bensmail, Rocco Langone, Carolina Varon, Johan A. K. Suykens |
IJCNN | 5 |
| 2016 | Multi-label semi-supervised learning using regularized kernel spectral clusteringabstractOften in real-world applications such as web page categorization, automatic image annotations and protein function prediction, each instance is associated with multiple labels (categories) simultaneously. In addition, due to the labeling cost one usually deals with a large amount of unlabeled data while the fraction of labeled data points will typically be small. In this paper, we propose a multi-label semi-supervised kernel spectral clustering learning algorithm that learns from both labeled and unlabeled instances. The kernel spectral clustering algorithm (KSC) serves as a core model and the information of labeled data points is integrated into the model via regularization terms. The propagation of the multiple labels to unlabeled data points is achieved by incorporating the mutual correlation between (similarity across) labels as well as encouraging the model output to be as close as possible to the given ground-truth of the labeled data points. Thanks to the Nyström approximation method, an explicit feature map is constructed and the optimization problem is solved in the primal. Experimental results demonstrate the effectiveness of the proposed approaches on real multi-label datasets. Siamak Mehrkanoon, Johan A. K. Suykens |
IJCNN | 2 |
| 2016 | Estimating the unknown time delay in chemical processes
Siamak Mehrkanoon, Yuri A. W. Shardt, Johan A. K. Suykens, Steven X. Ding |
Eng. Appl. Artif. Intell. | 3 |
| 2016 | Reweighted stochastic learning
Vilen Jumutc, Johan A. K. Suykens |
Neurocomputing | 2 |
| 2016 | Fast and scalable Lasso via stochastic Frank-Wolfe methods with a convergence guarantee
Emanuele Frandi, Ricardo Ñanculef, Stefano Lodi, Claudio Sartori 0001, Johan A. K. Suykens |
Mach. Learn. | 5 |
| 2016 | Kernelized Elastic Net Regularization: Generalization Bounds, and Sparse RecoveryabstractKernelized elastic net regularization (KENReg) is a kernelization of the well-known elastic net regularization (Zou & Hastie, 2005). The kernel in KENReg is not required to be a Mercer kernel since it learns from a kernelized dictionary in the coefficient space. Feng, Yang, Zhao, Lv, and Suykens (2014) showed that KENReg has some nice properties including stability, sparseness, and generalization. In this letter, we continue our study on KENReg by conducting a refined learning theory analysis. This letter makes the following three main contributions. First, we present refined error analysis on the generalization performance of KENReg. The main difficulty of analyzing the generalization error of KENReg lies in characterizing the population version of its empirical target function. We overcome this by introducing a weighted Banach space associated with the elastic net regularization. We are then able to conduct elaborated learning theory analysis and obtain fast convergence rates under proper complexity and regularity assumptions. Second, we study the sparse recovery problem in KENReg with fixed design and show that the kernelization may improve the sparse recovery ability compared to the classical elastic net regularization. Finally, we discuss the interplay among different properties of KENReg that include sparseness, stability, and generalization. We show that the stability of KENReg leads to generalization, and its sparseness confidence can be derived from generalization. Moreover, KENReg is stable and can be simultaneously sparse, which makes it attractive theoretically and practically. Yunlong Feng, Shao-Gao Lv, Hanyuan Hang, Johan A. K. Suykens |
Neural Comput. | 4 |
| 2016 | Robust Support Vector Machines for Classification with Nonconvex and Smooth LossesabstractThis letter addresses the robustness problem when learning a large margin classifier in the presence of label noise. In our study, we achieve this purpose by proposing robustified large margin support vector machines. The robustness of the proposed robust support vector classifiers (RSVC), which is interpreted from a weighted viewpoint in this work, is due to the use of nonconvex classification losses. Besides the robustness, we also show that the proposed RSCV is simultaneously smooth, which again benefits from using smooth classification losses. The idea of proposing RSVC comes from M-estimation in statistics since the proposed robust and smooth classification losses can be taken as one-sided cost functions in robust statistics. Its Fisher consistency property and generalization ability are also investigated. Besides the robustness and smoothness, another nice property of RSVC lies in the fact that its solution can be obtained by solving weighted squared hinge loss-based support vector machine problems iteratively. We further show that in each iteration, it is a quadratic programming problem in its dual space and can be solved by using state-of-the-art methods. We thus propose an iteratively reweighted type algorithm and provide a constructive proof of its convergence to a stationary point. Effectiveness of the proposed classifiers is verified on both artificial and real data sets. Yunlong Feng, Xiaolin Huang, Siamak Mehrkanoon, Johan A. K. Suykens |
Neural Comput. | 5 |
| 2016 | Learning Theory Estimates with Observations from General Stationary Stochastic ProcessesabstractThis letter investigates the supervised learning problem with observations drawn from certain general stationary stochastic processes. Here by general, we mean that many stationary stochastic processes can be included. We show that when the stochastic processes satisfy a generalized Bernstein-type inequality, a unified treatment on analyzing the learning schemes with various mixing processes can be conducted and a sharp oracle inequality for generic regularized empirical risk minimization schemes can be established. The obtained oracle inequality is then applied to derive convergence rates for several learning schemes such as empirical risk minimization (ERM), least squares support vector machines (LS-SVMs) using given generic kernels, and SVMs using gaussian kernels for both least squares and quantile regression. It turns out that for independent and identically distributed (i.i.d.) processes, our learning rates for ERM recover the optimal rates. For non-i.i.d. processes, including geometrically [Formula: see text]-mixing Markov processes, geometrically [Formula: see text]-mixing processes with restricted decay, [Formula: see text]-mixing processes, and (time-reversed) geometrically [Formula: see text]-mixing processes, our learning rates for SVMs with gaussian kernels match, up to some arbitrarily small extra term in the exponent, the optimal rates. For the remaining cases, our rates are at least close to the optimal rates. As a by-product, the assumed generalized Bernstein-type inequality also provides an interpretation of the so-called effective number of observations for various mixing processes. Hanyuan Hang, Yunlong Feng, Ingo Steinwart, Johan A. K. Suykens |
Neural Comput. | 4 |
| 2016 | Coordinate Descent Algorithm for Ramp Loss Linear Programming Support Vector Machines
Xiangming Xi, Xiaolin Huang, Johan A. K. Suykens, Shuning Wang |
Neural Process. Lett. | 3 |
| 2016 | Efficient evolutionary spectral clustering
Rocco Langone, Marc Van Barel, Johan A. K. Suykens |
Pattern Recognit. Lett. | 3 |
| 2016 | Robust Gradient Learning With ApplicationsabstractThis paper addresses the robust gradient learning (RGL) problem. Gradient learning models aim at learning the gradient vector of some target functions in supervised learning problems, which can be further used to applications, such as variable selection, coordinate covariance estimation, and supervised dimension reduction. However, existing GL models are not robust to outliers or heavy-tailed noise. This paper provides an RGL framework to address this problem in both regression and classification. This is achieved by introducing a robust regression loss function and proposing a robust classification loss. Moreover, our RGL algorithm works in an instance-based kernelized dictionary instead of some fixed reproducing kernel Hilbert space, which may provide more flexibility. To solve the proposed nonconvex model, a simple computational algorithm based on gradient descent is provided and the convergence of the proposed method is also analyzed. We then apply the proposed RGL model to applications, such as nonlinear variable selection and coordinate covariance estimation. The efficiency of our proposed model is verified on both synthetic and real data sets. Yunlong Feng, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Robust Low-Rank Tensor Recovery With Regularized Redescending M-EstimatorabstractThis paper addresses the robust low-rank tensor recovery problems. Tensor recovery aims at reconstructing a low-rank tensor from some linear measurements, which finds applications in image processing, pattern recognition, multitask learning, and so on. In real-world applications, data might be contaminated by sparse gross errors. However, the existing approaches may not be very robust to outliers. To resolve this problem, this paper proposes approaches based on the regularized redescending M-estimators, which have been introduced in robust statistics. The robustness of the proposed approaches is achieved by the regularized redescending M-estimators. However, the nonconvexity also leads to a computational difficulty. To handle this problem, we develop algorithms based on proximal and linearized block coordinate descent methods. By explicitly deriving the Lipschitz constant of the gradient of the data-fitting risk, the descent property of the algorithms is present. Moreover, we verify that the objective functions of the proposed approaches satisfy the Kurdyka-Łojasiewicz property, which establishes the global convergence of the algorithms. The numerical experiments on synthetic data as well as real data verify that our approaches are robust in the presence of outliers and still effective in the absence of outliers. Yunlong Feng, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Regularized and sparse stochastic k-means for distributed large-scale clusteringabstractIn this paper we present a novel clustering approach based on the stochastic learning paradigm and regularization with l1-norms. Our approach is an extension of the widely acknowledged K-Means algorithm. We introduce a simple regularized dual averaging scheme for learning prototype vectors (centroids) with l1-norms in a stochastic mode. In our approach we distribute the learning of individual prototype vectors for each cluster, and the re-assignment of cluster memberships is performed only for a fixed number of outer iterations. The latter approach is exactly the same as in original K-Means algorithm and aims at re-shuffling the pool of samples per cluster according to the learned centroids. We report an extended evaluation and comparison of our approach with respect to various clustering techniques like randomized K-Means and Proximal Plane Clustering. Our experimental studies indicate the usefulness of the proposed methods for obtaining better prototype vectors and corresponding cluster memberships while being able to perform feature selection by l1-norm minimization. Vilen Jumutc, Rocco Langone, Johan A. K. Suykens |
IEEE BigData | 3 |
| 2015 | Ranking Overlap and Outlier Points in Data using Soft Kernel Spectral Clustering
Raghvendra Mall, Rocco Langone, Johan A. K. Suykens |
ESANN | 3 |
| 2015 | A PARTAN-accelerated Frank-Wolfe algorithm for large-scale SVM classificationabstractFrank-Wolfe algorithms have recently regained the attention of the Machine Learning community. Their solid theoretical properties and sparsity guarantees make them a suitable choice for a wide range of problems in this field. In addition, several variants of the basic procedure exist that improve its theoretical properties and practical performance. In this paper, we investigate the application of some of these techniques to Machine Learning, focusing in particular on a Parallel Tangent (PARTAN) variant of the FW algorithm for SVM classification, which has not been previously suggested or studied for this type of problem. We provide experiments both in a standard setting and using a stochastic speed-up technique, showing that the considered algorithms obtain promising results on several medium and large-scale benchmark datasets. Emanuele Frandi, Ricardo Ñanculef, Johan A. K. Suykens |
IJCNN | 3 |
| 2015 | Black-box modeling for temperature prediction in weather forecastingabstractAccurate weather forecasting is one of most challenging tasks that deals with a large amount of observations and features. In this paper, a black-box modeling technique is proposed for temperature forecasting. Due to the high dimensionality of data, feature selection is done in two steps with k-Nearest Neighbors and Elastic net. Next, Least Squares Support Vector Machine regression is applied to generate the forecasting model. In the experimental results, the influence of each part of this procedure on the performance is investigated and compared with “Weather underground” results. For the case study, the prediction of the temperature in Brussels is considered. It is shown that black-box modeling has a good and competitive accuracy with current state-of-the-art methods for temperature prediction. Zahra Karevan, Siamak Mehrkanoon, Johan A. K. Suykens |
IJCNN | 3 |
| 2015 | Kernel spectral document clustering using unsupervised precision-recall metricsabstractKernel Spectral Clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. The KSC model is built on a small subset of data using a proper training, model selection and a test phase. The clustering model is obtained using the dual solution of the problem and has a powerful out-of-sample extensions property which allows cluster affiliation for previously unseen data points. In the model selection phase, we estimate the appropriate number of clusters using a metric that evaluates the quality of the clusters. Traditional quality indices like inertia, Davies-Bouldin (DB) index and silhouette (SIL) are known to be method-dependent and not perform well in case of complex heterogeneous data like textual data. In this paper, we utilize the quality evaluation techniques based on an unsupervised version of Precision, Recall and F-measure proposed in [1] to come up with a new kernel spectral document clustering (KSDC) model which generates homogeneous clusters of documents. We compare the quality of the clusters obtained by the proposed KSDC technique with k-means and neural gas algorithm, which are more oriented towards these metrics, on several real world textual data. Raghvendra Mall, Johan A. K. Suykens |
IJCNN | 2 |
| 2015 | Hierarchical semi-supervised clustering using KSC based modelabstractThis paper introduces a methodology to incorporate the label information in discovering the underlying clusters in a hierarchical setting using multi-class semi-supervised clustering algorithm. The method aims at revealing the relationship between clusters given few labels associated to some of the clusters. The problem is formulated as a regularized kernel spectral clustering algorithm in the primal-dual setting. The available labels are incorporated in different levels of hierarchy from top to bottom. As we advance towards the lowers levels in the tree all the previously added labels are used in the generation of the new levels of hierarchy. The model is trained on a subset of the data and then applied to the rest of the data in a learning framework. Thanks to the previously learned model, the out-of-sample extension property of the model allows then to predict the memberships of a new point. A combination of an internal clustering quality index and classification accuracy is used for model selection. Experiments are conducted on synthetic data and real image segmentation problems to show the applicability of the proposed approach. Siamak Mehrkanoon, Oscar Mauricio Agudelo, Raghvendra Mall, Johan A. K. Suykens |
IJCNN | 4 |
| 2015 | LS-SVM based spectral clustering and regression for predicting maintenance of industrial machines
Rocco Langone, Carlos Alzate, Bart De Ketelaere, Jonas Vlasselaer, Wannes Meert, Johan A. K. Suykens |
Eng. Appl. Artif. Intell. | 6 |
| 2015 | A robust ensemble approach to learn from positive and unlabeled data using SVM base models
Marc Claesen, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
Neurocomputing | 3 |
| 2015 | Sequential minimal optimization for SVM with pinball loss
Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens |
Neurocomputing | 3 |
| 2015 | Learning solutions to partial differential equations using LS-SVM
Siamak Mehrkanoon, Johan A. K. Suykens |
Neurocomputing | 2 |
| 2015 | Learning with the maximum correntropy criterion induced losses for regression
Yunlong Feng, Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens |
J. Mach. Learn. Res. | 5 |
| 2015 | Incremental multi-class semi-supervised clustering regularized by Kalman filtering
Siamak Mehrkanoon, Oscar Mauricio Agudelo, Johan A. K. Suykens |
Neural Networks | 3 |
| 2015 | Identifying intervals for hierarchical clustering using the Gershgorin circle theorem
Raghvendra Mall, Siamak Mehrkanoon, Johan A. K. Suykens |
Pattern Recognit. Lett. | 3 |
| 2015 | Two-level ℓ1 minimization for compressed sensing
Xiaolin Huang, Yipeng Liu 0001, Lei Shi 0010, Sabine Van Huffel, Johan A. K. Suykens |
Signal Process. | 5 |
| 2015 | Signal recovery for jointly sparse vectors with different sensing matrices
Li Li 0013, Xiaolin Huang, Johan A. K. Suykens |
Signal Process. | 3 |
| 2015 | A Rank-One Tensor Updating Algorithm for Tensor Completionabstract© 1994-2012 IEEE. In this letter, we propose a rank-one tensor updating algorithm for solving tensor completion problems. Unlike the existing methods which penalize the tensor by using the sum of nuclear norms of unfolding matrices, our optimization model directly employs the tensor nuclear norm which is studied recently. Under the framework of the conditional gradient method, we show that at each iteration, solving the proposed model amounts to computing the tensor spectral norm and the related rank-one tensor. Because the problem of finding the related rank-one tensor is NP-hard, we propose a subroutine to solve it approximately, which is of low computational complexity. Experimental results on real datasets show that our algorithm is efficient and effective. Yunlong Feng, Johan A. K. Suykens |
IEEE Signal Process. Lett. | 3 |
| 2015 | Very Sparse LSSVM Reductions for Large-Scale DataabstractLeast squares support vector machines (LSSVMs) have been widely applied for classification and regression with comparable performance with SVMs. The LSSVM model lacks sparsity and is unable to handle large-scale data due to computational and memory constraints. A primal fixed-size LSSVM (PFS-LSSVM) introduce sparsity using Nyström approximation with a set of prototype vectors (PVs). The PFS-LSSVM model solves an overdetermined system of linear equations in the primal. However, this solution is not the sparsest. We investigate the sparsity-error tradeoff by introducing a second level of sparsity. This is done by means of L0 -norm-based reductions by iteratively sparsifying LSSVM and PFS-LSSVM models. The exact choice of the cardinality for the initial PV set is not important then as the final model is highly sparse. The proposed method overcomes the problem of memory constraints and high computational costs resulting in highly sparse reductions to LSSVM models. The approximations of the two models allow to scale the models to large-scale datasets. Experiments on real-world classification and regression data sets from the UCI repository illustrate that these approaches achieve sparse models without a significant tradeoff in errors. Raghvendra Mall, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Multiclass Semisupervised Learning Based Upon Kernel Spectral ClusteringabstractThis paper proposes a multiclass semisupervised learning algorithm by using kernel spectral clustering (KSC) as a core model. A regularized KSC is formulated to estimate the class memberships of data points in a semisupervised setting using the one-versus-all strategy while both labeled and unlabeled data points are present in the learning process. The propagation of the labels to a large amount of unlabeled data points is achieved by adding the regularization terms to the cost function of the KSC formulation. In other words, imposing the regularization term enforces certain desired memberships. The model is then obtained by solving a linear system in the dual. Furthermore, the optimal embedding dimension is designed for semisupervised clustering. This plays a key role when one deals with a large number of clusters. Siamak Mehrkanoon, Carlos Alzate, Raghvendra Mall, Rocco Langone, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2015 | Noise Level Estimation for Model Selection in Kernel PCA DenoisingabstractOne of the main challenges in unsupervised learning is to find suitable values for the model parameters. In kernel principal component analysis (kPCA), for example, these are the number of components, the kernel, and its parameters. This paper presents a model selection criterion based on distance distributions (MDDs). This criterion can be used to find the number of components and the σ(2) parameter of radial basis function kernels by means of spectral comparison between information and noise. The noise content is estimated from the statistical moments of the distribution of distances in the original dataset. This allows for a type of randomization of the dataset, without actually having to permute the data points or generate artificial datasets. After comparing the eigenvalues computed from the estimated noise with the ones from the input dataset, information is retained and maximized by a set of model parameters. In addition to the model selection criterion, this paper proposes a modification to the fixed-size method and uses the incomplete Cholesky factorization, both of which are used to solve kPCA in large-scale applications. These two approaches, together with the model selection MDD, were tested in toy examples and real life applications, and it is shown that they outperform other known algorithms. Carolina Varon, Carlos Alzate, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Representative subsets for big data learning using k-NN graphsabstractIn this paper we propose a deterministic method to obtain subsets from big data which are a good representative of the inherent structure in the data. We first convert the large scale dataset into a sparse undirected k-NN graph using a distributed network generation framework that we propose in this paper. After obtaining the k-NN graph we exploit the fast and unique representative subset (FURS) selection method [1], [2] to deterministically obtain a subset for this big data network. The FURS selection technique selects nodes from different dense regions in the graph retaining the natural community structure. We then locate the points in the original big data corresponding to the selected nodes and compare the obtained subset with subsets acquired from state-of-the-art subset selection techniques. We evaluate the quality of the selected subset on several synthetic and real-life datasets for different learning tasks including big data classification and big data clustering. Raghvendra Mall, Vilen Jumutc, Rocco Langone, Johan A. K. Suykens |
IEEE BigData | 4 |
| 2014 | New bilinear formulation to semi-supervised classification based on Kernel Spectral ClusteringabstractIn this paper we present a novel semi-supervised classification approach which combines bilinear formulation for non-parallel binary classifiers based upon Kernel Spectral Clustering. The cornerstone of our approach is a bilinear term introduced into the primal formulation of semi-supervised classification problem. In addition we perform separate manifold regularization for each individual classifier. The latter relates to the Kernel Spectral Clustering unsupervised counterpart which helps to obtain more precise and generalizable classification boundaries. We derive the dual problem which can be effectively translated into a linear system of equations and then solved without introducing extra costs. In our experiments we show the usefulness and report considerable improvements in performance with respect to other semi-supervised approaches, like Laplacian SVMs and other KSC-based models. Vilen Jumutc, Johan A. K. Suykens |
CIDM | 2 |
| 2014 | Alarm prediction in industrial machines using autoregressive LS-SVM modelsabstractIn industrial machines different alarms are embedded in machines controllers. They make use of sensors and machine states to indicate to end-users various information (e.g. diagnostics or need of maintenance) or to put machines in a specific mode (e.g. shut-down when thermal protection is activated). More specifically, the alarms are often triggered based on comparing sensors data to a threshold defined in the controllers software. In batch production machines, triggering an alarm (e.g. thermal protection) in the middle of a batch production is crucial for the quality of the produced batch and results into a high production loss. This situation can be avoided if the settings of the production machine (e.g. production speed) is adjusted accordingly based on the temperature monitoring. Therefore, predicting a temperature alarm and adjusting the production speed to avoid triggering the alarm seems logical. In this paper we show the effectiveness of Least Squares Support Vector Machines (LS-SVMs) in predicting the evolution of the temperature in a steel production machine and, as a consequence, possible alarms due to overheating. Firstly, in an offline fashion, we develop a nonlinear autoregressive (NAR) model, where a systematic model selection procedure allows to carefully tune the model parameters. Afterwards, the NAR model is used online to forecast the future temperature trend. Finally, a classifier which uses as input the outcomes of the NAR model allows to foresee future alarms. Rocco Langone, Carlos Alzate, Abdellatif Bey-Temsamani, Johan A. K. Suykens |
CIDM | 4 |
| 2014 | Clustering data over time using kernel spectral clustering with memoryabstractThis paper discusses the problem of clustering data changing over time, a research domain that is attracting increasing attention due to the increased availability of streaming data in the Web 2.0 era. In the analysis conducted throughout the paper we make use of the kernel spectral clustering with memory (MKSC) algorithm, which is developed in a constrained optimization setting. Since the objective function of the MKSC model is designed to explicitly incorporate temporal smoothness, the algorithm belongs to the family of evolutionary clustering methods. Experiments over a number of real and synthetic datasets provide very interesting insights in the dynamics of the clusters evolution. Specifically, MKSC is able to handle objects leaving and entering over time, and recognize events like continuing, shrinking, growing, splitting, merging, dissolving and forming of clusters. Moreover, we discover how one of the regularization constants of the MKSC model, referred as the smoothness parameter, can be used as a change indicator measure. Finally, some possible visualizations of the cluster dynamics are proposed. Rocco Langone, Raghvendra Mall, Johan A. K. Suykens |
CIDM | 3 |
| 2014 | Agglomerative hierarchical kernel spectral data clusteringabstractIn this paper we extend the agglomerative hierarchical kernel spectral clustering (AH-KSC [1]) technique from networks to datasets and images. The kernel spectral clustering (KSC) technique builds a clustering model in a primal-dual optimization framework. The dual solution leads to an eigen-decomposition. The clustering model consists of kernel evaluations, projections onto the eigenvectors and a powerful out-of-sample extension property. We first estimate the optimal model parameters using the balanced angular fitting (BAF) [2] criterion. We then exploit the eigen-projections corresponding to these parameters to automatically identify a set of increasing distance thresholds. These distance thresholds provide the clusters at different levels of hierarchy in the dataset which are merged in an agglomerative fashion as shown in [1], [4]. We showcase the effectiveness of the AH-KSC method on several datasets and real world images. We compare the AH-KSC method with several agglomerative hierarchical clustering techniques and overcome the issues of hierarchical KSC technique proposed in [5]. Raghvendra Mall, Rocco Langone, Johan A. K. Suykens |
CIDM | 3 |
| 2014 | Reweighted l1 Dual Averaging Approach for Sparse Stochastic Learning
Vilen Jumutc, Johan A. K. Suykens |
ESANN | 2 |
| 2014 | Agglomerative hierarchical kernel spectral clustering for large scale networks
Raghvendra Mall, Rocco Langone, Johan A. K. Suykens |
ESANN | 3 |
| 2014 | Optimal Data Projection for Kernel Spectral Clustering
Diego Hernán Peluffo-Ordóñez, Carlos Alzate, Johan A. K. Suykens, Germán Castellanos-Domínguez |
ESANN | 3 |
| 2014 | SVD truncation schemes for fixed-size kernel modelsabstractIn this paper, two schemes for reducing the effective number of parameters are presented. To do this, different versions of Fixed-Size Kernel models based on Fixed-Size Least Squares Support Vector Machines (FS-LSSVM) are employed. The schemes include Fixed-Size Ordinary Least Squares (FS-OLS) and Fixed-Size Ridge Regression (FS-RR) with their respective truncations through Singular Value Decomposition (SVD). When these schemes are applied to the Silverbox and Wiener-Hammerstein data sets in system identification, it was found that a great deal of the complexity of the model could be reduced in a trade-off with the generalization performance. Ricardo Castro-Garcia, Siamak Mehrkanoon, Anna Marconato, Johan Schoukens, Johan A. K. Suykens |
IJCNN | 5 |
| 2014 | Optimal reduced sets for sparse kernel spectral clusteringabstractKernel spectral clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. It results in a clustering model using the dual solution of the problem. It has a powerful out-of-sample extension property leading to good clustering generalization w.r.t. the unseen data points. The out-of-sample extension property allows to build a sparse model on a small training set and introduces the first level of sparsity. The clustering dual model is expressed in terms of non-sparse kernel expansions where every point in the training set contributes. The goal is to find reduced set of training points which can best approximate the original solution. In this paper a second level of sparsity is introduced in order to reduce the time complexity of the computationally expensive out-of-sample extension. In this paper we investigate various penalty based reduced set techniques including the Group Lasso, L0, L1+ L0penalization and compare the amount of sparsity gained w.r.t. a previous L1penalization technique. We observe that the optimal results in terms of sparsity corresponds to the Group Lasso penalization technique in majority of the cases. We showcase the effectiveness of the proposed approaches on several real world datasets and an image segmentation dataset. Raghvendra Mall, Siamak Mehrkanoon, Rocco Langone, Johan A. K. Suykens |
IJCNN | 4 |
| 2014 | Large scale semi-supervised learning using KSC based modelabstractOften in practice one deals with a large amount of unlabeled data, while the fraction of labeled data points will typically be small. Therefore one prefers to apply a semi-supervised algorithm, which uses both labeled and unlabeled data points in the learning process, to have a better performance. Considering the large amount of unlabeled data, making a semi-supervised algorithm scalable is an important task. In this paper we adopt a recently proposed multi-class semi-supervised KSC based algorithm (MSS-KSC) and make it scalable by means of two different approaches. The first one is based on the Nyström approximation method which provides a finite dimensional feature map that can then be used to solve the optimization problem in the primal. The second approach is based on the reduced kernel technique that solves the problem in the dual by reducing the dimensionality of the kernel matrix to a rectangular kernel. Experimental results demonstrate the scalability and efficiency of the proposed approaches on real datasets. Siamak Mehrkanoon, Johan A. K. Suykens |
IJCNN | 2 |
| 2014 | Reweighted l 2-Regularized Dual Averaging Approach for Highly Sparse Stochastic Learning
Vilen Jumutc, Johan A. K. Suykens |
ISNN | 2 |
| 2014 | Predicting breast cancer using an expression values weighted clinical classifierabstractBACKGROUND: Clinical data, such as patient history, laboratory analysis, ultrasound parameters-which are the basis of day-to-day clinical decision support-are often used to guide the clinical management of cancer in the presence of microarray data. Several data fusion techniques are available to integrate genomics or proteomics data, but only a few studies have created a single prediction model using both gene expression and clinical data. These studies often remain inconclusive regarding an obtained improvement in prediction performance. To improve clinical management, these data should be fully exploited. This requires efficient algorithms to integrate these data sets and design a final classifier. LS-SVM classifiers and generalized eigenvalue/singular value decompositions are successfully used in many bioinformatics applications for prediction tasks. While bringing up the benefits of these two techniques, we propose a machine learning approach, a weighted LS-SVM classifier to integrate two data sources: microarray and clinical parameters. RESULTS: We compared and evaluated the proposed methods on five breast cancer case studies. Compared to LS-SVM classifier on individual data sets, generalized eigenvalue decomposition (GEVD) and kernel GEVD, the proposed weighted LS-SVM classifier offers good prediction performance, in terms of test area under ROC Curve (AUC), on all breast cancer case studies. CONCLUSIONS: Thus a clinical classifier weighted with microarray data set results in significantly improved diagnosis, prognosis and prediction responses to therapy. The proposed model has been shown as a promising mathematical framework in both data fusion and non-linear classification problems. Minta Thomas, Kris De Brabanter, Johan A. K. Suykens, Bart De Moor |
BMC Bioinform. | 3 |
| 2014 | Incremental kernel spectral clustering for online learning of non-stationary data
Rocco Langone, Oscar Mauricio Agudelo, Bart De Moor, Johan A. K. Suykens |
Neurocomputing | 4 |
| 2014 | Non-parallel support vector classifiers with different loss functions
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens |
Neurocomputing | 3 |
| 2014 | QoS prediction for web service compositions using kernel-based quantile estimation with online adaptation of the constant offset
Dries Geebelen, Kristof Geebelen, Eddy Truyen, Sam Michiels, Johan A. K. Suykens, Joos Vandewalle, Wouter Joosen |
Inf. Sci. | 5 |
| 2014 | EnsembleSVM: a library for ensemble learning using support vector machines
Marc Claesen, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
J. Mach. Learn. Res. | 3 |
| 2014 | Ramp loss linear programming support vector machine
Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens |
J. Mach. Learn. Res. | 3 |
| 2014 | Learning with tensors: a framework based on convex optimization and spectral regularization
Marco Signoretto, Quoc Tran-Dinh, Lieven De Lathauwer, Johan A. K. Suykens |
Mach. Learn. | 4 |
| 2014 | Support Vector Machine Classifier With Pinball LossabstractTraditionally, the hinge loss is used to construct support vector machine (SVM) classifiers. The hinge loss is related to the shortest distance between sets and the corresponding classifier is hence sensitive to noise and unstable for re-sampling. In contrast, the pinball loss is related to the quantile distance and the result is less sensitive. The pinball loss has been deeply studied and widely applied in regression but it has not been used for classification. In this paper, we propose a SVM classifier with the pinball loss, called pin-SVM, and investigate its properties, including noise insensitivity, robustness, and misclassification error. Besides, insensitive zone is applied to the pin-SVM for a sparse model. Compared to the SVM with the hinge loss, the proposed pin-SVM has the same computational complexity and enjoys noise insensitivity and re-sampling stability. Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Multi-Class Supervised Novelty DetectionabstractIn this paper we study the problem of finding a support of unknown high-dimensional distributions in the presence of labeling information, called Supervised Novelty Detection (SND). The One-Class Support Vector Machine (SVM) is a widely used kernel-based technique to address this problem. However with the latter approach it is difficult to model a mixture of distributions from which the support might be constituted. We address this issue by presenting a new class of SVM-like algorithms which help to approach multi-class classification and novelty detection from a new perspective. We introduce a new coupling term between classes which leverages the problem of finding a good decision boundary while preserving the compactness of a support with the l2-norm penalty. First we present our optimization objective in the primal and then derive a dual QP formulation of the problem. Next we propose a Least-Squares formulation which results in a linear system which drastically reduces computational costs. Finally we derive a Pegasos-based formulation which can effectively cope with large data sets that cannot be handled by many existing QP solvers. We complete our paper with experiments that validate the usefulness and practical importance of the proposed methods both in classification and novelty detection settings. Vilen Jumutc, Johan A. K. Suykens |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Self-tuned kernel spectral clustering for large scale networksabstractWe propose a parameter-free kernel spectral clustering model for large scale complex networks. The kernel spectral clustering (KSC) method works by creating a model on a subgraph of the complex network. The model requires a kernel function which can have parameters and the number of communities k has be detected in the large scale network. We exploit the structure of the projections in the eigenspace to automatically identify the number of clusters. We use the concept of entropy and balanced clusters for this purpose. We show the effectiveness of the proposed approach by comparing the cluster memberships w.r.t. several large scale community detection techniques like Louvain, Infomap and Bigclam methods. We conducted experiments on several synthetic networks of varying size and mixing parameter along with large scale real world experiments to show the efficiency of the proposed approach. Raghvendra Mall, Rocco Langone, Johan A. K. Suykens |
IEEE BigData | 3 |
| 2013 | Supervised Novelty DetectionabstractIn this paper we present a novel approach and a new machine learning problem, called Supervised Novelty Detection (SND). This problem extends the One-Class Support Vector Machine setting for binary classification while keeping the nice properties of novelty detection problem at hand. To tackle this we approach binary classification from a new perspective using two different estimators and a coupled regularization term. It involves optimization over a different objective and a doubled set of Lagrange multipliers. One might consider our approach as a joint estimation of the support for different probability distributions per class where an ultimate goal is to separate classes with the largest possible angle between the normal vectors to the decision hyperplanes in the feature space. Regarding an obvious novelty of our problem we report and compare the results along the lines of standard C-SVM, LS-SVM and One-Class SVM. Experiments have demonstrated promising results that validate the usefulness of the proposed method. Vilen Jumutc, Johan A. K. Suykens |
CIDM | 2 |
| 2013 | Kernel spectral clustering for predicting maintenance of industrial machinesabstractEarly and accurate fault detection in modern industrial machines is crucial in order to minimize downtime, increase the safety of plant operations, and reduce manufacturing costs. The process monitoring techniques that have been most effective in practice are based on the analysis of historical process data. In this paper we present a novel approach that uses Kernel Spectral Clustering (KSC) on the sensor data to distinguish between normal operating condition and abnormal situations. In other words, the main contribution is to show how KSC can be a valid tool also for outlier detection, a field where other techniques are more popular. KSC is a state-of-the-art unsupervised learning technique with out-of-sample ability and a systematic model selection scheme. Thanks to the abovementioned characteristics and the capability of discovering complex clustering boundaries, KSC is able to detect in advance the need of maintenance actions in the analyzed machine. Rocco Langone, Carlos Alzate, Bart De Ketelaere, Johan A. K. Suykens |
CIDM | 4 |
| 2013 | DynOpt: Incorporating dynamics into mean-variance portfolio optimizationabstractMean-variance (MV) portfolio theory leads to relatively simple and elegant numerical problems. Nonetheless, the approach has been criticized for treating the market parameters as if they were constant over time. We propose a novel convex optimization problem that extends an existing MV formulation with chance constraint(s) by accounting for the portfolio dynamics. The core idea is to consider a multiperiod scenario where portfolio weights are implicitly regarded as the output of a state-space dynamical system driven by external inputs. The approach leverages a result on realization theory and uses the nuclear norm to penalize complex dynamical behaviors. The proposed ideas are illustrated by two case studies. Marco Signoretto, Johan A. K. Suykens |
CIFEr | 2 |
| 2013 | Fixed-size Pegasos for hinge and pinball loss SVMabstractPegasos has become a widely acknowledged algorithm for learning linear Support Vector Machines. It utilizes properties of hinge loss and theory of strongly convex optimization problems for fast convergence rates and lower computational and memory costs. In this paper we adopt the recently proposed pinball loss for the Pegasos algorithm and show some advantages of using it in a variety of classification problems. First we present the newly derived Pegasos optimization objective with respect to pinball loss and analyze its properties and convergence rates. Additionally we present extensions of the Pegasos algorithm applied to the kernel-induced and Nyström approximated feature maps which introduce non-linearity in the input space. This is done using a Fixed-Size kernel method approach. Second we give experimental results for publicly available UCI datasets to justify the advantages and the importance of pinball loss for achieving a better classification accuracy and greater numerical stability in the partially or fully stochastic setting. Finally we conclude our paper with a brief discussion of the applicability of pinball loss to real-life problems. Vilen Jumutc, Xiaolin Huang, Johan A. K. Suykens |
IJCNN | 3 |
| 2013 | Soft kernel spectral clusteringabstractIn this paper we propose an algorithm for soft (or fuzzy) clustering. In soft clustering each point is not assigned to a single cluster (like in hard clustering), but it can belong to every cluster with a different degree of membership. Generally speaking, this property is desirable in order to improve the interpretability of the results. Our starting point is a state-of-the art technique called kernel spectral clustering (KSC). Instead of using the hard assignment method present therein, we suggest a fuzzy assignment based on the cosine distance from the cluster prototypes. We then call the new method soft kernel spectral clustering (SKSC). We also introduce a related model selection technique, called average membership strength criterion, which solves the drawbacks of the previously proposed method (namely balanced linefit). We apply the new algorithm to synthetic and real datasets, for image segmentation and community detection on networks. We show that in many cases SKSC outperforms KSC. Rocco Langone, Raghvendra Mall, Johan A. K. Suykens |
IJCNN | 3 |
| 2013 | Non-parallel semi-supervised classification based on kernel spectral clusteringabstractIn this paper, a non-parallel semi-supervised algorithm based on kernel spectral clustering is formulated. The prior knowledge about the labels is incorporated into the kernel spectral clustering formulation via adding regularization terms. In contrast with the existing multi-plane classifiers such as Multisurface Proximal Support Vector Machine (GEPSVM) and Twin Support Vector Machines (TWSVM) and its least squares version (LSTSVM) we will not use a kernel-generated surface. Instead we apply the kernel trick in the dual. Therefore as opposed to conventional non-parallel classifiers one does not need to formulate two different primal problems for the linear and nonlinear case separately. The proposed method will generate two non-parallel hyperplanes which then are used for out-of-sample extension. Experimental results demonstrate the efficiency of the proposed method over existing methods. Siamak Mehrkanoon, Johan A. K. Suykens |
IJCNN | 2 |
| 2013 | Kernel spectral clustering for dynamic data using multiple kernel learningabstractIn this paper we propose a kernel spectral clustering-based technique to catch the different regimes experienced by a time-varying system. Our method is based on a multiple kernel learning approach, which is a linear combination of kernels. The calculation of the linear combination coefficients is done by determining a ranking vector that quantifies the overall dynamical behavior of the analyzed data sequence over-time. This vector can be calculated from the eigenvectors provided by the the solution of the kernel spectral clustering problem. We apply the proposed technique to a trial from the Graphics Lab Motion Capture Database from Carnegie Mellon University, as well as to a synthetic example, namely three moving Gaussian clouds. For comparison purposes, some conventional spectral clustering techniques are also considered, namely, kernel k-means and min-cuts. Also, standard k-means. The normalized mutual information and adjusted random index metrics are used to quantify the clustering performance. Results show the usefulness of proposed technique to track dynamic data, even being able to detect hidden objects. Diego Hernán Peluffo-Ordóñez, Sergio García-Vega, Rocco Langone, Johan A. K. Suykens, Germán Castellanos-Domínguez |
IJCNN | 4 |
| 2013 | Sparse Reductions for Fixed-Size Least Squares Support Vector Machines on Large Scale Data
Raghvendra Mall, Johan A. K. Suykens |
PAKDD (1) | 2 |
| 2013 | The skweezee system: enabling the design and the programming of squeeze interactionsabstractThe Skweezee System is an easy, flexible and open system for designing and developing squeeze-based, gestural interactions. It consists of Skweezees, which are soft objects, filled with conductive padding, that can be deformed or squeezed by applying pressure. These objects contain a number of electrodes that are dispersed over the shape. The electrodes sense the shape shifting of the conductive filling by measuring the changing resistance between every possible pair of electrodes. In addition, the Skweezee System contains user-friendly software that allows end-users to define and to record their own squeeze gestures. These gestures are distinguished using a Support Vector Machine (SVM) classifier. In this paper we introduce the concept and the underlying technology of the Skweezee System and we demonstrate the robustness of the SVM based classifier via two experimental user studies. The results of these studies demonstrate accuracies of 81% (8 gestures, user-defined) to 97% (3 gestures, user-defined), with an accuracy of 90% for 7 pre-defined gestures. Karen Vanderloock, Vero Vanden Abeele, Johan A. K. Suykens, Lucca Geurts |
UIST | 3 |
| 2013 | Load forecasting using a multivariate meta-learning system
Marin Matijas, Johan A. K. Suykens, Slavko Krajcar |
Expert Syst. Appl. | 2 |
| 2013 | Risk group detection and survival function estimation for interval coded survival methods
Vanya Van Belle, Patrick Neven, Vernon Harvey, Sabine Van Huffel, Johan A. K. Suykens, Stephen P. Boyd |
Neurocomputing | 5 |
| 2013 | Support vector machines with piecewise linear feature mapping
Xiaolin Huang, Siamak Mehrkanoon, Johan A. K. Suykens |
Neurocomputing | 3 |
| 2013 | Hinging Hyperplanes for Time-Series SegmentationabstractDivision of a time series into segments is a common technique for time-series processing, and is known as segmentation. Segmentation is traditionally done by linear interpolation in order to guarantee the continuity of the reconstructed time series. The interpolation-based segmentation methods may perform poorly for data with a level of noise because interpolation is noise sensitive. To handle the problem, this paper establishes an explicit expression for segmentation from a compact representation for piecewise linear functions using hinging hyperplanes. This expression enables the use of regression to obtain a continuous reconstructed signal and, as a consequence, application of advanced techniques in segmentation. In this paper, a least squares support vector machine with lasso using a hinging feature map is given and analyzed, based on which a segmentation algorithm and its online version are established. Numerical experiments conducted on synthetic and real-world datasets demonstrate the advantages of our methods compared to existing segmentation algorithms. Xiaolin Huang, Marin Matijas, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2012 | Interval coded scoring systems for survival analysis
Vanya Van Belle, Sabine Van Huffel, Johan A. K. Suykens, Stephen P. Boyd |
ESANN | 3 |
| 2012 | Joint Regression and Linear Combination of Time Series for Optimal Prediction
Dries Geebelen, Kim Batselier, Philippe Dreesen, Marco Signoretto, Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
ESANN | 5 |
| 2012 | A semi-supervised formulation to binary kernel spectral clusteringabstractA semi-supervised formulation to binary kernel spectral clustering is presented. The formulation fits in a constrained optimization setting with primal and dual model representations. The clustering model can be applied naturally to out-of-sample points allowing model selection and achieving good generalization capabilities. The proposed method incorporates labeled information into the core binary kernel spectral clustering by adding an extra term into the objective function together with a regularization constant. The resulting dual problem is no longer an eigenvalue problem as in the case of the original core model but a linear system. A model selection criterion combining a cluster distortion measure on the unlabeled part and the classification accuracy on the labeled part is also presented. This criterion can be used to obtain clustering parameters such that the clustering model evaluated at validation points display a desirable structure. Simulation results with toy data and real benchmark datasets show the applicability of the proposed method. Carlos Alzate, Johan A. K. Suykens |
IJCNN | 2 |
| 2012 | Robustness of kernel based regression: Influence and weight functionsabstractIt has been shown that kernel based regression (KBR) with a least squares loss has some undesirable properties from robustness point of view. KBR with more robust loss functions, e.g. Huber or Logistic losses, often give rise to more complicated computations. In classical statistics, robustness is improved by reweighting the original estimate. We study the influence of reweighting the LS-KBR estimate using three well-known weight functions and one new weight function called Myriad. Our results give practical guidelines in order to choose the weights, providing robustness and fast convergence. It turns out that Logistic and Myriad weights are suitable reweighting schemes when outliers are present in the data. In fact, the Myriad shows better performance over the others in the presence of extreme outliers (e.g. Cauchy distributed errors). These findings are then illustrated on toy example as well as on a real life data sets. Finally, we establish an empirical maxbias curve to demonstrate the ability of the proposed methodology. Kris De Brabanter, Jos De Brabanter, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
IJCNN | 3 |
| 2012 | Kernel spectral clustering for community detection in complex networksabstractThis paper is related to community detection in complex networks. We show the use of kernel spectral clustering for the analysis of unweighted networks. We employ the primal-dual framework and make use of out-of-sample extension. In the latter the assignment rule for the new nodes is based on a model learned in the training phase. We propose a method to extract from a network a small subgraph representative for its overall community structure. We use a model selection procedure based on the modularity statistic which is novel, because modularity is commonly used only at a training level. We demonstrate the effectiveness of our model on synthetic networks and benchmark data from real networks (power grid network and protein interaction network of yeast). Finally, we compare our model with the Nyström method, showing that our approach is better in terms of quality of the discovered partitions and needs less computation time. Rocco Langone, Carlos Alzate, Johan A. K. Suykens |
IJCNN | 3 |
| 2012 | Towards the detection of error-related potentials and its integration in the context of a P300 speller brain-computer interface
Adrien Combaz, Nikolay Chumerin, Nikolay V. Manyakov, Arne Robben, Johan A. K. Suykens, Marc M. Van Hulle |
Neurocomputing | 5 |
| 2012 | Hierarchical kernel spectral clustering
Carlos Alzate, Johan A. K. Suykens |
Neural Networks | 2 |
| 2012 | Optimized Data Fusion for Kernel k-Means ClusteringabstractThis paper presents a novel optimized kernel k-means algorithm (OKKC) to combine multiple data sources for clustering analysis. The algorithm uses an alternating minimization framework to optimize the cluster membership and kernel coefficients as a nonconvex problem. In the proposed algorithm, the problem to optimize the cluster membership and the problem to optimize the kernel coefficients are all based on the same Rayleigh quotient objective; therefore the proposed algorithm converges locally. OKKC has a simpler procedure and lower complexity than other algorithms proposed in the literature. Simulated and real-life data fusion applications are experimentally studied, and the results validate that the proposed algorithm has comparable performance, moreover, it is more efficient on large-scale data sets. (The Matlab implementation of OKKC algorithm is downloadable from http://homes.esat.kuleuven.be/~sistawww/bio/syu/okkc.html.). Léon-Charles Tranchevent, Xinhai Liu, Wolfgang Glänzel, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2012 | Confidence bands for least squares support vector machine classifiers: A regression approach
Kris De Brabanter, Peter Karsmakers, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Pattern Recognit. | 4 |
| 2012 | Reducing the Number of Support Vectors of SVM Classifiers Using the Smoothed Separable Case ApproximationabstractIn this brief, we propose a new method to reduce the number of support vectors of support vector machine (SVM) classifiers. We formulate the approximation of an SVM solution as a classification problem that is separable in the feature space. Due to the separability, the hard-margin SVM can be used to solve it. This approach, which we call the separable case approximation (SCA), is very similar to the cross-training algorithm explained in , which is inspired by editing algorithms . The norm of the weight vector achieved by SCA can, however, become arbitrarily large. For that reason, we propose an algorithm, called the smoothed SCA (SSCA), that additionally upper-bounds the weight vector of the pruned solution and, for the commonly used kernels, reduces the number of support vectors even more. The lower the chosen upper bound, the larger this extra reduction becomes. Upper-bounding the weight vector is important because it ensures numerical stability, reduces the time to find the pruned solution, and avoids overfitting during the approximation phase. On the examined datasets, SSCA drastically reduces the number of support vectors. Dries Geebelen, Johan A. K. Suykens, Joos Vandewalle |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Approximate Solutions to Ordinary Differential Equations Using Least Squares Support Vector MachinesabstractIn this paper, a new approach based on least squares support vector machines (LS-SVMs) is proposed for solving linear and nonlinear ordinary differential equations (ODEs). The approximate solution is presented in closed form by means of LS-SVMs, whose parameters are adjusted to minimize an appropriate error function. For the linear and nonlinear cases, these parameters are obtained by solving a system of linear and nonlinear equations, respectively. The method is well suited to solving mildly stiff, nonstiff, and singular ODEs with initial and boundary conditions. Numerical results demonstrate the efficiency of the proposed method over existing methods. Siamak Mehrkanoon, Tillmann Falck, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2011 | Sparse LS-SVMs with L0 - norm minimization
Jorge López Lázaro, Kris De Brabanter, José R. Dorronsoro, Johan A. K. Suykens |
ESANN | 4 |
| 2011 | Symbolic computing of LS-SVM based models
Siamak Mehrkanoon, Carlos Alzate, Johan A. K. Suykens |
ESANN | 4 |
| 2011 | Automatic Seizure Detection Incorporating Structural Information
Borbála Hunyadi, Maarten De Vos, Marco Signoretto, Johan A. K. Suykens, Wim Van Paesschen, Sabine Van Huffel |
ICANN (1) | 4 |
| 2011 | Out-of-sample eigenvectors in kernel spectral clusteringabstractA method to estimate eigenvectors for out-of-sample data in the context of kernel spectral clustering is presented. The proposed method is within a constrained optimization framework with primal and dual model representations. This formulation allows the clustering model to be extended naturally to out-of-sample points together with the possibility to perform model selection in a learning setting. A model selection methodology based on the Fisher criterion is also presented. The proposed criterion can be used to select clustering parameters such that the out-of-sample eigenvector space show a desirable structure. This special structure appears when the clusters are well-formed and the clustering parameters have been chosen properly. Simulation results with toy examples and images show the applicability of the proposed method and model selection criterion. Carlos Alzate, Johan A. K. Suykens |
IJCNN | 2 |
| 2011 | Modularity-based model selection for kernel spectral clusteringabstractA proper way of choosing the tuning parameters in a kernel model has a fundamental importance in determining the success of the model for a particular task. This paper is related to model selection in the framework of community detection on weighted and unweighted networks by means of a kernel spectral clustering model. Here we propose a new method based on Modularity (a popular measure of community structure in a network) which can deal with quite general situations (i.e. overlapping communities with different sizes). Thus we use Modularity criterion for model selection and not at the training level, which is the case of all the clustering algorithms proposed so far in the literature. Rocco Langone, Carlos Alzate, Johan A. K. Suykens |
IJCNN | 3 |
| 2011 | Support vector methods for survival analysis: a comparison between ranking and regression approaches
Vanya Van Belle, Kristiaan Pelckmans, Sabine Van Huffel, Johan A. K. Suykens |
Artif. Intell. Medicine | 4 |
| 2011 | Improved performance on high-dimensional survival data by application of Survival-SVMabstractMOTIVATION: New application areas of survival analysis as for example based on micro-array expression data call for novel tools able to handle high-dimensional data. While classical (semi-) parametric techniques as based on likelihood or partial likelihood functions are omnipresent in clinical studies, they are often inadequate for modelling in case when there are less observations than features in the data. Support vector machines (svms) and extensions are in general found particularly useful for such cases, both conceptually (non-parametric approach), computationally (boiling down to a convex program which can be solved efficiently), theoretically (for its intrinsic relation with learning theory) as well as empirically. This article discusses such an extension of svms which is tuned towards survival data. A particularly useful feature is that this method can incorporate such additional structure as additive models, positivity constraints of the parameters or regression constraints. RESULTS: Besides discussion of the proposed methods, an empirical case study is conducted on both clinical as well as micro-array gene expression data in the context of cancer studies. Results are expressed based on the logrank statistic, concordance index and the hazard ratio. The reported performances indicate that the present method yields better models for high-dimensional data, while it gives results which are comparable to what classical techniques based on a proportional hazard model give for clinical data. Vanya Van Belle, Kristiaan Pelckmans, Sabine Van Huffel, Johan A. K. Suykens |
Bioinform. | 4 |
| 2011 | Optimized data fusion for K-means Laplacian clusteringabstractMOTIVATION: We propose a novel algorithm to combine multiple kernels and Laplacians for clustering analysis. The new algorithm is formulated on a Rayleigh quotient objective function and is solved as a bi-level alternating minimization procedure. Using the proposed algorithm, the coefficients of kernels and Laplacians can be optimized automatically. RESULTS: Three variants of the algorithm are proposed. The performance is systematically validated on two real-life data fusion applications. The proposed Optimized Kernel Laplacian Clustering (OKLC) algorithms perform significantly better than other methods. Moreover, the coefficients of kernels and Laplacians optimized by OKLC show some correlation with the rank of performance of individual data source. Though in our evaluation the K values are predefined, in practical studies, the optimal cluster number can be consistently estimated from the eigenspectrum of the combined kernel Laplacian matrix. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/oklc.html. Xinhai Liu, Léon-Charles Tranchevent, Wolfgang Glänzel, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
Bioinform. | 5 |
| 2011 | Sparse kernel spectral clustering models for large-scale data analysis
Carlos Alzate, Johan A. K. Suykens |
Neurocomputing | 2 |
| 2011 | Learning Transformation Models for Ranking and Survival Analysis
Vanya Van Belle, Kristiaan Pelckmans, Johan A. K. Suykens, Sabine Van Huffel |
J. Mach. Learn. Res. | 3 |
| 2011 | Kernel Regression in the Presence of Correlated Errors
Kris De Brabanter, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
J. Mach. Learn. Res. | 3 |
| 2011 | Sparse conjugate directions pursuit with application to fixed-size kernel modelsabstractThis work studies an optimization scheme for computing sparse approximate solutions of over-determined linear systems. Sparse Conjugate Directions Pursuit (SCDP) aims to construct a solution using only a small number of nonzero (i.e. nonsparse) coefficients. Motivations of this work can be found in a setting of machine learning where sparse models typically exhibit better generalization performance, lead to fast evaluations, and might be exploited to define scalable algorithms. The main idea is to build up iteratively a conjugate set of vectors of increasing cardinality, in each iteration solving a small linear subsystem. By exploiting the structure of this conjugate basis, an algorithm is found (i) converging in at most D iterations for D -dimensional systems, (ii) with computational complexity close to the classical conjugate gradient algorithm, and (iii) which is especially efficient when a few iterations suffice to produce a good approximation. As an example, the application of SCDP to Fixed-Size Least Squares Support Vector Machines (FS-LSSVM) is discussed resulting in a scheme which efficiently finds a good model size for the FS-LSSVM setting, and is scalable to large-scale machine learning tasks. The algorithm is empirically verified in a classification context. Further discussion includes algorithmic issues such as component selection criteria, computational analysis, influence of additional hyper-parameters, and determination of a suitable stopping criterion. Peter Karsmakers, Kristiaan Pelckmans, Kris De Brabanter, Hugo Van hamme, Johan A. K. Suykens |
Mach. Learn. | 5 |
| 2011 | A kernel-based framework to tensorial data analysis
Marco Signoretto, Lieven De Lathauwer, Johan A. K. Suykens |
Neural Networks | 3 |
| 2011 | First and Second Order SMO Algorithms for LS-SVM Classifiers
Jorge López Lázaro, Johan A. K. Suykens |
Neural Process. Lett. | 2 |
| 2011 | Tensor Versus Matrix Completion: A Comparison With Application to Spectral DataabstractTensor completion recently emerged as a generalization of matrix completion for higher order arrays. This problem formulation allows one to exploit the structure of data that intrinsically have multiple dimensions. In this work, we recall a convex formulation for minimum (multilinear) ranks completion of arrays of arbitrary order. Successively we focus on completion of partially observed spectral images; the latter can be naturally represented as third order tensors and typically exhibit intraband correlations. We compare different convex formulations and assess them through case studies. Marco Signoretto, Raf Van de Plas, Bart De Moor, Johan A. K. Suykens |
IEEE Signal Process. Lett. | 4 |
| 2011 | Approximate Confidence and Prediction Intervals for Least Squares Support Vector RegressionabstractBias-corrected approximate 100(1-α)% pointwise and simultaneous confidence and prediction intervals for least squares support vector machines are proposed. A simple way of determining the bias without estimating higher order derivatives is formulated. A variance estimator is developed that works well in the homoscedastic and heteroscedastic case. In order to produce simultaneous confidence intervals, a simple Šidák correction and a more involved correction (based on upcrossing theory) are used. The obtained confidence intervals are compared to a state-of-the-art bootstrap-based method. Simulations show that the proposed method obtains similar intervals compared to the bootstrap at a lower computational cost. Kris De Brabanter, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Neural Networks | 3 |
| 2010 | Highly sparse kernel spectral clustering with predictive out-of-sample extensions
Carlos Alzate, Johan A. K. Suykens |
ESANN | 2 |
| 2010 | On the use of a clinical kernel in survival analysis
Vanya Van Belle, Kristiaan Pelckmans, Johan A. K. Suykens, Sabine Van Huffel |
ESANN | 3 |
| 2010 | Kernel-Based Learning from Infinite Dimensional 2-Way Tensors
Marco Signoretto, Lieven De Lathauwer, Johan A. K. Suykens |
ICANN (2) | 3 |
| 2010 | Polynomial componentwise LS-SVM: Fast variable selection using low rank updatesabstractThis paper describes a Least Squares Support Vector Machines (LS-SVM) approach to estimate additive models as a sum of non-linear components. In particular, this work discusses the low rank matrix modifications for componentwise polynomial kernels, which allow the factors of the modified kernel-matrix to be directly updated. The main concept refers to the use of a valid explicit feature map for polynomial kernels in an additive setting. By exploiting the structure of such feature map the model parameters of the classification/regression problem can be easily modified and updated when new variables are added. Therefore, the low rank updates constitute an algorithmic tool to efficiently obtain the model parameters once the system has been altered in some minimal sense. Such strategy allows, for instance, the development of algorithms for sequential variable ranking in high dimensional settings, while non-linearity is provided by the polynomial feature map. Moreover relevant variables can be robustly ranked using the closed form of the leave-one-out (LOO) error estimator, obtained as a by-product of the low rank modifications. Fabian Ojeda, Tillmann Falck, Bart De Moor, Johan A. K. Suykens |
IJCNN | 4 |
| 2010 | L2-norm multiple kernel learning and its application to biomedical data fusionabstractBACKGROUND: This paper introduces the notion of optimizing different norms in the dual problem of support vector machines with multiple kernels. The selection of norms yields different extensions of multiple kernel learning (MKL) such as L(infinity), L1, and L2 MKL. In particular, L2 MKL is a novel method that leads to non-sparse optimal kernel coefficients, which is different from the sparse kernel coefficients optimized by the existing L(infinity) MKL method. In real biomedical applications, L2 MKL may have more advantages over sparse integration method for thoroughly combining complementary information in heterogeneous data sources. RESULTS: We provide a theoretical analysis of the relationship between the L2 optimization of kernels in the dual problem with the L2 coefficient regularization in the primal problem. Understanding the dual L2 problem grants a unified view on MKL and enables us to extend the L2 method to a wide range of machine learning problems. We implement L2 MKL for ranking and classification problems and compare its performance with the sparse L(infinity) and the averaging L1 MKL methods. The experiments are carried out on six real biomedical data sets and two large scale UCI data sets. L2 MKL yields better performance on most of the benchmark data sets. In particular, we propose a novel L2 MKL least squares support vector machine (LSSVM) algorithm, which is shown to be an efficient and promising classifier for large scale data sets processing. CONCLUSIONS: This paper extends the statistical framework of genomic data fusion based on MKL. Allowing non-sparse weights on the data sources is an attractive option in settings where we believe most data sources to be relevant to the problem at hand and want to avoid a "winner-takes-all" effect seen in L(infinity) MKL, which can be detrimental to the performance in prospective studies. The notion of optimizing L2 kernels can be straightforwardly extended to ranking, classification, regression, and clustering algorithms. To tackle the computational burden of MKL, this paper proposes several novel LSSVM based MKL algorithms. Systematic comparison on real data sets shows that LSSVM MKL has comparable performance as the conventional SVM MKL algorithms. Moreover, large scale numerical experiments indicate that when cast as semi-infinite programming, LSSVM MKL can be solved more efficiently than SVM MKL. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/l2lssvm.html. Tillmann Falck, Anneleen Daemen, Léon-Charles Tranchevent, Johan A. K. Suykens, Bart De Moor, Yves Moreau |
BMC Bioinform. | 5 |
| 2010 | Multiway Spectral Clustering with Out-of-Sample Extensions through Weighted Kernel PCAabstractA new formulation for multiway spectral clustering is proposed. This method corresponds to a weighted kernel principal component analysis (PCA) approach based on primal-dual least-squares support vector machine (LS-SVM) formulations. The formulation allows the extension to out-of-sample points. In this way, the proposed clustering model can be trained, validated, and tested. The clustering information is contained on the eigendecomposition of a modified similarity matrix derived from the data. This eigenvalue problem corresponds to the dual solution of a primal optimization problem formulated in a high-dimensional feature space. A model selection criterion called the Balanced Line Fit (BLF) is also proposed. This criterion is based on the out-of-sample extension and exploits the structure of the eigenvectors and the corresponding projections when the clusters are well formed. The BLF criterion can be used to obtain clustering parameters in a learning framework. Experimental results with difficult toy problems and image segmentation show improved performance in terms of generalization to new samples and computation times. Carlos Alzate, Johan A. K. Suykens |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Coupled Simulated AnnealingabstractWe present a new class of methods for the global optimization of continuous variables based on simulated annealing (SA). The coupled SA (CSA) class is characterized by a set of parallel SA processes coupled by their acceptance probabilities. The coupling is performed by a term in the acceptance probability function, which is a function of the energies of the current states of all SA processes. A particular CSA instance method is distinguished by the form of its coupling term and acceptance probability. In this paper, we present three CSA instance methods and compare them with the uncoupled case, i.e., multistart SA. The primary objective of the coupling in CSA is to create cooperative behavior via information exchange. This aim helps in the decision of whether uphill moves will be accepted. In addition, coupling can provide information that can be used online to steer the overall optimization process toward the global optimum. We present an example where we use the acceptance temperature to control the variance of the acceptance probabilities with a simple control scheme. This approach leads to much better optimization efficiency, because it reduces the sensitivity of the algorithm to initialization parameters while guiding the optimization process to quasioptimal runs. We present the results of extensive experiments and show that the addition of the coupling and the variance control leads to considerable improvements with respect to the uncoupled case and a more recently proposed distributed version of SA. Samuel Xavier de Souza, Johan A. K. Suykens, Joos Vandewalle, Désiré Bollé |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2009 | Differentiation between brain metastases and glioblastoma multiforme based on MRI, MRS and MRSIabstractBrain metastases and glioblastoma multiforme are the most aggressive and common brain tumours in adults and they require a different clinical management. Anatomical magnetic resonance imaging (MRI) or clinical history, cannot always clearly distinguish between them. This study describes and verifies the use of magnetic resonance spectroscopy (MRS) and magnetic resonance spectroscopic imaging (MRSI) in combination with MRI for differential diagnosis of glioblastomas and metastases. Feature selection methods are applied to the magnetic resonance (MR) spectra of 121 patients and relevant features are detected. Different classification methods are used to distinguish glioblastoma multiforme and metastasis based on the single-voxel MR spectra, but no reliable differentiation is obtained: the accuracy varies from 50 to 78%. Next, MRSI and MRI data from 10 patients (5 glioblastomas, 5 solitary metastases) are used for differentiation purposes. The combination of multivoxel MR data and MRI data suggests a more clear differentiation between glioblastoma multiforme and brain metastasis. The results are visualized based on nosologic images, which are generated by including spectroscopic information in the segmented MR image. The methodology offers a new way that may support clinicians in decision making. Jan Luts, Johan A. K. Suykens, Sabine Van Huffel, Teresa Laudadio, Sofie Van Cauter, Uwe Himmelreich, Enrique Molla, Jose Piquer, M. Carmen Martínez-Bisbal, Bernardo Celda |
CBMS | 2 |
| 2009 | Transductively Learning from Positive Examples Only
Kristiaan Pelckmans, Johan A. K. Suykens |
ESANN | 2 |
| 2009 | Identifying Customer Profiles in Power Load Time Series Using Spectral Clustering
Carlos Alzate, Marcelo Espinoza, Bart De Moor, Johan A. K. Suykens |
ICANN (2) | 4 |
| 2009 | MINLIP: Efficient Learning of Transformation Models
Vanya Van Belle, Kristiaan Pelckmans, Johan A. K. Suykens, Sabine Van Huffel |
ICANN (1) | 3 |
| 2009 | Robustness of Kernel Based Regression: A Comparison of Iterative Weighting Schemes
Kris De Brabanter, Kristiaan Pelckmans, Jos De Brabanter, Michiel Debruyne, Johan A. K. Suykens, Mia Hubert, Bart De Moor |
ICANN (1) | 5 |
| 2009 | Feature Extraction and Classification of EEG Signals for Rapid P300 Mind SpellingabstractThe Mind Speller is a Brain-Computer Interface which enables subjects to spell text on a computer screen by detecting P300 Event-Related Potentials in their electroencephalograms. This BCI application is of particular interest for disabled patients who have lost all means of verbal and motor communication. We report on the implementation of a feature extraction procedure on a new ultra low-power 8-channel wireless EEG device. The feature extraction procedure is based on downsampled EEG signal epochs, the Student's t-statistic of the Continuous Wavelet Transform, and the Common Spatial Pattern technique. For classification, we use a linear Least-Squares Support Vector Machine. The results show that subjects are potentially able to communicate a character in less than ten seconds with an accuracy of 94.5%, which is more than twice as fast as the state of the art. In addition since our EEG device is wireless it offers an increased comfort to the subject. Adrien Combaz, Nikolay V. Manyakov, Nikolay Chumerin, Johan A. K. Suykens, Marc M. Van Hulle |
ICMLA | 4 |
| 2009 | A regularized formulation for spectral clustering with pairwise constraintsabstractA regularized method to incorporate prior knowledge into spectral clustering in the form of pairwise constraints is proposed. This method is based on a weighted kernel principal component analysis (PCA) interpretation of spectral clustering with primal-dual least squares support vector machines (LS-SVM) formulations. The weighted kernel PCA framework allows incorporating pairwise constraints into the primal problem leading to a dual eigenvalue problem involving a modified kernel matrix. This modification on the metric is a regularized rank-1 downdate of the original kernel matrix. The clustering model can also be extended to out-of-sample points which becomes important for generalization, predictive purposes and large-scale data. An extension of an existing model selection criterion is also proposed. This extension introduces an additional term to the criterion measuring the constraint fit. Simulation results with toy examples and an image segmentation problem show the applicability of the proposed method. Carlos Alzate, Johan A. K. Suykens |
IJCNN | 2 |
| 2009 | Least conservative support and tolerance tubesabstractThis paper studies a distribution-free estimator of the conditional support and tolerance intervals of a distributions underlying a set of paired independent and identically distributed (i.i.d.) observations. The key ingredients are (a) an appropriate notion of risk which measures what probability mass is not captured by the estimate, (b) a uniform concentration inequality for the empirical risk based on a compression argument, and (c) the derivation of a lower bound to the mutual information, dictating how to maximize the informativeness of the estimator. For this result we extend Fano's inequality to the bivariate case. Kristiaan Pelckmans, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Inf. Theory | 3 |
| 2008 | Survival SVM: a practical scalable algorithm
Vanya Van Belle, Kristiaan Pelckmans, Johan A. K. Suykens, Sabine Van Huffel |
ESANN | 3 |
| 2008 | Quadratically Constrained Quadratic Programming for Subspace Selection in Kernel Regression Estimation
Marco Signoretto, Kristiaan Pelckmans, Johan A. K. Suykens |
ICANN (1) | 3 |
| 2008 | Sparse kernel models for spectral clustering using the incomplete Cholesky decompositionabstractA new sparse kernel model for spectral clustering is presented. This method is based on the incomplete Cholesky decomposition and can be used to efficiently solve large-scale spectral clustering problems. The formulation arises from a weighted kernel principal component analysis (PCA) interpretation of spectral clustering. The interpretation is within a constrained optimization framework with primal and dual model representations allowing the clustering model to be extended to out-of-sample points. The incomplete Cholesky decomposition is used to compute low-rank approximations of a modified affinity matrix derived from the data which contains cluster information. A reduced set method is also presented to compute efficiently the cluster indicators for out-of-sample data. Simulation results with large-scale toy datasets and images show improved performance in terms of computational complexity. Carlos Alzate, Johan A. K. Suykens |
IJCNN | 2 |
| 2008 | A regularized kernel CCA contrast function for ICA
Carlos Alzate, Johan A. K. Suykens |
Neural Networks | 2 |
| 2008 | Low rank updated LS-SVM classifiers for fast variable selection
Fabian Ojeda, Johan A. K. Suykens, Bart De Moor |
Neural Networks | 2 |
| 2008 | Kernel Component Analysis Using an Epsilon-Insensitive Robust Loss FunctionabstractKernel principal component analysis (PCA) is a technique to perform feature extraction in a high-dimensional feature space, which is nonlinearly related to the original input space. The kernel PCA formulation corresponds to an eigendecomposition of the kernel matrix: eigenvectors with large eigenvalues correspond to the principal components in the feature space. Starting from the least squares support vector machine (LS-SVM) formulation to kernel PCA, we extend it to a generalized form of kernel component analysis (KCA) with a general underlying loss function made explicit. For classical kernel PCA, the underlying loss function is L(2) . In this generalized form, one can plug in also other loss functions. In the context of robust statistics, it is known that the L(2) loss function is not robust because its influence function is not bounded. Therefore, outliers can skew the solution from the desired one. Another issue with kernel PCA is the lack of sparseness: the principal components are dense expansions in terms of kernel functions. In this paper, we introduce robustness and sparseness into kernel component analysis by using an epsilon-insensitive robust loss function. We propose two different algorithms. The first method solves a set of nonlinear equations with kernel PCA as starting points. The second method uses a simplified iterative weighting procedure that leads to solving a sequence of generalized eigenvalue problems. Simulations with toy and real-life data show improvements in terms of robustness together with a sparse representation. Carlos Alzate, Johan A. K. Suykens |
IEEE Trans. Neural Networks | 2 |
| 2008 | Data Visualization and Dimensionality Reduction Using Kernel Maps With a Reference PointabstractIn this paper, a new kernel-based method for data visualization and dimensionality reduction is proposed. A reference point is considered corresponding to additional constraints taken in the problem formulation. In contrast with the class of kernel eigenmap methods, the solution (coordinates in the low-dimensional space) is characterized by a linear system instead of an eigenvalue problem. The kernel maps with a reference point are generated from a least squares support vector machine (LS-SVM) core part that is extended with an additional regularization term for preserving local mutual distances together with reference point constraints. The kernel maps possess primal and dual model representations and provide out-of-sample extensions, e.g., for validation-based tuning. The method is illustrated on toy problems and real-life data sets. Johan A. K. Suykens |
IEEE Trans. Neural Networks | 1 |
| 2007 | Convex optimization for the design of learning machines
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ESANN | 2 |
| 2007 | Comparing Methods for Multi-class Probabilities in Medical Decision Making Using LS-SVMs and Kernel Logistic Regression
Ben Van Calster, Jan Luts, Johan A. K. Suykens, George Condous, Tom Bourne, Dirk Timmerman, Sabine Van Huffel |
ICANN (2) | 3 |
| 2007 | State-of-the-Art and Evolution in Public Data Sets and Competitions for System Identification, Time Series Prediction and Pattern RecognitionabstractIt is the aim of reproducible research to provide mechanisms for objective comparison of methods, algorithms, software and procedures in various research topics. In this paper, we discuss the role of data sets, benchmarks and competitions in the fields of system identification, time series prediction, classification, and pattern recognition in view of creating an environment of reproducible research. Important elements are the data sets, their origin, and the comparison measures that will be used to rank the performance of the methods. The issues are discussed, a comparison is made and recommendations are given. Joos Vandewalle, Johan A. K. Suykens, Bart De Moor, Amaury Lendasse |
ICASSP (4) | 2 |
| 2007 | ICA through an LS-SVM based Kernel CCA Measure for IndependenceabstractA new measure for independence based on canonical correlation in high dimensional feature spaces is presented. This measure can be used as a contrast function for independent component analysis (ICA). The formulation fits in the least squares support vector machines (LS-SVM) framework as a primal-dual interpretation of kernel canonical correlation analysis (CCA) in the context of constrained optimization problems. Regularization is incorporated naturally in the primal formulation leading to a dual generalized eigenvalue problem. Due to the primal-dual nature of the proposed approach, the measure for independence can be calculated for out-of-sample data points which is important for parameter selection ensuring statistical reliability of the estimated measure. Simulations results with small toy datasets performing model selection on a validation set showed good performance avoiding overfltting. Experiments with image demixing using approximated kernel matrices via incomplete Cholesky decomposition showed good results together with a reduced computational cost. Carlos Alzate, Johan A. K. Suykens |
IJCNN | 2 |
| 2007 | Multi-class kernel logistic regression: a fixed-size implementationabstractThis research studies a practical iterative algorithm for multi-class kernel logistic regression (KLR). Starting from the negative penalized log likelihood criterium we show that the optimization problem in each iteration can be solved by a weighted version of least squares support vector machines (LS-SVMs). In this derivation it turns out that the global regularization term is reflected as a usual regularization in each separate step. In the LS-SVM framework, fixed-size LS-SVM is known to perform well on large data sets. We therefore implement this model to solve large scale multi-class KLR problems with estimation in the primal space. To reduce the size of the Hessian, an alternating descent version of Newton's method is used which has the extra advantage that it can be easily used in a distributed computing environment. It is investigated how a multi-class kernel logistic regression model compares to a one-versus-all coding scheme. Peter Karsmakers, Kristiaan Pelckmans, Johan A. K. Suykens |
IJCNN | 3 |
| 2007 | Variable selection by rank-one updates for least squares support vector machinesabstractLeast squares support vector machines (LS-SVM) classifiers are a class of simple, yet powerful, kernel methods whose solution follows from a set of linear equations. Here, forward and backward algorithms, based on this technique, are proposed for fast and efficient variable selection. By exploiting the structure of the LS-SVM solution a closed form expression for the leave-one-out (LOO) estimator, useful for selecting variables, is obtained. For inclusion or removal of a new variable, rank-one adjustments in the kernel matrix (linear kernel) allow for updating, rather than recomputing, the LS-SVM solution. The proposed approach is applied to microarray data for gene selection. Simulations clearly show lower computational complexity along with good stability on the generalization performance when compared to other related algorithms. Fabian Ojeda, Johan A. K. Suykens, Bart De Moor |
IJCNN | 2 |
| 2007 | Fixed-size kernel logistic regression for phoneme classificationabstractKernel logistic regression (KLR) is a popular non-linear classification technique. Unlike an empirical risk minimization approach such as employed by Support Vector Machines (SVMs), KLR yields probabilistic outcomes based on a maximum likelihood argument which are particularly important in speech recognition. Different from other KLR implementations we use a Nyström approximation to solve large scale problems with estimation in the primal space such as done in fixed-size Least Squares Support Vector Machines (LS-SVMs). In the speech experiments it is investigated how a natural KLR extension to multi-class classification compares to binary KLR models coupled via a one-versus-one coding scheme. Moreover, a comparison to SVMs is made. Index Terms: phoneme classification, kernel logistic regression, large-scale, multi-class Peter Karsmakers, Kristiaan Pelckmans, Johan A. K. Suykens, Hugo Van hamme |
INTERSPEECH | 3 |
| 2007 | A Risk Minimization Principle for a Class of Parzen EstimatorsabstractThis paper explores the use of a Maximal Average Margin (MAM) optimality principle for the design of learning algorithms. It is shown that the application of this risk minimization principle results in a class of (computationally) simple learning machines similar to the classical Parzen window classifier. A direct relation with the Rademacher complexities is established, as such facilitating analysis and providing a notion of certainty of prediction. This analysis is related to Support Vector Machines by means of a margin transformation. The power of the MAM principle is illustrated further by application to ordinal regression tasks, resulting in an $O(n)$ algorithm able to process large datasets in reasonable time. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
NIPS | 2 |
| 2007 | A combined MRI and MRSI based multiclass system for brain tumour recognition using LS-SVMs with class probabilities and feature selection
Jan Luts, Arend Heerschap, Johan A. K. Suykens, Sabine Van Huffel |
Artif. Intell. Medicine | 3 |
| 2007 | Efficiently updating and tracking the dominant kernel principal components
Luc Hoegaerts, Lieven De Lathauwer, Ivan Goethals, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
Neural Networks | 4 |
| 2007 | Bagging Linear Sparse Bayesian Learning Models for Variable Selection in Cancer DiagnosisabstractThis paper investigates variable selection (VS) and classification for biomedical datasets with a small sample size and a very high input dimension. The sequential sparse Bayesian learning methods with linear bases are used as the basic VS algorithm. Selected variables are fed to the kernel-based probabilistic classifiers: Bayesian least squares support vector machines (BayLS-SVMs) and relevance vector machines (RVMs). We employ the bagging techniques for both VS and model building in order to improve the reliability of the selected variables and the predictive performance. This modeling strategy is applied to real-life medical classification problems, including two binary cancer diagnosis problems based on microarray data and a brain tumor multiclass classification problem using spectra acquired via magnetic resonance spectroscopy. The work is experimentally compared to other VS methods. It is shown that the use of bagging can improve the reliability and stability of both VS and model prediction. Chuan Lu, Andy Devos, Johan A. K. Suykens, Carles Arús, Sabine Van Huffel |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2007 | A Convex Approach to Validation-Based Learning of the Regularization ConstantabstractThis letter investigates a tight convex relaxation to the problem of tuning the regularization constant with respect to a validation based criterion. A number of algorithms is covered including ridge regression, regularization networks, smoothing splines, and least squares support vector machines (LS-SVMs) for regression. This convex approach allows the application of reliable and efficient tools, thereby improving computational cost and automatization of the learning method. It is shown that all solutions of the relaxation allow an interpretation in terms of a solution to a weighted LS-SVM. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
IEEE Trans. Neural Networks | 2 |
| 2006 | A Weighted Kernel PCA Formulation with Out-of-Sample Extensions for Spectral Clustering MethodsabstractA new formulation to spectral clustering methods based on the weighted kernel principal component analysis is presented. This formulation fits in the Least Squares Support Vector Machines (LS-SVM) framework as a primal-dual interpretation in the context of constrained optimization problems. Starting from the LS-SVM formulation to kernel PCA, a weighted approach is derived. An advantage of this method is the possibility to apply the trained clustering model to out-of-sample (test) data points without using approximation techniques such as the Nystrom method. Links with some existing spectral clustering techniques are given, showing that these techniques are particular cases of weighted kernel PCA. Simulation results with toy and real-life data show improvements in terms of generalization to new samples. Carlos Alzate, Johan A. K. Suykens |
IJCNN | 2 |
| 2006 | Multi-scroll and hypercube attractors from Josephson junctionsabstractIn this paper Josephson junctions are used in order to generate n-scroll and n-scroll hypercube attractors. We propose to use of the Josephson junction in a general Jerk circuit in such a way that there is no need for synthesizing the nonlinearity towards n-scroll and n-scroll hypercube attractors. The results are illustrated with computer simulations. Müstak E. Yalçin, Johan A. K. Suykens, Joos Vandewalle |
ISCAS | 2 |
| 2006 | A process model to develop an internal rating system: Sovereign credit ratings
Tony Van Gestel, Bart Baesens, Peter Van Dijcke, Joao Garcia, Johan A. K. Suykens, Jan Vanthienen |
Decis. Support Syst. | 5 |
| 2006 | Additive Regularization Trade-Off: Fusion of Training and Validation Levels in Kernel Methods
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
Mach. Learn. | 2 |
| 2005 | Componentwise Support Vector Machines for Structure Detection
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ICANN (2) | 2 |
| 2005 | Extending kernel principal component analysis to general underlying loss functionsabstractKernel principal component analysis can be considered as a natural nonlinear generalization of PCA because it performs linear PCA in a kernel induced feature space. It allows us to extract nonlinear structures in the input data. The classical kernel PCA formulation leads to an eigendecomposition of the kernel matrix: eigenvectors with large eigenvalue correspond to the principal components in the feature space. Starting from the least squares support vector machine (LS-SVM) formulation to kernel PCA we extend it to general underlying loss functions. For classical kernel PCA, the underlying loss function is L/sub 2/. In this approach, one can easily plug in other loss functions and solve a nonlinear optimization problem to achieve desirable properties. Simulations with Huber's loss function for robustness and quadratic epsilon insensitive loss function for sparseness demonstrate the flexibility of our approach. Carlos Alzate, Johan A. K. Suykens |
IJCNN | 2 |
| 2005 | Maximal variation and missing values for componentwise support vector machinesabstractThis paper proposes primal-dual kernel machine classifiers based on worst-case analysis of a finite set of observations including missing values of the inputs. Key ingredients are the use of a componentwise support vector machine (cSVM) and an empirical measure of maximal variation of the components to bind the influence of the component which cannot be evaluated due to missing values. A regularization term based on the L/sub 1/ norm of the maximal variation is used to obtain a mechanism for structure detection in that context. An efficient implementation using the hierarchical kernel machines framework is elaborated. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor, Jos De Brabanter |
IJCNN | 2 |
| 2005 | M@CBETH: a microarray classification benchmarking toolabstractMicroarray classification can be useful to support clinical management decisions for individual patients in, for example, oncology. However, comparing classifiers and selecting the best for each microarray dataset can be a tedious and non-straightforward task. The M@CBETH (a MicroArray Classification BEnchmarking Tool on a Host server) web service offers the microarray community a simple tool for making optimal two-class predictions. M@CBETH aims at finding the best prediction among different classification methods by using randomizations of the benchmarking dataset. The M@CBETH web service intends to introduce an optimal use of clinical microarray data classification. Nathalie Pochet, Frizo A. L. Janssens, Frank De Smet, Kathleen Marchal, Johan A. K. Suykens, Bart De Moor |
Bioinform. | 5 |
| 2005 | Subset based least squares subspace regression in RKHS
Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
Neurocomputing | 2 |
| 2005 | The differogram: Non-parametric noise variance estimation and its use for model selection
Kristiaan Pelckmans, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Neurocomputing | 3 |
| 2005 | Building sparse representations and structure determination on LS-SVM substrates
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
Neurocomputing | 2 |
| 2005 | Handling missing values in support vector machine classifiers
Kristiaan Pelckmans, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Neural Networks | 3 |
| 2005 | Primal-Dual Monotone Kernel Regression
Kristiaan Pelckmans, Marcelo Espinoza, Jos De Brabanter, Johan A. K. Suykens, Bart De Moor |
Neural Process. Lett. | 4 |
| 2004 | Sparse LS-SVMs using additive regularization with a penalized validation criterion
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ESANN | 2 |
| 2004 | A Comparison of Pruning Algorithms for Sparse Least Squares Support Vector Machines
Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
ICONIP | 2 |
| 2004 | Morozov, Ivanov and Tikhonov Regularization Based LS-SVMs
Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
ICONIP | 2 |
| 2004 | Primal space sparse kernel partial least squares regression for large scale problemsabstractKernel based methods suffer from exceeding time and memory requirements when applied on large datasets since the involved optimization problems typically scale polynomially in the number of data samples. As a remedy we propose both working on a reduced set (for fast evaluation) and at the same time keeping the number of model parameters small (for fast training). Departing from the Nystrom based feature approximation we describe fixed-size least squares support vector machine in the context of primal space least squares regression, to extend it with a supervised counterpart, sparse kernel partial least squares. The model is illustrated on a large scale example. Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
IJCNN | 2 |
| 2004 | Regularization constants in LS-SVMs: a fast estimate via convex optimizationabstractThe tuning of the regularization constant in applications of least squares support vector machines (LS-SVMs) for regression and classification is considered. The formulation of the LS-SVM training and regularization constant tuning problem (w.r.t. the validation performance) is considered as a single constrained optimization problem. In the formulation with Tikhonov regularization the problem of estimation the weights, validation errors and the regularization constants is a non-convex problem. The main result of This work is a conversion of the nonlinear constraints into a set of linear constraints, which turns the problem into a convex one. This is done based upon a simple Nadaraya-Watson kernel estimator via approximating the LS-SVM smoother matrix by the Nadaraya-Watson smoother. The paper further illustrates how to use this initial estimate towards grid search or local search methods. Numerical examples show considerable speed-ups by the proposed method. Kristiaan Pelckmans, Johan A. K. Suykens, Bart De Moor |
IJCNN | 2 |
| 2004 | Brain tumor classification based on long echo proton MRS signals
Lukas, Andy Devos, Johan A. K. Suykens, Leentje Vanhamme, Franklyn A. Howe, Carles Majós, Àngel Moreno-Torres, M. Van Der Graaf, Anne Rosemary Tate, Carles Arús, Sabine Van Huffel |
Artif. Intell. Medicine | 3 |
| 2004 | Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reductionabstractMOTIVATION: Microarrays are capable of determining the expression levels of thousands of genes simultaneously. In combination with classification methods, this technology can be useful to support clinical management decisions for individual patients, e.g. in oncology. The aim of this paper is to systematically benchmark the role of non-linear versus linear techniques and dimensionality reduction methods. RESULTS: A systematic benchmarking study is performed by comparing linear versions of standard classification and dimensionality reduction techniques with their non-linear versions based on non-linear kernel functions with a radial basis function (RBF) kernel. A total of 9 binary cancer classification problems, derived from 7 publicly available microarray datasets, and 20 randomizations of each problem are examined. CONCLUSIONS: Three main conclusions can be formulated based on the performances on independent test sets. (1) When performing classification with least squares support vector machines (LS-SVMs) (without dimensionality reduction), RBF kernels can be used without risking too much overfitting. The results obtained with well-tuned RBF kernels are never worse and sometimes even statistically significantly better compared to results obtained with a linear kernel in terms of test set receiver operating characteristic and test set accuracy performances. (2) Even for classification with linear classifiers like LS-SVM with linear kernel, using regularization is very important. (3) When performing kernel principal component analysis (kernel PCA) before classification, using an RBF kernel for kernel PCA tends to result in overfitting, especially when using supervised feature selection. It has been observed that an optimal selection of a large number of features is often an indication for overfitting. Kernel PCA with linear kernel gives better results. Nathalie Pochet, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
Bioinform. | 3 |
| 2004 | Benchmarking Least Squares Support Vector Machine ClassifiersabstractIn Support Vector Machines (SVMs), the solution of the classification problem is characterized by a (convex) quadratic programming (QP) problem. In a modified version of SVMs, called Least Squares SVM classifiers (LS-SVMs), a least squares cost function is proposed so as to obtain a linear set of equations in the dual space. While the SVM classifier has a large margin interpretation, the LS-SVM formulation is related in this paper to a ridge regression approach for classification with binary targets and to Fisher's linear discriminant analysis in the feature space. Multiclass categorization problems are represented by a set of binary classifiers using different output coding schemes. While regularization is used to control the effective number of parameters of the LS-SVM classifier, the sparseness property of SVMs is lost due to the choice of the 2-norm. Sparseness can be imposed in a second stage by gradually pruning the support value spectrum and optimizing the hyperparameters during the sparse approximation procedure. In this paper, twenty public domain benchmark datasets are used to evaluate the test set performance of LS-SVM classifiers with linear, polynomial and radial basis function (RBF) kernels. Both the SVM and LS-SVM classifier with RBF kernel in combination with standard cross-validation procedures for hyperparameter selection achieve comparable test set performances. These SVM and LS-SVM performances are consistently very good when compared to a variety of methods described in the literature including decision tree based algorithms, statistical algorithms and instance based learning methods. We show on ten UCI datasets that the LS-SVM sparse approximation procedure can be successfully applied. Tony Van Gestel, Johan A. K. Suykens, Bart Baesens, Stijn Viaene, Jan Vanthienen, Guido Dedene, Bart De Moor, Joos Vandewalle |
Mach. Learn. | 2 |
| 2003 | Classification of Ovarian Tumors Using Bayesian Least Squares Support Vector Machines
Chuan Lu, Tony Van Gestel, Johan A. K. Suykens, Sabine Van Huffel, Dirk Timmerman, Ignace Vergote |
AIME | 3 |
| 2003 | Bankruptcy prediction with least squares support vector machine classifiersabstractClassification algorithms like linear discriminant analysis and logistic regression are popular linear techniques for modelling and predicting corporate distress. These techniques aim at finding an optimal linear combination of explanatory input variables, such as, e.g., solvency and liquidity ratios, in order to analyse, model and predict corporate default risk. Recently, performant kernel based nonlinear classification techniques, like support vector machines, least squares support vector machines and kernel fisher discriminant analysis, have been developed. Basically, these methods map the inputs first in a nonlinear way to a high dimensional kernel-induced feature space, in which a linear classifier is constructed in the second step. Practical expressions are obtained in the so-called dual space by application of Mercer's theorem. In this paper, we explain the relations between linear and nonlinear kernel based classification and illustrate their performance on predicting bankruptcy of mid-cap firms in Belgium and the Netherlands. Tony Van Gestel, Bart Baesens, Johan A. K. Suykens, Marcelo Espinoza, Dirk-Emma Baestaens, Jan Vanthienen, Bart De Moor |
CIFEr | 3 |
| 2003 | Kernel PLS variants for regression
Luc Hoegaerts, Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
ESANN | 2 |
| 2003 | Preoperative prediction of malignancy of ovarian tumors using least squares support vector machines
Chuan Lu, Tony Van Gestel, Johan A. K. Suykens, Sabine Van Huffel, Ignace Vergote, Dirk Timmerman |
Artif. Intell. Medicine | 3 |
| 2003 | A support vector machine formulation to PCA analysis and its kernel versionabstractIn this paper, we present a simple and straightforward primal-dual support vector machine formulation to the problem of principal component analysis (PCA) in dual variables. By considering a mapping to a high-dimensional feature space and application of the kernel trick (Mercer theorem), kernel PCA is obtained as introduced by Scholkopf et al. (2002). While least squares support vector machine classifiers have a natural link with the kernel Fisher discriminant analysis (minimizing the within class scatter around targets +1 and -1), for PCA analysis one can take the interpretation of a one-class modeling problem with zero target value around which one maximizes the variance. The score variables are interpreted as error variables within the problem formulation. In this way primal-dual constrained optimization problem interpretations to the linear and kernel PCA analysis are obtained in a similar style as for least square-support vector machine classifiers. Johan A. K. Suykens, Tony Van Gestel, Joos Vandewalle, Bart De Moor |
IEEE Trans. Neural Networks | 1 |
| 2002 | Prediction of mental development of preterm newborns at birth time using LS-SVM
Lieveke Ameye, Chuan Lu, Lukas, Jos De Brabanter, Johan A. K. Suykens, Sabine Van Huffel, Hans Daniels, Gunnar Naulaers, Hugo Devlieger |
ESANN | 5 |
| 2002 | The use of LS-SVM in the classification of brain tumors based on Magnetic Resonance Spectroscopy signals
Lukas, Andy Devos, Johan A. K. Suykens, Leentje Vanhamme, Sabine Van Huffel, Anne Rosemary Tate, Carles Majós, Carles Arús |
ESANN | 3 |
| 2002 | Robust Cross-Validation Score Function for Non-linear Function Estimation
Jos De Brabanter, Kristiaan Pelckmans, Johan A. K. Suykens, Joos Vandewalle |
ICANN | 3 |
| 2002 | Compactly Supported RBF Kernels for Sparsifying the Gram Matrix in LS-SVM Regression Models
Bart Hamers, Johan A. K. Suykens, Bart De Moor |
ICANN | 2 |
| 2002 | Weighted least squares support vector machines: robustness and sparse approximation
Johan A. K. Suykens, Jos De Brabanter, Lukas, Joos Vandewalle |
Neurocomputing | 1 |
| 2002 | Bayesian Framework for Least-Squares Support Vector Machine Classifiers, Gaussian Processes, and Kernel Fisher Discriminant AnalysisabstractThe Bayesian evidence framework has been successfully applied to the design of multilayer perceptrons (MLPs) in the work of MacKay. Nevertheless, the training of MLPs suffers from drawbacks like the nonconvex optimization problem and the choice of the number of hidden units. In support vector machines (SVMs) for classification, as introduced by Vapnik, a nonlinear decision boundary is obtained by mapping the input vector first in a nonlinear way to a high-dimensional kernel-induced feature space in which a linear large margin classifier is constructed. Practical expressions are formulated in the dual space in terms of the related kernel function, and the solution follows from a (convex) quadratic programming (QP) problem. In least-squares SVMs (LS-SVMs), the SVM problem formulation is modified by introducing a least-squares cost function and equality instead of inequality constraints, and the solution follows from a linear system in the dual space. Implicitly, the least-squares formulation corresponds to a regression formulation and is also related to kernel Fisher discriminant analysis. The least-squares regression formulation has advantages for deriving analytic expressions in a Bayesian evidence framework, in contrast to the classification formulations used, for example, in gaussian processes (GPs). The LS-SVM formulation has clear primal-dual interpretations, and without the bias term, one explicitly constructs a model that yields the same expressions as have been obtained with GPs for regression. In this article, the Bayesian evidence framework is combined with the LS-SVM classifier formulation. Starting from the feature space formulation, analytic expressions are obtained in the dual space on the different levels of Bayesian inference, while posterior class probabilities are obtained by marginalizing over the model parameters. Empirical results obtained on 10 public domain data sets show that the LS-SVM classifier designed within the Bayesian evidence framework consistently yields good generalization performances. Tony Van Gestel, Johan A. K. Suykens, Gert R. G. Lanckriet, Annemie Lambrechts, Bart De Moor, Joos Vandewalle |
Neural Comput. | 2 |
| 2002 | Multiclass LS SVMs Moderated Outputs and Coding Decoding Schemes
Tony Van Gestel, Johan A. K. Suykens, Gert R. G. Lanckriet, Annemie Lambrechts, Bart De Moor, Joos Vandewalle |
Neural Process. Lett. | 2 |
| 2001 | Automatic relevance determination for Least Squares Support Vector Machines classifiers
Tony Van Gestel, Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
ESANN | 2 |
| 2001 | Kernel Canonical Correlation Analysis and Least Squares Support Vector Machines
Tony Van Gestel, Johan A. K. Suykens, Jos De Brabanter, Bart De Moor, Joos Vandewalle |
ICANN | 2 |
| 2001 | Knowledge discovery in a direct marketing case using least squares support vector machinesabstractWe study the problem of repeat-purchase modeling in a direct marketing setting using Belgian data. More specifically, we investigate the detection and qualification of the most relevant explanatory variables for predicting purchase incidence. The analysis is based on a wrapped form of input selection using a sensitivity based pruning heuristic to guide a greedy, stepwise, and backward traversal of the input space. For this purpose, we make use of a powerful and promising least squares support vector machine (LS-SVM) classifier formulation. This study extends beyond the standard recency frequency monetary (RFM) modeling semantics in two ways: (1) by including alternative operationalizations of the RFM variables, and (2) by adding several other (non-RFM) predictors. Results indicate that elimination of redundant/irrelevant inputs allows significant reduction of model complexity. The empirical findings also highlight the importance of frequency and monetary variables, while the recency variable category seems to be of somewhat lesser importance to the case at hand. Results also point to the added value of including non-RFM variables for improving customer profiling. More specifically, customer/company interaction, measured using indicators of information requests and complaints, and merchandise returns provide additional predictive power to purchase incidence modeling for database marketing. © 2001 John Wiley & Sons, Inc. Stijn Viaene, Bart Baesens, Tony Van Gestel, Johan A. K. Suykens, Dirk Van den Poel, Jan Vanthienen, Bart De Moor, Guido Dedene |
Int. J. Intell. Syst. | 4 |
| 2001 | Improved Long-Term Temperature Prediction by Chaining of Neural NetworksabstractWhen an artificial neural network (ANN) is trained to predict signals p steps ahead, the quality of the prediction typically decreases for large values of p. In this paper, we compare two methods for prediction with ANNs: the classical recursion of one-step ahead predictors and a new kind of chain structure. When applying both techniques to the prediction of the temperature at the end of a blast furnace, we conclude that the chaining approach leads to an improved prediction of the temperature and avoidance of instabilities, since the chained networks gradually take the prediction of their predecessors in the chain as an extra input. It is observed that instabilities might occur in the iterative case, which does not happen with the chaining approach. To select relevant inputs and decrease the number of weights in this approach, Automatic Relevance Determination (ARD) for multilayer perceptrons is applied. Michel Duhoux, Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
Int. J. Neural Syst. | 2 |
| 2001 | Optimal control by least squares support vector machines
Johan A. K. Suykens, Joos Vandewalle, Bart De Moor |
Neural Networks | 1 |
| 2001 | Financial time series prediction using least squares support vector machines within the evidence frameworkabstractThe Bayesian evidence framework is applied in this paper to least squares support vector machine (LS-SVM) regression in order to infer nonlinear models for predicting a financial time series and the related volatility. On the first level of inference, a statistical framework is related to the LS-SVM formulation which allows one to include the time-varying volatility of the market by an appropriate choice of several hyper-parameters. The hyper-parameters of the model are inferred on the second level of inference. The inferred hyper-parameters, related to the volatility, are used to construct a volatility model within the evidence framework. Model comparison is performed on the third level of inference in order to automatically tune the parameters of the kernel function and to select the relevant inputs. The LS-SVM formulation allows one to derive analytic expressions in the feature space and practical expressions are obtained in the dual space replacing the inner product by the related kernel function using Mercer's theorem. The one step ahead prediction performances obtained on the prediction of the weekly 90-day T-bill rate and the daily DAX30 closing prices show that significant out of sample sign predictions can be made with respect to the Pesaran-Timmerman test statistic. Tony Van Gestel, Johan A. K. Suykens, Dirk-Emma Baestaens, Annemie Lambrechts, Gert R. G. Lanckriet, Bruno Vandaele, Bart De Moor, Joos Vandewalle |
IEEE Trans. Neural Networks | 2 |
| 2000 | Sparse least squares Support Vector Machine classifiers
Johan A. K. Suykens, Lukas, Joos Vandewalle |
ESANN | 1 |
| 2000 | The K.U.Leuven competition data: a challenge for advanced neural network techniques
Johan A. K. Suykens, Joos Vandewalle |
ESANN | 1 |
| 2000 | Sparse approximation using least squares support vector machinesabstractIn least squares support vector machines (LS-SVMs) for function estimation Vapnik's /spl epsiv/-insensitive loss function has been replaced by a cost function which corresponds to a form of ridge regression. In this way nonlinear function estimation is done by solving a linear set of equations instead of solving a quadratic programming problem. The LS-SVM formulation also involves less tuning parameters. However, a drawback is that sparseness is lost in the LS-SVM case. In this paper we investigate imposing sparseness by pruning support values from the sorted support value spectrum which results from the solution to the linear system. Johan A. K. Suykens, Lukas, Joos Vandewalle |
ISCAS | 1 |
| 2000 | An empirical assessment of kernel type performance for least squares support vector machine classifiersabstractRecently, a modified version of support vector machines (SVMs), least-squares SVM (LS-SVM) classifiers, has been introduced, which is closely related to a form of ridge regression-type SVMs. In LS-SVMs, the classifier is obtained as the solution to a linear system instead of a quadratic programming problem. In this paper, UCI (University of California at Irvine) benchmark data sets are used to evaluate the performance of LS-SVM classifiers with linear, polynomial and radial basis function (RBF) kernels. The hyperparameters of the LS-SVM problem formulation are tuned using a 10-fold cross-validation procedure and a grid search mechanism. When comparing the performance of a nonlinear (RBF or polynomial) LS-SVM classifier with that of a linear LS-SVM, additional insight can be gained into the degree of nonlinearity of the classification problem at hand. Using a statistical motivation, it is concluded that RBF LS-SVM classifiers consistently yield among the best results for each data set. Bart Baesens, Stijn Viaene, Tony Van Gestel, Johan A. K. Suykens, Guido Dedene, Bart De Moor, Jan Vanthienen |
KES | 4 |
| 2000 | Knowledge Discovery Using Least Squares Support Vector Machine Classifiers: A Direct Marketing Case
Stijn Viaene, Bart Baesens, Tony Van Gestel, Johan A. K. Suykens, Dirk Van den Poel, Jan Vanthienen, Bart De Moor, Guido Dedene |
PKDD | 4 |
| 2000 | Robust local stability of multilayer recurrent neural networksabstractIn this paper we derive a condition for robust local stability of multilayer recurrent neural networks with two hidden layers. The stability condition follows from linking theory about linearization, robustness analysis of linear systems under nonlinear perturbation and matrix inequalities. A characterization of the basin of attraction of the origin is given in terms of the level set of a quadratic Lyapunov function. In a similar way like for NL theory, local stability is imposed around the origin and the apparent basin of attraction is made large by applying the criterion, while the proven basin of attraction is relatively small due to conservatism of the criterion. Modifying dynamic backpropagation by the new stability condition is discussed and illustrated by simulation examples. Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 1999 | Multiclass least squares support vector machinesabstractWe present an extension of least squares support vector machines (LS-SVMs) to the multiclass case. While standard SVM solutions involve solving quadratic or linear programming problems, the least squares version of SVMs corresponds to solving a set of linear equations, due to equality instead of inequality constraints in the problem formulation. In LS-SVMs the Mercer condition is still applicable. Hence several type of kernels such as polynomial, RBFs and MLPs can be used. The multiclass case that we discuss here is related to classical neural net approaches for classification where multi-classes are encoded by considering multiple outputs for the network. Efficient methods for solving large scale LS-SVMs are available. Johan A. K. Suykens, Joos Vandewalle |
IJCNN | 1 |
| 1999 | Continuous time NLq theory: absolute stability criteriaabstractWe present absolute stability (global asymptotic) criteria for continuous time multilayer recurrent neural networks with two hidden layers. Such forms arise when considering recurrent neural models and neural controllers for a given plant, both parametrized by multilayer perceptrons with one-hidden layer. The one-hidden layer case corresponds to systems in Lur'e form. These results are related to the NLq theory which is a stability theory for q-layered discrete time multilayer recurrent neural networks with conditions for global asymptotic stability and input-output stability with finite L/sub 2/-gain. The criteria can be used to constrain dynamic backpropagation in order to impose closed-loop stability for neural control schemes. Johan A. K. Suykens, Joos Vandewalle |
IJCNN | 1 |
| 1999 | Least Squares Support Vector Machine Classifiers
Johan A. K. Suykens, Joos Vandewalle |
Neural Process. Lett. | 1 |
| 1999 | Training multilayer perceptron classifiers based on a modified support vector methodabstractIn this paper we describe a training method for one hidden layer multilayer perceptron classifier which is based on the idea of support vector machines (SVM's). An upper bound on the Vapnik-Chervonenkis (VC) dimension is iteratively minimized over the interconnection matrix of the hidden layer and its bias vector. The output weights are determined according to the support vector method, but without making use of the classifier form which is related to Mercer's condition. The method is illustrated on a two-spiral classification problem. Johan A. K. Suykens, Joos Vandewalle |
IEEE Trans. Neural Networks | 1 |
| 1998 | Improved generalization ability of neurocontrollers by imposing NLq stability constraints
Johan A. K. Suykens, Joos Vandewalle |
ESANN | 1 |
| 1998 | On-Line Learning Fokker-Planck Machine
Johan A. K. Suykens, Herman Verrelst, Joos Vandewalle |
Neural Process. Lett. | 1 |
| 1997 | NLq Theory: A Neural Control Framework with Global Asymptotic Stability Criteria
Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
Neural Networks | 1 |
| 1996 | Modelling the Belgian Gas Consumption Using Neural Networks
Johan A. K. Suykens, Philippe Lemmerling, Wouter Favoreel, Bart De Moor, M. Crepel, P. Briol |
Neural Process. Lett. | 1 |
| 1995 | NLq theory: unifications in the theory of neural networks, systems and control
Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
ESANN | 1 |
| 1995 | Generalized Cellular Neural Networks Represented in he NLq FrameworkabstractThe aim of this paper is to show that discrete time Generalized Cellular Neural Networks, with feedforward, feedback or cascade interconnections between CNNs can be represented as NL/sub q/s. NL/sub q/s are nonlinear systems in state space form with the typical feature of having a number of q layers with alternating linear and nonlinear operators that satisfy a sector condition. It can be shown that many systems and problems arising in neural networks, systems and control are special cases of NL/sub q/s. Sufficient conditions for global asymptotic stability and dissipativity with finite L/sub 2/-gain are available. For q=1 the criteria are closely related to known results in H/sub /spl infin// and /spl mu/ control theory. Johan A. K. Suykens, Joos Vandewalle |
ISCAS | 1 |
| 1994 | Static and dynamic stabilizing neural controllers, applicable to transition between equilibrium points
Johan A. K. Suykens, Bart De Moor, Joos Vandewalle |
Neural Networks | 1 |