Siamak Mehrkanoon

dblp:96/9472 · DBLP profile ↗
← Back
46ranked-venue papers
20as first author
18since 2021 · last 2025
0000-0002-0516-0391ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 20 first-author · 18 since 2021
YearPublicationVenuePosition
2025 SSA-UNet: Advanced Precipitation Nowcasting via Channel Shuffling
abstract
Weather forecasting is essential for facilitating diverse socio-economic activity and environmental conservation initiatives. Deep learning techniques are increasingly being explored as complementary approaches to Numerical Weather Prediction (NWP) models, offering potential benefits such as reduced complexity and enhanced adaptability in specific applications. This work presents a novel design, Small Shuffled Attention UNet (SSA-UNet), which enhances SmaAt-UNet’s architecture by including a shuffle channeling mechanism to optimize performance and diminish complexity. To assess its efficacy, this architecture and its reduced variant are examined and trained on two datasets: a Dutch precipitation dataset from 2016 to 2019, and a French cloud cover dataset containing radar images from 2017 to 2018. Three output configurations of the proposed architecture are evaluated, yielding outputs of 1, 6, and 12 precipitation maps, respectively. To better understand how this model operates and produces its predictions, a gradient-based approach called Grad-CAM is used to analyze the outputs generated. The analysis of heatmaps generated by Grad-CAM facilitated the identification of regions within the input maps that the model considers most informative for generating its predictions. The implementation of SSA-UNet can be found on our Github1.
Marco Turzi, Siamak Mehrkanoon
IJCNN2
2025 Graph Dual-stream Convolutional Attention Fusion for precipitation nowcasting
abstract
Accurate precipitation nowcasting is crucial for applications such as flood prediction, disaster management, agriculture optimization, and transportation management. While many studies have approached this task using sequence-to-sequence models, most focus on single regions, ignoring correlations between disjoint areas. We reformulate precipitation nowcasting as a spatiotemporal graph sequence problem. Specifically, we propose Graph Dual-stream Convolutional Attention Fusion, a novel extension of the graph attention network. Our model’s dual-stream design employs distinct attention mechanisms for spatial and temporal interactions, capturing their unique dynamics. A gated fusion module integrates both streams, leveraging spatial and temporal information for improved predictive accuracy. Additionally, our framework enhances graph attention by directly processing three-dimensional tensors within graph nodes, removing the need for reshaping. This capability enables handling complex, high-dimensional data and exploiting higher-order correlations between data dimensions. Depthwise-separable convolutions are also incorporated to refine local feature extraction and efficiently manage high-dimensional inputs. We evaluate our model using seven years of precipitation data from Copernicus Climate Change Services, covering Europe and neighboring regions. Experimental results demonstrate superior performance of our approach compared to other models. Moreover, visualizations of seasonal spatial and temporal attention scores provide insights into the most significant connections between regions and time steps.
Lorand Vatamany, Siamak Mehrkanoon
Eng. Appl. Artif. Intell.2
2025 TransCORALNet: A two-stream transformer CORAL networks for supply chain credit assessment cold start
abstract
Supply chain credit assessment is critical for financial decision-making due to limited historical data for new borrowers and the domain shift between segment industries. Existing models often struggle with challenges such as domain shift, cold start, imbalanced classes, and lack of interpretability . This paper proposes an interpretable two-stream transformer CORAL network (TransCORALNet) for supply chain credit assessment, designed to address these challenges. The two-stream domain adaptation architecture with correlation alignment (CORAL) loss serves as the core model and is equipped with a transformer, which provides insights into the learned features and allows efficient parallelization during training. Thanks to the domain adaptation capability of the proposed model, the domain shift between the source and target domains is minimized. Furthermore, we employ Local Interpretable Model-agnostic Explanations (LIME) to provide additional insights into the model predictions and identify the key features contributing to supply chain credit assessment decisions. Experimental results on a real-world dataset demonstrate the superiority of TransCORALNet over several state-of-the-art baselines in terms of accuracy. The code is available on GitHub. 1
Jie Shi 0006, Arno Siebes, Siamak Mehrkanoon
Expert Syst. Appl.3
2025 BAST-Mamba: Binaural Audio Spectrogram Mamba Transformer for binaural sound localization
abstract
Accurate sound localization in reverberant environments is essential for human auditory perception. Recently, Convolutional Neural Networks (CNNs) have been used to model the binaural human auditory pathway. However, CNNs face limitations in capturing global acoustic features. To address this issue, we propose a novel end-to-end Binaural Audio Spectrogram Mamba Transformer (BAST-Mamba) model to predict sound azimuth in both anechoic and reverberant conditions. We explore two implementation modes: BAST-Mamba-SP and BAST-Mamba-NSP, which correspond to shared and non-shared parameter configurations, respectively. Our best model BAST-Mamba-SP, equipped with subtraction-based interaural integration and a hybrid loss function, achieves a state-of-the-art angular distance (AD) error of 0.89°and mean squared error of 0.0004, significantly outperforming baseline models. The model demonstrates generalization across acoustic environments, robust hemifield symmetry and high accurate real-time localization performance ( < 4°AD at 300 ms). Moderate noise augmentation at 30 dB SNR yields the strongest noise resilience. Explainability analyses highlight consistent frequency focus in the 2–3 kHz and 5.5–6.5 kHz bands, aligning with known neurophysiological cues. These results validate the potential of neurobiologically inspired Transformer for robust, high-precision sound localization and offer new insights into human sound localization.
Sheng Kuang, Jie Shi 0006, Kiki van der Heijden, Siamak Mehrkanoon
Neurocomputing4
2024 Dual Stream Graph Transformer Fusion Networks for Enhanced Brain Decoding
abstract
This paper presents the novel Dual Stream Graph-Transformer Fusion (DS-GTF) architecture designed specifically for classifying task-based Magnetoencephalography (MEG) data.In the spatial stream, inputs are initially represented as graphs, which are then passed through graph attention networks (GAT) to extract spatial patterns.Two methods, TopK and Thresholded Adjacency are introduced for initializing the adjacency matrix used in the GAT.In the temporal stream, the Transformer Encoder receives concatenated windowed input MEG data and learns new temporal representations.The learned temporal and spatial representations from both streams are fused before reaching the output layer.Experimental results demonstrate an enhancement in classification performance and a reduction in standard deviation across multiple test subjects compared to other examined models.
Lucas Goené, Siamak Mehrkanoon
ESANN2
2024 A novel dual-stream time-frequency contrastive pretext tasks framework for sleep stage classification
abstract
Self-supervised learning addresses the challenge encountered by many supervised methods, i.e. the requirement of large amounts of annotated data. This challenge is particularly pronounced in fields such as the electroencephalography (EEG) research domain. Self-supervised learning operates instead by utilizing pseudo-labels, which are generated by pretext tasks, to obtain a rich and meaningful data representation. In this study, we aim at introducing a dual-stream pretext task architecture that operates both in the time and frequency domains. In particular, we have examined the incorporation of the novel Frequency Similarity (FS) pretext task into two existing pretext tasks, Relative Positioning (RP) and Temporal Shuffling (TS). We assess the accuracy of these models using the Physionet Challenge 2018 (PC18) dataset in the context of the downstream task sleep stage classification. The inclusion of FS resulted in a notable improvement in downstream task accuracy, with a 1.28 percent improvement on RP and a 2.02 percent improvement on TS. Furthermore, when visualizing the learned embeddings using Uniform Manifold Approximation and Projection (UMAP), distinct clusters emerge, indicating that the learned representations carry meaningful information.
Sergio Kazatzidis, Siamak Mehrkanoon
IJCNN2
2024 GA-SmaAt-GNet: Generative adversarial small attention GNet for extreme precipitation nowcasting
Eloy Reulen, Jie Shi 0006, Siamak Mehrkanoon
Knowl. Based Syst.3
2023 SAR-UNet: Small Attention Residual UNet for Explainable Nowcasting Tasks
abstract
The accuracy and explainability of data-driven now-casting models are of great importance in many socio-economic sectors reliant on weather-dependent decision making. This paper proposes a novel architecture called Small Attention Residual UNet (SAR-UNet) for precipitation and cloud cover nowcasting. Here, SmaAt-UNet is used as a core model and is further equipped with residual connections, parallel to the depthwise separable convolutions. The proposed SAR-UNet model is evaluated on two datasets, i.e., Dutch precipitation maps ranging from 2016 to 2019 and French cloud cover binary images from 2017 to 2018. The obtained results show that SAR-UNet outperforms other examined models in precipitation nowcasting from 30 to 180 minutes in the future as well as cloud cover nowcasting in the next 90 minutes. Furthermore, we provide additional insights on the nowcasts made by our proposed model using Grad-CAM, a visual explanation technique, which is employed on different levels of the encoder and decoder paths of the SAR-UNet model and produces heatmaps highlighting the critical regions in the input image as well as intermediate representations to the precipitation. The heatmaps generated by Grad-CAM reveal the interactions between the residual connections and the depthwise separable convolutions inside of the multiple depthwise separable blocks placed throughout the network architecture.
Mathieu Renault, Siamak Mehrkanoon
IJCNN2
2023 MSCDA: Multi-level semantic-guided contrast improves unsupervised domain adaptation for breast MRI segmentation in small datasets
abstract
Deep learning (DL) applied to breast tissue segmentation in magnetic resonance imaging (MRI) has received increased attention in the last decade, however, the domain shift which arises from different vendors, acquisition protocols, and biological heterogeneity, remains an important but challenging obstacle on the path towards clinical implementation. In this paper, we propose a novel Multi-level Semantic-guided Contrastive Domain Adaptation (MSCDA) framework to address this issue in an unsupervised manner. Our approach incorporates self-training with contrastive learning to align feature representations between domains. In particular, we extend the contrastive loss by incorporating pixel-to-pixel, pixel-to-centroid, and centroid-to-centroid contrasts to better exploit the underlying semantic information of the image at different levels. To resolve the data imbalance problem, we utilize a category-wise cross-domain sampling strategy to sample anchors from target images and build a hybrid memory bank to store samples from source images. We have validated MSCDA with a challenging task of cross-domain breast MRI segmentation between datasets of healthy volunteers and invasive breast cancer patients. Extensive experiments show that MSCDA effectively improves the model's feature alignment capabilities between domains, outperforming state-of-the-art methods. Furthermore, the framework is shown to be label-efficient, achieving good performance with a smaller source dataset. The code is publicly available at https://github.com/ShengKuangCN/MSCDA.
Sheng Kuang, Henry C. Woodruff, Renee Granzier, Thiemo J. A. van Nijnatten, Marc Lobbes, Marjolein L. Smidt, Philippe Lambin, Siamak Mehrkanoon
Neural Networks8
2022 AA-TransUNet: Attention Augmented TransUNet For Nowcasting Tasks
abstract
Data driven modeling based approaches have recently gained a lot of attention in many challenging meteoro-logical applications including weather element forecasting. This paper introduces a novel data-driven predictive model based on TransUNet for precipitation nowcasting task. The TransUNet model which combines the Transformer and U-Net models has been previously successfully applied in medical segmentation tasks. Here, TransUnet is used as a core model and is further equipped with Convolutional Block Attention Modules (CBAM) and Depthwise-separable Convolution (DSC). The proposed Attention Augmented TransUNet (AA- TransUNet) model is evaluated on two distinct datasets: the Dutch precipitation map dataset and the French cloud cover dataset. The obtained results show that the proposed model outperforms other examined models on both tested datasets. Furthermore, the uncertainty analysis of the proposed AA-TransUNet is provided to give additional insights on its predictions.
Siamak Mehrkanoon
IJCNN2
2022 GCN-FFNN: A two-stream deep model for learning solution to partial differential equations
abstract
This paper introduces a novel two-stream deep model based on graph convolutional network (GCN) architecture and feed-forward neural networks (FFNN) for learning the solution of nonlinear partial differential equations (PDEs). The model aims at incorporating both graph and grid input representations using two streams corresponding to GCN and FFNN models, respectively. Each stream layer receives and processes its input representation. As opposed to FFNN which receives a grid-like structure, the GCN stream layer operates on graph input data where the neighborhood information is incorporated through the adjacency matrix of the graph. In this way, the proposed GCN-FFNN model learns from two types of input representations, i.e. grid and graph data, obtained via the discretization of the PDE domain. The GCN-FFNN model is trained in two phases. In the first phase, the model parameters of each stream are trained separately. Both streams employ the same error function to adjust their parameters by enforcing the models to satisfy the given PDE as well as its initial and boundary conditions on grid or graph collocation (training) data. In the second phase, the learned parameters of two-stream layers are frozen and their learned representation solutions are fed to fully connected layers whose parameters are learned using the previously used error function. The learned GCN-FFNN model is tested on test data located both inside and outside the PDE domain. The obtained numerical results demonstrate the applicability and efficiency of the proposed GCN-FFNN model over individual GCN and FFNN models on 1D-Burgers, 1D-Schrödinger, 2D-Burgers, and 2D-Schrödinger equations.
Onur Bilgin 0001, Thomas Vergutz, Siamak Mehrkanoon
Neurocomputing3
2022 Goal-driven, neurobiological-inspired convolutional neural network models of human spatial hearing
abstract
The human brain effortlessly solves the complex computational task of sound localization using a mixture of spatial cues. How the brain performs this task in naturalistic listening environments (e.g. with reverberation) is not well understood. In the present paper, we build on the success of deep neural networks at solving complex and high-dimensional problems [1] to develop goal-driven, neurobiological-inspired convolutional neural network (CNN) models of human spatial hearing. After training, we visualize and quantify feature representations in intermediate layers to gain insights into the representational mechanisms underlying sound location encoding in CNNs. Our results show that neurobiological-inspired CNN models trained on real-life sounds spatialized with human binaural hearing characteristics can accurately predict sound location in the horizontal plane. CNN localization acuity across the azimuth resembles human sound localization acuity, but CNN models outperform human sound localization in the back. Training models with different objective functions - that is, minimizing either Euclidean or angular distance - modulates localization acuity in particular ways. Moreover, different implementations of binaural integration result in unique patterns of localization errors that resemble behavioral observations in humans. Finally, feature representations reveal a gradient of spatial selectivity across network layers, starting with broad spatial representations in early layers and progressing to sparse, highly selective spatial representations in deeper layers. In sum, our results show that neurobiological-inspired CNNs are a valid approach to modeling human spatial hearing. This work paves the way for future studies combining neural network models with empirical measurements of neural activity to unravel the complex computational mechanisms underlying neural sound location encoding in the human auditory pathway.
Kiki van der Heijden, Siamak Mehrkanoon
Neurocomputing2
2022 Deep coastal sea elements forecasting using UNet-based models
abstract
Due to the recent development of deep learning techniques applied to satellite imagery, weather forecasting that uses remote sensing data has also been the subject of major progress. The present paper investigates multiple hours ahead coastal sea elements forecasting in the Netherlands using UNet based architectures. The hourly satellite image data from the Copernicus observation program spanned over a period of two years has been used to train the models and make the forecasting, including seasonal forecasting. Here, we propose 3D dimension Reducer UNet (3DDR-UNet), a variation of the UNet architecture, and further extend this novel model using residual connections, parallel convolutions and asymmetric convolutions which result in introducing three additional architectures, i.e. Res-3DDR-UNet, InceptionRes-3DDR-UNet and AsymmInceptionRes-3DDR-UNet respectively. In particular, we show that the architecture equipped with parallel and asymmetric convolutions as well as skip connections outperforms the other three discussed models.
Jesús García Fernández, Ismail Alaoui Abdellaoui, Siamak Mehrkanoon
Knowl. Based Syst.3
2021 Enhancing brain decoding using attention augmented deep neural networks
abstract
Neuroimaging techniques have shown to be valuable when studying brain activity.This paper uses Magnetoencephalography (MEG) data, provided by the Human Connectome Project (HCP), and different deep learning models to perform brain decoding.Specifically, we investigate to which extent one can infer the task performed by a subject based on its MEG data.In order to capture the most relevant features of the signals, self and global attention are incorporated into our models.The obtained results show that the inclusion of attention improves the performance and generalization of the models across subjects.Attention mechanisms allow the models to capture long-range dependencies, and highlight/suppress relevant/irrelevant parts of the input.The models used in this paper are equipped with two types of attention: self and global.
Ismail Alaoui Abdellaoui, Jesús García Fernández, Caner Sahinli, Siamak Mehrkanoon
ESANN4
2021 Deep Graph Convolutional Networks for Wind Speed Prediction
abstract
In this paper, we introduce a new model for wind speed prediction based on spatio-temporal graph convolutional networks.Here, weather stations are treated as nodes of a graph with a learnable adjacency matrix, which determines the strength of relations between the stations based on the historical weather data.The self-loop connection is added to the learnt adjacency matrix and its strength is controlled by additional learnable parameter.Experiments performed on real datasets collected from weather stations located in Denmark and the Netherlands show that our proposed model outperforms previously developed baseline models on the referenced datasets.
Tomasz Stanczyk, Siamak Mehrkanoon
ESANN2
2021 Exploring automatic liver tumor segmentation using deep learning
abstract
The segmentation of liver tumors is crucial for diagnosis, treatment planning and treatment evaluation. Due to the setbacks that the manual segmentation brings, automatic segmentation has recently gained a lot of attention. In this work, we explore various deep learning based approaches to address automatic liver tumor segmentation. We use the data from the Liver Tumor Segmentation challenge (LiTS). In particular, the considered models here are UNet-based architectures. In addition, we investigate the influence of incorporating extra elements to the pipeline such as attention mechanisms, model ensemble, test-time inference as well as an additional model to reject false positives, over the final performance. The obtained results show that the 3D-UNet architecture, together with ensemble learning methods, performs more accurate predictions than the other examined approaches.
Jesús García Fernández, Valerio Fortunati, Siamak Mehrkanoon
IJCNN3
2021 Broad-UNet: Multi-scale feature learning for nowcasting tasks
abstract
Weather nowcasting consists of predicting meteorological components in the short term at high spatial resolutions. Due to its influence in many human activities, accurate nowcasting has recently gained plenty of attention. In this paper, we treat the nowcasting problem as an image-to-image translation problem using satellite imagery. We introduce Broad-UNet, a novel architecture based on the core UNet model, to efficiently address this problem. In particular, the proposed Broad-UNet is equipped with asymmetric parallel convolutions as well as Atrous Spatial Pyramid Pooling (ASPP) module. In this way, the Broad-UNet model learns more complex patterns by combining multi-scale features while using fewer parameters than the core UNet model. The proposed model is applied on two different nowcasting tasks, i.e. precipitation maps and cloud cover nowcasting. The obtained numerical results show that the introduced Broad-UNet model performs more accurate predictions compared to the other examined architectures.
Jesús García Fernández, Siamak Mehrkanoon
Neural Networks2
2021 SmaAt-UNet: Precipitation nowcasting using a small attention-UNet architecture
abstract
Weather forecasting is dominated by numerical weather prediction that tries to model accurately the physical properties of the atmosphere. A downside of numerical weather prediction is that it is lacking the ability for short-term forecasts using the latest available information. By using a data-driven neural network approach we show that it is possible to produce an accurate precipitation nowcast. To this end, we propose SmaAt-UNet, an efficient convolutional neural networks-based on the well known UNet architecture equipped with attention modules and depthwise-separable convolutions. We evaluate our approaches on a real-life datasets using precipitation maps from the region of the Netherlands and binary images of cloud coverage of France. The experimental results show that in terms of prediction performance, the proposed model is comparable to other examined models while only using a quarter of the trainable parameters.
Kevin Trebing, Tomasz Stanczyk, Siamak Mehrkanoon
Pattern Recognit. Lett.3
2020 A Real-time PCB Defect Detector Based on Supervised and Semi-supervised Learning
Sanli Tang, Siamak Mehrkanoon, Xiaolin Huang, Jie Yang 0002
ESANN3
2020 Modelling human sound localization with deep neural networks
Kiki van der Heijden, Siamak Mehrkanoon
ESANN2
2020 Learning from partially labeled data
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens
ESANN1
2019 Deep shared representation learning for weather elements forecasting
Siamak Mehrkanoon
Knowl. Based Syst.1
2019 Deep neural-kernel blocks
Siamak Mehrkanoon
Neural Networks1
2019 Cross-domain neural-kernel networks
Siamak Mehrkanoon
Pattern Recognit. Lett.1
2018 Shallow and Deep Models for Domain Adaptation problems
Siamak Mehrkanoon, Matthew B. Blaschko, Johan A. K. Suykens
ESANN1
2018 Deep hybrid neural-kernel networks using random Fourier features
Siamak Mehrkanoon, Johan A. K. Suykens
Neurocomputing1
2018 Indefinite kernel spectral learning
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens
Pattern Recognit.1
2018 Regularized Semipaired Kernel CCA for Domain Adaptation
abstract
Domain adaptation learning is one of the fundamental research topics in pattern recognition and machine learning. This paper introduces a regularized semipaired kernel canonical correlation analysis formulation for learning a latent space for the domain adaptation problem. The optimization problem is formulated in the primal-dual least squares support vector machine setting where side information can be readily incorporated through regularization terms. The proposed model learns a joint representation of the data set across different domains by solving a generalized eigenvalue problem or linear system of equations in the dual. The approach is naturally equipped with out-of-sample extension property, which plays an important role for model selection. Furthermore, the Nyström approximation technique is used to make the computational issues due to the large size of the matrices involved in the eigendecomposition feasible. The learned latent space of the source domain is fed to a multiclass semisupervised kernel spectral clustering model that can learn from both labeled and unlabeled data points of the source domain in order to classify the data instances of the target domain. Experimental results are given to illustrate the effectiveness of the proposed approaches on synthetic and real-life data sets.
Siamak Mehrkanoon, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2017 Scalable Hybrid Deep Neural Kernel Networks
Siamak Mehrkanoon, Andreas Zell, Johan A. K. Suykens
ESANN1
2016 Multi-label semi-supervised learning using regularized kernel spectral clustering
abstract
Often in real-world applications such as web page categorization, automatic image annotations and protein function prediction, each instance is associated with multiple labels (categories) simultaneously. In addition, due to the labeling cost one usually deals with a large amount of unlabeled data while the fraction of labeled data points will typically be small. In this paper, we propose a multi-label semi-supervised kernel spectral clustering learning algorithm that learns from both labeled and unlabeled instances. The kernel spectral clustering algorithm (KSC) serves as a core model and the information of labeled data points is integrated into the model via regularization terms. The propagation of the multiple labels to unlabeled data points is achieved by incorporating the mutual correlation between (similarity across) labels as well as encouraging the model output to be as close as possible to the given ground-truth of the labeled data points. Thanks to the Nyström approximation method, an explicit feature map is constructed and the optimization problem is solved in the primal. Experimental results demonstrate the effectiveness of the proposed approaches on real multi-label datasets.
Siamak Mehrkanoon, Johan A. K. Suykens
IJCNN1
2016 Estimating the unknown time delay in chemical processes
Siamak Mehrkanoon, Yuri A. W. Shardt, Johan A. K. Suykens, Steven X. Ding
Eng. Appl. Artif. Intell.1
2016 Robust Support Vector Machines for Classification with Nonconvex and Smooth Losses
abstract
This letter addresses the robustness problem when learning a large margin classifier in the presence of label noise. In our study, we achieve this purpose by proposing robustified large margin support vector machines. The robustness of the proposed robust support vector classifiers (RSVC), which is interpreted from a weighted viewpoint in this work, is due to the use of nonconvex classification losses. Besides the robustness, we also show that the proposed RSCV is simultaneously smooth, which again benefits from using smooth classification losses. The idea of proposing RSVC comes from M-estimation in statistics since the proposed robust and smooth classification losses can be taken as one-sided cost functions in robust statistics. Its Fisher consistency property and generalization ability are also investigated. Besides the robustness and smoothness, another nice property of RSVC lies in the fact that its solution can be obtained by solving weighted squared hinge loss-based support vector machine problems iteratively. We further show that in each iteration, it is a quadratic programming problem in its dual space and can be solved by using state-of-the-art methods. We thus propose an iteratively reweighted type algorithm and provide a constructive proof of its convergence to a stationary point. Effectiveness of the proposed classifiers is verified on both artificial and real data sets.
Yunlong Feng, Xiaolin Huang, Siamak Mehrkanoon, Johan A. K. Suykens
Neural Comput.4
2015 Black-box modeling for temperature prediction in weather forecasting
abstract
Accurate weather forecasting is one of most challenging tasks that deals with a large amount of observations and features. In this paper, a black-box modeling technique is proposed for temperature forecasting. Due to the high dimensionality of data, feature selection is done in two steps with k-Nearest Neighbors and Elastic net. Next, Least Squares Support Vector Machine regression is applied to generate the forecasting model. In the experimental results, the influence of each part of this procedure on the performance is investigated and compared with “Weather underground” results. For the case study, the prediction of the temperature in Brussels is considered. It is shown that black-box modeling has a good and competitive accuracy with current state-of-the-art methods for temperature prediction.
Zahra Karevan, Siamak Mehrkanoon, Johan A. K. Suykens
IJCNN2
2015 Hierarchical semi-supervised clustering using KSC based model
abstract
This paper introduces a methodology to incorporate the label information in discovering the underlying clusters in a hierarchical setting using multi-class semi-supervised clustering algorithm. The method aims at revealing the relationship between clusters given few labels associated to some of the clusters. The problem is formulated as a regularized kernel spectral clustering algorithm in the primal-dual setting. The available labels are incorporated in different levels of hierarchy from top to bottom. As we advance towards the lowers levels in the tree all the previously added labels are used in the generation of the new levels of hierarchy. The model is trained on a subset of the data and then applied to the rest of the data in a learning framework. Thanks to the previously learned model, the out-of-sample extension property of the model allows then to predict the memberships of a new point. A combination of an internal clustering quality index and classification accuracy is used for model selection. Experiments are conducted on synthetic data and real image segmentation problems to show the applicability of the proposed approach.
Siamak Mehrkanoon, Oscar Mauricio Agudelo, Raghvendra Mall, Johan A. K. Suykens
IJCNN1
2015 Learning solutions to partial differential equations using LS-SVM
Siamak Mehrkanoon, Johan A. K. Suykens
Neurocomputing1
2015 Incremental multi-class semi-supervised clustering regularized by Kalman filtering
Siamak Mehrkanoon, Oscar Mauricio Agudelo, Johan A. K. Suykens
Neural Networks1
2015 Identifying intervals for hierarchical clustering using the Gershgorin circle theorem
Raghvendra Mall, Siamak Mehrkanoon, Johan A. K. Suykens
Pattern Recognit. Lett.2
2015 Multiclass Semisupervised Learning Based Upon Kernel Spectral Clustering
abstract
This paper proposes a multiclass semisupervised learning algorithm by using kernel spectral clustering (KSC) as a core model. A regularized KSC is formulated to estimate the class memberships of data points in a semisupervised setting using the one-versus-all strategy while both labeled and unlabeled data points are present in the learning process. The propagation of the labels to a large amount of unlabeled data points is achieved by adding the regularization terms to the cost function of the KSC formulation. In other words, imposing the regularization term enforces certain desired memberships. The model is then obtained by solving a linear system in the dual. Furthermore, the optimal embedding dimension is designed for semisupervised clustering. This plays a key role when one deals with a large number of clusters.
Siamak Mehrkanoon, Carlos Alzate, Raghvendra Mall, Rocco Langone, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2014 SVD truncation schemes for fixed-size kernel models
abstract
In this paper, two schemes for reducing the effective number of parameters are presented. To do this, different versions of Fixed-Size Kernel models based on Fixed-Size Least Squares Support Vector Machines (FS-LSSVM) are employed. The schemes include Fixed-Size Ordinary Least Squares (FS-OLS) and Fixed-Size Ridge Regression (FS-RR) with their respective truncations through Singular Value Decomposition (SVD). When these schemes are applied to the Silverbox and Wiener-Hammerstein data sets in system identification, it was found that a great deal of the complexity of the model could be reduced in a trade-off with the generalization performance.
Ricardo Castro-Garcia, Siamak Mehrkanoon, Anna Marconato, Johan Schoukens, Johan A. K. Suykens
IJCNN2
2014 Optimal reduced sets for sparse kernel spectral clustering
abstract
Kernel spectral clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. It results in a clustering model using the dual solution of the problem. It has a powerful out-of-sample extension property leading to good clustering generalization w.r.t. the unseen data points. The out-of-sample extension property allows to build a sparse model on a small training set and introduces the first level of sparsity. The clustering dual model is expressed in terms of non-sparse kernel expansions where every point in the training set contributes. The goal is to find reduced set of training points which can best approximate the original solution. In this paper a second level of sparsity is introduced in order to reduce the time complexity of the computationally expensive out-of-sample extension. In this paper we investigate various penalty based reduced set techniques including the Group Lasso, L0, L1+ L0penalization and compare the amount of sparsity gained w.r.t. a previous L1penalization technique. We observe that the optimal results in terms of sparsity corresponds to the Group Lasso penalization technique in majority of the cases. We showcase the effectiveness of the proposed approaches on several real world datasets and an image segmentation dataset.
Raghvendra Mall, Siamak Mehrkanoon, Rocco Langone, Johan A. K. Suykens
IJCNN2
2014 Large scale semi-supervised learning using KSC based model
abstract
Often in practice one deals with a large amount of unlabeled data, while the fraction of labeled data points will typically be small. Therefore one prefers to apply a semi-supervised algorithm, which uses both labeled and unlabeled data points in the learning process, to have a better performance. Considering the large amount of unlabeled data, making a semi-supervised algorithm scalable is an important task. In this paper we adopt a recently proposed multi-class semi-supervised KSC based algorithm (MSS-KSC) and make it scalable by means of two different approaches. The first one is based on the Nyström approximation method which provides a finite dimensional feature map that can then be used to solve the optimization problem in the primal. The second approach is based on the reduced kernel technique that solves the problem in the dual by reducing the dimensionality of the kernel matrix to a rectangular kernel. Experimental results demonstrate the scalability and efficiency of the proposed approaches on real datasets.
Siamak Mehrkanoon, Johan A. K. Suykens
IJCNN1
2014 Non-parallel support vector classifiers with different loss functions
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens
Neurocomputing1
2013 Non-parallel semi-supervised classification based on kernel spectral clustering
abstract
In this paper, a non-parallel semi-supervised algorithm based on kernel spectral clustering is formulated. The prior knowledge about the labels is incorporated into the kernel spectral clustering formulation via adding regularization terms. In contrast with the existing multi-plane classifiers such as Multisurface Proximal Support Vector Machine (GEPSVM) and Twin Support Vector Machines (TWSVM) and its least squares version (LSTSVM) we will not use a kernel-generated surface. Instead we apply the kernel trick in the dual. Therefore as opposed to conventional non-parallel classifiers one does not need to formulate two different primal problems for the linear and nonlinear case separately. The proposed method will generate two non-parallel hyperplanes which then are used for out-of-sample extension. Experimental results demonstrate the efficiency of the proposed method over existing methods.
Siamak Mehrkanoon, Johan A. K. Suykens
IJCNN1
2013 Support vector machines with piecewise linear feature mapping
Xiaolin Huang, Siamak Mehrkanoon, Johan A. K. Suykens
Neurocomputing2
2012 Approximate Solutions to Ordinary Differential Equations Using Least Squares Support Vector Machines
abstract
In this paper, a new approach based on least squares support vector machines (LS-SVMs) is proposed for solving linear and nonlinear ordinary differential equations (ODEs). The approximate solution is presented in closed form by means of LS-SVMs, whose parameters are adjusted to minimize an appropriate error function. For the linear and nonlinear cases, these parameters are obtained by solving a system of linear and nonlinear equations, respectively. The method is well suited to solving mildly stiff, nonstiff, and singular ODEs with initial and boundary conditions. Numerical results demonstrate the efficiency of the proposed method over existing methods.
Siamak Mehrkanoon, Tillmann Falck, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2011 Symbolic computing of LS-SVM based models
Siamak Mehrkanoon, Carlos Alzate, Johan A. K. Suykens
ESANN1