Jan Platos

dblp:64/2617 · DBLP profile ↗
← Back
66ranked-venue papers
8as first author
13since 2021 · last 2024
0000-0002-8481-0136ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 first-authorSystems, architecture and hardware · 5 · 1 first-author · 2 since 2021Security and privacy · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Computer networks · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Abstractive text summarization model combining a hierarchical attention mechanism and multiobjective reinforcement learning
Jan Platos
Expert Syst. Appl.2
2024 A hierarchical overlapping community detection method based on closed trail distance and maximal cliques
abstract
An important feature of real networks is their hierarchy and the existence of overlapping communities. Hierarchical agglomerative clustering is one way to determine the hierarchy of a network. To ensure the existence of overlapping communities, it is appropriate to choose the base elements for clustering – edges, cliques, etc. These base elements can then have common vertices and naturally provide the possibility of overlap. The proposed community detection method uses hierarchical agglomerative clustering on the 2-edge-connected component of the graph. Communities are constructed from maximal cliques as base elements. Novel dissimilarities for hierarchical agglomerative clustering were introduced for the merging of cliques. The dissimilarities use the size of the overlapped cliques and closed trail distance to express dissimilarity between communities in networks. The single linkage approach contains and extends the results of k-CPM. The proposed algorithm utilizing deterministic dissimilarity achieves comparable or superior outcomes compared to standard algorithms used for hierarchical or overlapping community detection.
Pavla Drázdilová, Petr Prokop, Jan Platos, Václav Snásel
Inf. Sci.3
2024 A natural gas consumption forecasting system for continual learning scenarios based on Hoeffding trees with change point detection mechanism
Radek Svoboda, Sebastián Basterrech, Jedrzej Kozal, Jan Platos, Michal Wozniak 0001
Knowl. Based Syst.4
2024 A quantum inspired differential evolution algorithm for automatic clustering of real life datasets
Alokananda Dey, Siddhartha Bhattacharyya 0001, Sandip Dey, Jan Platos, Václav Snásel
Multim. Tools Appl.4
2023 A Continual Learning System with Self Domain Shift Adaptation for Fake News Detection
abstract
Detecting fake news is currently one of the critical challenges facing modern societies. The problem is particularly relevant, as disinformation is readily used for political warfare but can also cause significant harm to the health of citizens, such as by promoting false data on the harmfulness of selected therapies. One way to combat disinformation is to treat fake news detection as a machine learning task. This paper presents such an approach, which additionally addresses an important problem related to the non-stationarity characteristics of the fake news. We elaborated a stream data with the simulation of domain shift based on two popular benchmark datasets dedicated to the fake news classification problem (Kaggle Fake News and Constraint@AAAI2021–COVID19 Fake News Detection). The proposed learning system works in a Continual Learning (CL) framework and integrates a self domain shift adaptation in a machine learning scheme. The method was built following state-of-the-art techniques, that includes Word2Vec as a feature extractor and the LSTM model as a classifier. The performance of the approach has been evaluated over the generated data stream. The convenience of our approach is showed in the results, where the accuracy gain with respect to a CL approach without domain adaptation is observed to be significant.
Sebastián Basterrech, Andrzej Kasprzak, Jan Platos, Michal Wozniak 0001
DSAA3
2023 Automatic clustering of colour images using quantum inspired meta-heuristic algorithms
Alokananda Dey, Siddhartha Bhattacharyya 0001, Sandip Dey, Jan Platos, Václav Snásel
Appl. Intell.4
2023 Differential Game-Based Optimal Control of Autonomous Vehicle Convoy
abstract
Group control of connected and autonomous vehicles on automated highways is challenging for the advanced driver assistance systems (ADAS) and the automated driving systems (ADS). This paper investigates the differential game-based approach to autonomous convoy control with the aim of deployment on automated highways. Under the noncooperative differential games, the coupled vehicles make their decisions independently while their states are interdependent. The receding horizon Nash equilibrium of the linear-quadratic differential game provides the convoy a distributed state-feedback control strategy. This approach suffers a fundamental issue that neither a Nash equilibrium’s existence nor the uniqueness is guaranteed. We convert the individual dynamics-based differential game to a relative dynamics-based optimal control problem that carries all the features of the differential game. The existence of a unique Nash control under the differential game corresponds to a unique solution to the optimal control problem. The latter is shown, as well as the asymptotic stability of the closed-loop system. Simulations illustrate the effectiveness of the presented convey control scheme and how it well suits automated highway driving scenarios.
Hossein Barghi Jond, Jan Platos
IEEE Trans. Intell. Transp. Syst.2
2022 Experimental Analysis on Dissimilarity Metrics and Sudden Concept Drift Detection
Sebastián Basterrech, Jan Platos, Gerardo Rubino, Michal Wozniak 0001
ISDA (3)2
2022 Automation of cleaning and ensembles for outliers detection in questionnaire data
abstract
This article is focused on the automatic detection of the corrupted or inappropriate responses in questionnaire data using unsupervised outliers detection. The questionnaire surveys are often used in psychology research to collect self-report data and their preprocessing takes a lot of manual effort. Unlike with numerical data where the distance-based outliers prevail, the records in questionnaires have to be assessed from various perspectives that do not relate so much. We identify the most frequent types of errors in questionnaires. For each of them, we suggest different outliers detection methods ranking the records with the usage of normalized scores. Considering the similarity between pairs of outlier scores (some are highly uncorrelated), we propose an ensemble based on the union of outliers detected by different methods. Our outlier detection framework consists of some well-known algorithms but we also propose novel approaches addressing the typical issues of questionnaires. The selected methods are based on distance, entropy, and probability. The experimental section describes the process of assembling the methods and selecting their parameters for the final model detecting significant outliers in the real-world HBSC dataset.
Vojtech Uher, Pavla Drázdilová, Jan Platos, Petr Badura
Expert Syst. Appl.3
2022 Editorial for the the special issue of WWW journal on Computational Aspects of Network Science (CAoNS)
Apostolos N. Papadopoulos, Richard Chbeir, Jan Platos, Václav Snásel
World Wide Web3
2021 Population data mobility retrieval at territory of Czechia in pandemic COVID-19 period
abstract
This article describes the methodology and the possibilities of collecting operation data in a mobile network provider. First, the architecture and the principles used in the system are described. The precision analysis of the population commuting in the region and during the pandemic and nonpandemic times. Moreover, several ideas about further utilization of the data will be formulated and described. Finally, a graph-based approach that describes the creation of the community structure between the people and the means of its analysis.
Jan Platos, Pavel Krömer, Miroslav Voznak, Václav Snásel
Concurr. Comput. Pract. Exp.1
2021 High-dimensional data classification model based on random projection and Bagging-support vector machine
abstract
Abstract Aiming at the long training time when classifying high‐dimensional data, a parallel classification model is proposed based on random projection and Bagging‐support vector machine (SVM) to process high‐dimensional data. The model first uses random projection to project the input data into the low‐dimensional space. Then, we used the Bagging method to construct multiple training data subsets and used SVM to train the training subset in parallel and generate several subclassifiers. Finally, various classifiers vote to determine the category of the test sample. The model has been verified using two standard datasets. The experimental results show that the model can significantly improve the training speed and classification performance of high‐dimensional data with little accuracy loss.
Jan Platos
Concurr. Comput. Pract. Exp.2
2021 Quantum inspired meta-heuristic approaches for automatic clustering of colour images
abstract
In this article, quantum inspired incarnations of two swarm based meta-heuristic algorithms, namely, Crow Search Optimization Algorithm and Intelligent Crow Search Optimization Algorithm have been proposed for automatic clustering of colour images. The performance and effectiveness of the proposed algorithms have been judged by experimenting on 15 Berkeley images and five publicly available real life images of different sizes. The validity of the proposed algorithms has been justified with the help of four different cluster validity indices, namely, Pakhira Bandyopadhyay Maulik, I-index, Silhouette and CS-measure. Moreover, Sobol's sensitivity analysis has been performed to tune the parameters of the proposed algorithms. The experimental results prove the superiority of proposed algorithms with respect to optimal fitness, computational time, convergence rate, accuracy, robustness, t -test and Friedman test. Finally, the efficacy of the proposed algorithms has been proved with the help of quantitative evaluation of segmentation evaluation metrics.
Alokananda Dey, Sandip Dey, Siddhartha Bhattacharyya 0001, Jan Platos, Václav Snásel
Int. J. Intell. Syst.4
2020 Self-organizing Migrating Algorithm for the Single Row Facility Layout Problem
abstract
Single row facility layout problem is an important problem encountered in facility design, factory construction, production optimization, and other areas. At the same time, it is a challenging NP-hard combinatorial optimization problem that has been addressed by many advanced algorithms. In practical scenarios, real-world problems can be cast as single row facility location problem instances with different high-level properties and efficient algorithms that can solve them are sought. This work uses a variant of the self-organizing migration algorithm developed recently for permutation problems to tackle the single row facility layout problem and evaluates its accuracy and performance.
Pavel Krömer, Jan Platos, Václav Snásel
CEC2
2020 Solving the single row facility layout problem by differential evolution
abstract
Differential evolution is an efficient evolutionary optimization paradigm that has shown a good ability to solve a variety of practical problems, including combinatorial optimization ones. Single row facility layout problem is an NP-hard permutation problem often found in facility design, factory construction, production optimization, and other areas. Real-world problems can be cast as large single row facility location problem instances with different high-level properties and efficient algorithms that can solve them efficiently are needed. In this work, the differential evolution is used to solve the single row facility location problem and the ability of three different variants of the algorithm to evolve solutions to various problem instances is studied.
Pavel Krömer, Jan Platos, Václav Snásel
GECCO2
2020 JPEG steganography with particle swarm optimization accelerated by AVX
abstract
Summary Digital steganography aims at hiding secret messages in digital data transmitted over insecure channels. The JPEG format is prevalent in digital communication, and images are often used as cover objects in digital steganography. Optimization methods can improve the properties of images with embedded secret but introduce additional computational complexity to their processing. AVX instructions available in modern CPUs are, in this work, used to accelerate data parallel operations that are part of image steganography with advanced optimizations.
Václav Snásel, Pavel Krömer, Jakub Safarik, Jan Platos
Concurr. Comput. Pract. Exp.4
2020 Lightweight Spectral-Spatial Squeeze-and- Excitation Residual Bag-of-Features Learning for Hyperspectral Classification
abstract
Of late, convolutional neural networks (CNNs) find great attention in hyperspectral image (HSI) classification since deep CNNs exhibit commendable performance for computer vision-related areas. CNNs have already proved to be very effective feature extractors, especially for the classification of large data sets composed of 2-D images. However, due to the existence of noisy or correlated spectral bands in the spectral domain and nonuniform pixels in the spatial neighborhood, HSI classification results are often degraded and unacceptable. However, the elementary CNN models often find intrinsic representation of pattern directly when employed to explore the HSI in the spectral-spatial domain. In this article, we design an end-to-end spectral-spatial squeeze-and-excitation (SE) residual bag-of-feature (S3EResBoF) learning framework for HSI classification that takes as input raw 3-D image cubes without engineering and builds a codebook representation of transform feature by motivating the feature maps facilitating classification by suppressing useless feature maps based on patterns present in the feature maps. To boost the classification performance and learn the joint spatial-spectral features, every residual block is connected to every other 3-D convolutional layer through an identity mapping followed by an SE block, thereby facilitating the rich gradients through backpropagation. Additionally, we introduce batch normalization on every convolutional layer (ConvBN) to regularize the convergence of the network and scale invariant BoF quantization for the measure of classification. The experiments conducted using three well-known HSI data sets and compared with the state-of-the-art classification methods reveal that S3EResBoF provides competitive performance in terms of both classification and computation time.
Swalpa Kumar Roy, Subhrasankar Chatterjee, Siddhartha Bhattacharyya 0001, Bidyut B. Chaudhuri, Jan Platos
IEEE Trans. Geosci. Remote. Sens.5
2020 High-Dimensional Text Clustering by Dimensionality Reduction and Improved Density Peak
abstract
This study focuses on high-dimensional text data clustering, given the inability of K-means to process high-dimensional data and the need to specify the number of clusters and randomly select the initial centers. We propose a Stacked-Random Projection dimensionality reduction framework and an enhanced K-means algorithm DPC-K-means based on the improved density peaks algorithm. The improved density peaks algorithm determines the number of clusters and the initial clustering centers of K-means. Our proposed algorithm is validated using seven text datasets. Experimental results show that this algorithm is suitable for clustering of text data by correcting the defects of K-means.
Jan Platos
Wirel. Commun. Mob. Comput.2
2019 Random Key Self-Organizing Migrating Algorithm for Permutation Problems
abstract
Self-organizing migrating algorithm (SOMA) is a modern stochastic optimization algorithm. It is built upon the principles of evolutionary and swarm computation and has been successfully applied to a variety of theoretical and practical optimization problems. The candidate solutions in SOMA are real- valued and the use of the algorithm for continuous optimization is straightforward. Its application to combinatorial optimization, on the other hand, requires a translation of candidate solutions from continuous search space to discrete problem solution space. In this work, a version of SOMA suitable for permutation problems is proposed and evaluated on two well-known hard permutation problems.
Pavel Krömer, Jan Janousek, Jan Platos
CEC3
2019 Enhancement of dronogram aid to visual interpretation of target objects via intuitionistic fuzzy hesitant sets
abstract
In this paper, we address the hesitant information in enhancement task often caused by differences in image contrast. Enhancement approaches generally use certain filters which generate artifacts or are unable to recover all the objects details in images. Typically, the contrast of an image quantifies a unique ratio between the amounts of black and white through a single pixel. However, contrast is better represented by a group of pixels. We have proposed a novel image enhancement scheme based on intuitionistic hesitant fuzzy sets (IHFSs) for drone images (dronogram) to facilitate better interpretations of target objects. First, a given dronogram is divided into foreground and background areas based on an estimated threshold from which the proposed model measures the amount of black/white intensity levels. Next, we fuzzify both of them and determine the hesitant score indicated by the distance between the two areas for each point in the fuzzy plane. Finally, a hyperbolic operator is adopted for each membership grade to improve the photographic quality leading to enhanced results via defuzzification. The proposed method is tested on a large drone image database. Results demonstrate better contrast enhancement, improved visual quality, and better recognition compared to the state-of-the-art methods.
Biswajit Biswas, Siddhartha Bhattacharyya 0001, Jan Platos, Václav Snásel
Inf. Sci.3
2018 LZ77 Like Lossy Transformation of Quality Scores
abstract
The current development in the Next-Generation Sequencing (NGS) technologies and the gradual growth of its use leads to the production of a huge amount of sequencing data. There is a need to efficiently transfer and store these data. This article introduces a novel lossy transformation algorithm for quality scores in sequencing data. Asymptotically, the algorithm preserves the likelihood of occurrence of particular quality score in individual positions of quality sequences. Such a model may be advantageous for sequencing data with very high coverage, such as targeted amplicon sequencing data. In experimental results, we show the comparison of characteristics of this algorithm with other algorithms performing lossy compression of the quality sequences. The proposed algorithm can be easily integrated into current sequencing pipelines. In this work we apply the algorithm to SAM files, which are then compressed into BAM files. The goal of the algorithm is to modify the data so that the subsequent Deflate algorithm application achieves a better compression ratio while minimizing the negative effects in the subsequent variant calling.
Michal Vasinek, Jan Platos
DCC2
2018 An acceleration of quasigroup operations by residue arithmetic
abstract
Summary Quasigroup operations are essential for a wide range of cryptographic procedures that includes cryptographic hash functions, electronic signatures, pseudorandom number generators, and stream and block ciphers. Quasigroup cryptography achieves high levels of security at low memory and computational costs by an iterative application of quasigroup operations to streams and blocks of data. The use of large quasigroups can further improve the strength of cryptographic operations. However, the order of used quasigroups is the main factor affecting the memory requirements of quasigroup cryptographic schemes. Alternative quasigroup representations that do not store their multiplication tables in computer memory yield increased computational costs. In any case, an efficient implementation of quasigroup operations is critical for practical applications of quasigroup cryptography. Residue number systems allow a fast, concurrent realization of addition and multiplication. In this work, residue arithmetic is used to accelerate quasigroup operations, and an efficient computational approach to their implementation, designed with respect to the extended instruction sets of modern processors, is proposed.
Pavel Krömer, Jan Platos, Jana Nowaková, Václav Snásel
Concurr. Comput. Pract. Exp.2
2016 Genetic algorithm for entropy-based feature subset selection
abstract
The data-driven society of today generates very large volumes of high-dimensional data. Its efficient processing by established methods represents an increasing challenge and novel advanced approaches are needed. Feature selection is a traditional data pre-processing strategy that can be used to reduce the volume and complexity of data. It selects a subset of data features so that data volume is reduced but its information content maintained. Evolutionary feature selection methods have already shown good ability to identify in very-high-dimensional data sets feature subsets according to selected criteria. Their efficiency depends, among others, on feature subset representation and objective function definition. This work employs a recent genetic algorithm for fixed-length subset selection to find feature subsets on the basis of their entropy, estimated by a fast data compression method. The reasonability of this new fitness criterion and the usefulness of selected feature subsets for practical data mining is evaluated using well-known data sets and several widely-used classification algorithms.
Pavel Krömer, Jan Platos
CEC2
2016 Evolutionary Feature Subset Selection with Compression-based Entropy Estimation
abstract
Modern massive data sets often comprise of millions of records and thousands of features. Their efficient processing by traditional methods represents an increasing challenge. Feature selection methods form a family of traditional instruments for data dimensionality reduction. They aim at selecting subsets of data features so that the loss of information, contained in the full data set, is minimized. Evolutionary feature selection methods have shown good ability to identify feature subsets in very-high-dimensional data sets. Their efficiency depends, among others, on a particular optimization algorithm, feature subset representation, and objective function definition. In this paper, two evolutionary methods for fixed-length subset selection are employed to find feature subsets on the basis of their entropy, estimated by a fast data compression algorithm. The reasonability of the fitness criterion, ability of the investigated methods to find good feature subsets, and the usefulness of selected feature subsets for practical data mining, is evaluated using two well-known data sets and several widely-used classification algorithms.
Pavel Krömer, Jan Platos
GECCO2
2015 Generalized Context Transformations - Enhanced Entropy Reduction
abstract
Context transformations is a very simple data transformation method that we presented recently and it is used to decrease uncertainty in input data. The transformation is based on exchange of two different di-grams. This paper is focused on new consequences of the relationships discovered subsequently. We were able to find a mathematical model which predicts the efficiency of each transformation. The new type of the transformation, Generalized context transformation, developed recently is more efficient than the previous one and it is able to remove almost all redundancy based on the symbols mutual information. The newly developed algorithm is computationally and entropic ally more efficient than the previous one.
Michal Vasinek, Jan Platos
DCC2
2014 Local representatives in weighted networks
abstract
The main features of current real-world networks are their large sizes and structures, which show varying degrees of importance of the nodes in their surroundings. The topic of evaluating the importance of the nodes offers many different approaches that usually work with unweighted networks. We present a novel, simple and straightforward approach for the evaluation of the network's nodes with a focus on local properties in their surroundings. The presented approach is intended for weighted networks where the weight can be interpreted as the proximity between the nodes. Our suggested x-representativeness then takes into account the degree of the node, its nearest neighbors and one other parameter which we call the x-representativeness base. Following that, we also present experiments with three different real-world networks. The aim of these experiments is to show that the x-representativeness can be used to deterministically reduce the network to differently sized samples of representatives, while maintaining the topological properties of the original network.
Sarka Zehnalova, Milos Kudelka, Jan Platos, Zdenek Horak
ASONAM3
2014 Genetic Algorithm for the Column Subset Selection Problem
abstract
The column subset selection problem is a well-known hard optimization problem of selecting an optimal subset of k columns from the matrix A (m x n), k
Pavel Krömer, Jan Platos, Václav Snásel
CISIS2
2014 Entropy Reduction Using Context Transformations
abstract
In this paper we propose a reversible string transformation method that could be used to decrease entropy in input messages. The method is based on context information included in the structure we call a context map. The method is used to manipulate with symbols distribution. Since the change in distribution of symbols is followed by change in entropy, we could use cases when entropy is decreasing and then employ entropy compression algorithm on the transformed string.
Michal Vasinek, Jan Platos
DCC2
2014 Genetic algorithm for sampling from scale-free data and networks
abstract
A variety of real-world data and networks can be described by a heavy-tailed probability distribution of its values, vertex degrees, or other significant properties, that follows the power law. Such a scale-free data and networks can be found in both natural phenomena such as protein interaction networks and gene regulation networks and man-made structures like the Internet, language, and various social networks. An efficient analysis of large scale data and networks is often impractical and various heuristic and metaheuristc sampling techniques are deployed to select smaller subsets of the data for analysis and visualisation. A key goal of data and network sampling is to select such a subset of the original data that would accurately represent the original data with respect to selected attributes. In this work we propose a novel genetic algorithm for scale-free data and network sampling and evaluate the algorithm in a series of computational experiments.
Pavel Krömer, Jan Platos
GECCO2
2014 Clustering using artificial bee colony on CUDA
abstract
Artificial bee colony is a meta-heuristic optimization algorithm based on the behavior of honey bee swarm. These bees work largely independently of other bees, making the algorithm suitable for parallel implementation. Within this paper, we introduce the algorithm itself and its subsequent parallelization utilizing the CUDA platform. The runtime speedup is demonstrated on several commonly used test functions for optimization. The algorithm is subsequently applied to the problem of clustering real data.
Jan Janousek, Jan Platos, Václav Snásel
SMC2
2014 Solving the p-median problem by a simple differential evolution
abstract
Differential evolution is a real-parameter metaheuristic optimization method with a history of successful applications in many different domains. The p-median problem is a well-known combinatorial optimization problem with several possible formulations and many practical applications in areas such as operational research and planning. It has been also used as a testbed for various heuristic and metaheuristic optimization algorithms. This work uses a simple variant of the differential evolution to solve the p-median problem and evaluates the efficiency of this method in a series of computational experiments.
Pavel Krömer, Jan Platos
SMC2
2014 Local representativeness in vector data
abstract
The amount of large-scale real data around us is increasing in size very quickly, as is the necessity to reduce its size by obtaining a representative sample. Such sample allows us to use a great variety of analytical methods, the direct application of which on original data would be unfeasible. Conventional sampling methods provide non-deterministic results trying to preserve selected characteristics of the input dataset. We present a novel, simple, straightforward and deterministic approach with the same goal. It is not sampling in the true sense but a reduction of vector data, which maintains very well internal data structures (clusters and density). The approach is based on analyzing the nearest neighbors. Our suggested x-representativeness then takes into account the local density of the data and nearest neighbors of individual data objects. Following that, we also present experiments with two different datasets. The aim of these experiments is to show that the x-representativeness can be used to deterministically reduce the datasets to differently sized samples of representatives, while maintaining properties of the original datasets.
Sarka Zehnalova, Milos Kudelka, Jan Platos
SMC3
2014 Data Parallel density-based genetic clustering on CUDA Architecture
abstract
SUMMARY Evolutionary clustering algorithms have been proven as a good ability to find clusters in data. Among their advantages belong the abilities to adapt to data and to determine the number of clusters automatically, thus requiring lessa prioriassumptions about analyzed objects than traditional clustering methods. Unfortunately, such a clustering by genetic algorithms and evolutionary algorithms in general suffers from high computational costs when it comes to recurrent fitness function evaluation. Computing on graphic processing units (GPUs) is a recent programming and development paradigm bringing high performance parallel computing closer to general audience. Modern general purpose GPUs are composed of tens to thousands of computational cores that can execute programs in parallel using the single instruction multiple data parallel processing approach. General purpose GPU programs need to be designed and implemented in a data parallel way and with respect to the architecture of target devices to fully utilize their high performance. This study presents a design, implementation, and evaluation of a data parallel genetic algorithm for density‐based clustering. The algorithm was implemented and evaluated on the nVidia Compute Unified Device Architecture (CUDA) platform. Copyright © 2013 John Wiley & Sons, Ltd.
Pavel Krömer, Jan Platos, Václav Snásel
Concurr. Comput. Pract. Exp.2
2013 Parallel Differential Evolution in Unified Parallel C
abstract
Distributed environments and emerging highly-parallel platforms provide a suitable hardware infrastructure for parallel Evolutionary Computation. Partitioned Global Address Space model is a well-known parallel computing model used to implement scalable algorithms for many-core systems and clusters. This study investigates the Unified Parallel C programming language as a tool for implementation of scalable evolutionary algorithms for high-dimensional problems. The design concepts and initial implementation are demonstrated on the Differential Evolution algorithm. The mapping of Differential Evolution concepts to Unified Parallel C features is presented and three variants of parallel Differential Evolution for many-core shared memory systems and clusters of computers with distributed memory are implemented and evaluated in the environment of a small real-world cluster.
Pavel Krömer, Jan Platos, Václav Snásel
IEEE Congress on Evolutionary Computation2
2013 Compression-based similarity in EEG signals
abstract
The electrical activity of brain or EEG signal is very complex data system that may be used to many different applications such as device control using mind. It is not easy to understand and detect useful signals in continuous EEG data stream. In this paper, we are describing an application of data compression which is able to recognize important patterns in this data. The proposed algorithm uses Lampel-Ziv complexity for complexity measurement and it is able to successfully detect patterns in EEG signal.
Michal Prilepok, Jan Platos, Václav Snásel, Ibrahim Salem Jahan
ISDA2
2013 Spam Detection using Data Compression and Signatures
abstract
In this article, we introduce a novel method for spam detection based on a combination of Bayesian filtering, signature trees, and data compression–based similarity. Bayesian filtering is one of the most popular and most efficient algorithms for dealing with spam detection. The problem with Bayesian filtering is that it is unable to classify any e-mail without doubt and sometimes spam e-mails are classified as regular e-mails. This novel method sorts out this problem by using signature trees and data compression–based similarity. The main result of this article is an up to 99% improvement in spam detection precision using this novel method.
Michal Prilepok, Petr Berek, Jan Platos, Václav Snásel
Cybern. Syst.3
2012 An ACO inspired weighting approach for the spectral partitioning of co-authorship networks
abstract
Spectral partitioning is a well known method in the area of graph and matrix analysis. Several approaches based on spectral partitioning and spectral clustering were used to detect structures in real world networks and databases. In this paper, we use the spectral partitioning to detect communities in a co-authorship network. The partitioning depends heavily on the weighting of the underlying network. We use an intuitive weighting scheme based on the ant colony optimization and show the communities found by spectral partitioning when using the ACO inspired weighting and when using trivial weighting based on the number of interactions between the authors.
Pavel Krömer, Václav Snásel, Jan Platos, Milos Kudelka, Zdenek Horak
IEEE Congress on Evolutionary Computation3
2012 Neural PCA and Maximum Likelihood Hebbian Learning on the GPU
Pavel Krömer, Emilio Corchado, Václav Snásel, Jan Platos, Laura García-Hernández
ICANN (2)4
2012 Photovoltaic Power Plant Output Estimation by Neural Networks and Fuzzy Inference
Lukás Prokop 0001, Stanislav Misák, Tomás Novosád, Pavel Krömer, Jan Platos, Václav Snásel
IDEAL5
2012 Genetic algorithm for clustering accelerated by the CUDA platform
abstract
Unsupervised clustering of large data sets is a complicated NP-hard task. Due to its complexity, various metaheuristic machine learning algorithms have been used to automate or aid the clustering process. Genetic and evolutionary algorithms have been deployed to find clusters in data sets with success. However, also evolutionary clustering suffers from the high computational demands when it comes to fitness function evaluation. The GPU computing is a recent programming and development paradigm introducing high performance parallel computing to general audience. This work presents an initial design and implementation of a genetic algorithm for density based clustering on the GPU using the nVidia CUDA platform.
Pavel Krömer, Jan Platos, Václav Snásel
SMC2
2012 Searching for optimal alphabet for data compression using simulated annealing
abstract
Data compression is very important today and it will be even more important in the future. Textual data use only limited alphabet - total number of used symbols (letters, numbers, diacritics, dots, spaces, etc.). In most languages, letters are joined into syllables and words. All three approaches are useful in text compression, but none of them is the best for any file. This paper describes a variant of algorithm for evolving alphabet from characters, 2-grams and 3-grams, which is optimal for compression of text files. We used Simulated Annealing for this evolution of the alphabet. The efficiency of the new variant will be tested on four compression algorithms. The achieved results are very promising.
Jan Platos, Pavel Krömer
SMC1
2012 A PSO-based document classification algorithm accelerated by the CUDA Platform
abstract
Document classification is a well-known problem that is focused on assigning predefined labels or categories to the documents found in the searched collection. Many classical algorithms were developed for solving of this problem. They usually have large time complexity and with increasing number of documents it is necessary to find algorithm which are able to find solution in reasonable time. Such algorithms are usually inspired by biological processes. Even such meta-heuristics algorithms become too slow when the number of documents is really large and it is necessary to optimize them for faster processing. This paper describes a document classification algorithm based on Particle Swarm Optimization with implementation of one and two GPUs.
Jan Platos, Václav Snásel, Tomás Jezowicz, Pavel Krömer, Ajith Abraham
SMC1
2012 Artificially evolved soft computing models for photovoltaic power plant output estimation
abstract
Renewable energy sources are becoming a significant part of todays energy mix. The unstable production of many renewable energy sources including photovoltaic and wind power plants puts increased demands on power transmission systems and on the power grid as a whole. Soft computing methods can contribute to the prediction of electric energy production of renewable resources and therefore to the reliability of the power transmission networks. This work compares two soft computing methods that utilize genetic programming to evolve predictors of a selected renewable energy resource that meets the real world criterion of high output variance and relatively large installed power (in context of the power distribution system of the Czech Republic).
Lukás Prokop 0001, Stanislav Misák, Tomás Novosád, Pavel Krömer, Jan Platos, Václav Snásel
SMC5
2012 Fast decoding algorithms for variable-lengths codes
Jirí Walder, Michal Krátký, Radim Baca, Jan Platos, Václav Snásel
Inf. Sci.4
2011 Differential evolution for the linear ordering problem implemented on CUDA
abstract
Linear Ordering Problem (LOP) is a well know NP-hard problem combinatorial optimization problem attractive for its complexity, rich library of test data and variety of real world applications. In this paper, we use differential evolution accelerated by the GPU using the nVidia CUDA platform to find good LOP solutions. The well known LOLIB library was used to evaluate the efficiency and precision of the approach in solving LOP instances.
Pavel Krömer, Jan Platos, Václav Snásel
IEEE Congress on Evolutionary Computation2
2011 Many-threaded implementation of differential evolution for the CUDA platform
abstract
Differential evolution is an efficient populational meta -- heuristic optimization algorithm successful in solving difficult real world problems. Due to the simplicity of its operations and data structures, it is suitable for a parallel implementation on multicore systems and on the GPU. In this paper, we design a simple yet highly parallel implementation of the differential evolution using the CUDA architecture. We demonstrate the speedup obtained by the proposed parallelization of the differential evolution on an NP hard combinatorial optimization problem and on a benchmark function of many variables.
Pavel Krömer, Václav Snásel, Jan Platos, Ajith Abraham
GECCO3
2011 Tension estimation by fuzzy predictors in heavy facilities
abstract
We use genetic programming to evolve accurate predictors (fuzzy rules) that are deployed to estimate the tension in a power plant generator. The meta-heuristic is compared to the finite element method that was used to compute estimated tension. In contrast to the finite element method, the fuzzy predictor (once found) approximates the tension in the facility quickly and with sufficient precision.
Jan Platos, Václav Snásel, Pavel Krömer, Petr Fiala
HIS1
2011 Optimizing alphabet using genetic algorithms
abstract
Data compression algorithms were usually designed for data processing symbol by symbol. The input symbols of these algorithms are usually taken from the ASCII table, i.e. the size of the input alphabet is 256 symbols which are representable by 8-bit numbers. Several other techniques were developed-syllable-based compression, which uses the syllable as a basic compression symbol, and word-based compression, which uses words as basic symbols. These three approaches are strictly bounded and no overlap is allowed. This may be a problem because it may be helpful to have an overlap between them and use a character-based approach with a few symbols as a sequence of characters. This paper describes an algorithm that looks for the optimal alphabet for different text files. The alphabet may contain characters and 2-grams.
Jan Platos, Pavel Krömer
ISDA1
2011 Towards the analysis of co-authorship networks by iterative spectral partitioning
abstract
Spectral partitioning is a well known method in the area of graph and matrix analysis. Several approaches based on spectral partitioning and spectral clustering were used to detect structures and mine data from real world networks. In this paper, we use a simple spectral decomposition to analyze a co-authorship network. We use a straightforward approach based on algebraic connectivity and characteristic valuation and show that even this simple form of spectral partitioning is useful for the analysis of relations and communities in a co-authorship networks.
Václav Snásel, Pavel Krömer, Jan Platos
ISDA3
2011 Fuzzy classification by evolutionary algorithms
abstract
Fuzzy sets and fuzzy logic can be used for efficient data classification by fuzzy rules and fuzzy classifiers. This paper presents an application of genetic programming to the evolution of fuzzy classifiers based on extended Boolean queries. Extended Boolean queries are well known concept in the area of fuzzy information retrieval. An extended Boolean query represents a complex soft search expression that defines a fuzzy set on the collection of searched documents. We interpret the data mining task as a fuzzy information retrieval problem and we apply a proven method for query induction from data to find useful fuzzy classifiers. The ability of the genetic programming to evolve useful fuzzy classifiers is demonstrated on two use cases in which we detect faulty products in a product processing plant and discover intrusions in a computer network.
Pavel Krömer, Jan Platos, Václav Snásel, Ajith Abraham
SMC2
2010 Towards intrusion detection by information retrieval and genetic programming
abstract
Fuzzy classifiers and fuzzy rules are powerful tools in data mining and knowledge discovery. In this work, intrusion detection is approached as a data mining task and genetic programming is deployed to evolve fuzzy classifiers for detection of intrusion and security problems. We train the fuzzy classifier on a data set modeled as a fuzzy information retrieval collection and investigate its ability to detect illegitimate actions. Proposed approach is experimentally evaluated on the popular KDD Cup intrusion detection data set.
Pavel Krömer, Jan Platos, Václav Snásel, Ajith Abraham
IAS2
2010 Fast intrusion detection system based on Flexible Neural Tree
abstract
Computer security is very important in these days. Computers are used probably in any industry and their protection against attacks is very important task. The protection usually consist in several levels. The first level is preventions. Intrusion detection system (IDS) may be used as next level. IDS is useful in detection of intrusions, but also in monitoring of security issues and the traffic. This paper present IDS based on Flexible Neural Trees. Flexible neural tree is hierarchical neural network, which is automatically created using evolutionary algorithms to solving of defined problem. This is very important, because it is not necessary to set the structure and the weights of neural networks prior the problem is solved. The accuracy of proposed technique is always above 98% and the speed of decision making process enable its using in real-time applications.
Tomás Novosád, Jan Platos, Václav Snásel, Ajith Abraham
IAS2
2010 Scaling IDS construction based on Non-negative Matrix factorization using GPU computing
abstract
Attacks on the computer infrastructures are becoming an increasingly serious problem. Whether it is banking, e-commerce businesses, health care, law enforcement, air transportation, or education, we are all becoming increasingly reliant upon the networked computers. The possibilities and opportunities are limitless; unfortunately, so too are the risks and chances of malicious intrusions. Intrusion detection is required as an additional wall for protecting systems despite of prevention techniques and is useful not only in detecting successful intrusions, but also in monitoring attempts to security, which provides important information for timely countermeasures. This paper presents some improvements to some of our previous approaches using a Non-negative Matrix factorization approach. To improve the performance (detection accuracy) and computational speed (scaling) a GPU implementation is detailed. Empirical results indicate that the speedup was up to 500x for the training phase and up to 190x for the testing phase.
Jan Platos, Pavel Krömer, Václav Snásel, Ajith Abraham
IAS1
2010 Evolutionary improvement of search queries and its parameters
abstract
The formulation of user queries is an important part of the information retrieval process. In the complex environment of the World Wide Web and other large data collections, it is often not easy for the users to express their information needs in an optimal way. In this paper, we investigate evolutionary algorithms (in particular genetic programming) as a tool for the optimization of user queries and seek for its good settings.
Pavel Krömer, Václav Snásel, Jan Platos, Ajith Abraham
HIS3
2010 Search personalization in hyperlinked environments by relevance propagation and ant colony optimization
abstract
Personalization is a promising way of improvement of the search services in large document collections and on the Web. User modeling is in the core of many personalization efforts because accurate user model can provide essential information for user specific search adjustments and result set processing. In this paper, we propose and study user modeling technique based on click-through data, relevance propagation and ant colony optimization.
Pavel Krömer, Václav Snásel, Jan Platos, Suhail S. J. Owais
HIS3
2010 Genetic Algorithms Evolving Quasigroups with Good Pseudorandom Properties
Václav Snásel, Jiri Dvorský, Eliska Ochodkova, Pavel Krömer, Jan Platos, Ajith Abraham
ICCSA (3)5
2010 Iris recognition on GPU with the usage of Non-Negative Matrix Factorization
abstract
In this paper, we describe an alternative method of the recognition of human irises with the usage of Non-Negative Matrix Factorization. The proposed method has been implemented on graphic processor unit (GPU) which makes the method usable in the real world due to short computation time.
Petr Gajdos, Jan Platos, Pavel Moravec 0001
ISDA2
2010 Data mining using NMF and generalized matrix inverse
abstract
Non-negative matrix factorization is an important method helpful in the analysis of high dimensional datasets. It has a number of applications including pattern recognition, data clustering, information retrieval or computer security. One its significant drawback lies in its computational complexity. In this paper, we introduce a new method allowing fast approximate transformation from input space to feature space defined by non-negative matrix factorization and discuss some examples of its application.
Pavel Krömer, Jan Platos, Václav Snásel
ISDA2
2009 Detecting Insider Attacks Using Non-negative Matrix Factorization
abstract
It is a fact that vast majority of attention is given to protecting against external threats, which are considered more dangerous. However, some industrial surveys have indicated they have had attacks reported internally. Insider Attacks are an unusual type of threat which are also serious and very common. Unlike an external intruder, in the case of internal attacks, the intruder is someone who has been entrusted with authorized access to the network. This paper presents a Non-negative Matrix factorization approach to detect inside attacks. Comparisons with other established pattern recognition techniques reveal that the Non-negative Matrix Factorization approach could be also an ideal candidate to detect internal threats.
Jan Platos, Václav Snásel, Pavel Krömer, Ajith Abraham
IAS1
2009 Differential Evolution and Genetic Algorithms for the Linear Ordering Problem
Václav Snásel, Pavel Krömer, Jan Platos
KES (1)3
2008 Matrix Factorization Approach for Feature Deduction and Design of Intrusion Detection Systems
abstract
Current Intrusion Detection Systems (IDS) examine all data features to detect intrusion or misuse patterns. Some of the features may be redundant or contribute little (if anything) to the detection process. The purpose of this research is to identify important input features in building an IDS that is computationally efficient and effective. This paper propose a novel matrix factorization approach for feature deduction and design of intrusion detection systems. Experiment results indicate that the proposed method is efficient.
Václav Snásel, Jan Platos, Pavel Krömer, Ajith Abraham
IAS2
2008 Implicit User Modelling Using Hybrid Meta-Heuristics
abstract
The requirements imposed on information retrieval systems are increasing steadily. The vast number of documents in today's large databases and espe-cially on World Wide Web causes notable problems when searching for concrete information. It is difficult to find satisfactory information that accurately matches user information needs even if it is present in the database. One of the key elements when searching the web is proper formulation of user queries. Search effectiveness can be seen as the accuracy of matching user information needs against the retrieved information. Personalized search applications can notably contribute to the improvement of web search effectiveness. In this paper, we investigate two user modelling and search optimization techniques based on genetic algorithms and ant colony optimization.
Pavel Krömer, Václav Snásel, Jan Platos, Ajith Abraham
HIS3
2008 Implementing Boolean Matrix Factorization
Roman Neruda, Václav Snásel, Jan Platos, Pavel Krömer, Dusan Húsek, Alexander A. Frolov
ICANN (1)3
2008 Compression of small text files
Jan Platos, Václav Snásel, Eyas El-Qawasmeh
Adv. Eng. Informatics1
2007 Optimizing Interleaver for Turbo Codes by Genetic Algorithms
abstract
Since the appearance in 1993, first approaching the Shannon limit, the Turbo Codes give a new direction for the channel encoding field, especially since they were adopted for multiple norms of telecommunications, such as deeper communication. To obtain an excellent performance it is necessary to design robust turbo code interleaver. We are investigating genetic algorithms as a promising optimization method to find good performing interleaver for the large frame sizes. In this paper, we present our work, compare with several previous approaches and present experimental results.
Pavel Krömer, Václav Snásel, Jan Platos, Nabil Ouddane
ICTAI (1)3
2007 Implicit User Modelling for Web Search Improvement
abstract
The requirements imposed on information retrieval systems are increasing steadily. The vast number of documents in today's large databases and especially on World Wide Web causes notable problems when searching for concrete information. It is difficult to find satisfactory information that accurately matches user information needs even if it is present in the database. One of the key elements when searching the web is proper formulation of user queries. Search effectiveness can be seen as the accuracy of matching user information needs against the retrieved information. Personalized search applications can notably contribute to the improvement of web search effectiveness. It has been shown, that genetic programming can evolve search queries towards users interests captured by the means of relevance. In this paper, we propose user modelling technique based on relevance estimation and provide experimental results in web search framework with evolutionary query optimization.
Pavel Krömer, Václav Snásel, Jan Platos, Suhail S. J. Owais
ISDA3