Václav Snásel

dblp:s/VaclavSnasel · DBLP profile ↗
← Back
31ranked-venue papers in the field
1as first author
9since 2021 · last 2026
0000-0002-9600-8319ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 11Database Systems & Data Management · 9Data Mining & Knowledge Discovery · 4 (1 first)Other / Interdisciplinary · 4Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 A lightweight feature selection method based on rankability
abstract
Feature selection, as one of the essential dimensionality reduction techniques, hasbecome one popular yet challenging area in the field, such as data mining andmachine learning. Unlike feature extraction (such as principle component analysis andnon-negative matrix factorization), preserving the entire information but losing thefeature relevance. Data processing in feature selection will lose information, leading tothe need to develop lightweight, efficient, and practical methods that preserve the datainformation as much as possible while performing dimensional reduction. In this paper,we propose a rankability-based feature selection method. The rankability concept wasproposed in 2019, similar to the entropy concept, and has not been studied widely yet.The proposed method is lightweight in terms of complexity, which requires no iterativeoptimization or auxiliary estimator tools.We experimented with sixteen datasets and compared our results with four otheralgorithms. The results show that our rankability-based feature selection methodoutperforms the fuzzy entropy-based method on five datasets in eight, and the averageaccuracy increased by 0.1482, 0.1078, and 0.1157, respectively. Then, in the varieddimension-reducing experiments, the proposed method shows superiority on fourdatasets out of eight and is competitive with others on two datasets out of eight.
Lingping Kong 0001, Juan D. Velásquez 0001, Irina Perfilieva, Millie Pant, Jeng-Shyang Pan 0001, Václav Snásel
Inf. Sci.6
2024 A hierarchical overlapping community detection method based on closed trail distance and maximal cliques
abstract
An important feature of real networks is their hierarchy and the existence of overlapping communities. Hierarchical agglomerative clustering is one way to determine the hierarchy of a network. To ensure the existence of overlapping communities, it is appropriate to choose the base elements for clustering – edges, cliques, etc. These base elements can then have common vertices and naturally provide the possibility of overlap. The proposed community detection method uses hierarchical agglomerative clustering on the 2-edge-connected component of the graph. Communities are constructed from maximal cliques as base elements. Novel dissimilarities for hierarchical agglomerative clustering were introduced for the merging of cliques. The dissimilarities use the size of the overlapped cliques and closed trail distance to express dissimilarity between communities in networks. The single linkage approach contains and extends the results of k-CPM. The proposed algorithm utilizing deterministic dissimilarity achieves comparable or superior outcomes compared to standard algorithms used for hierarchical or overlapping community detection.
Pavla Drázdilová, Petr Prokop, Jan Platos, Václav Snásel
Inf. Sci.4
2023 Low-rank and global-representation-key-based attention for graph transformer
abstract
Transformer architectures have been applied to graph-specific data such as protein structure and shopper lists, and they perform accurately on graph/node classification and prediction tasks. Researchers have proved that the attention matrix in Transformers has low-rank properties, and the self-attention plays a scoring role in the aggregation function of the Transformers. However, it can not solve the issues such as heterophily and over-smoothing. The low-rank properties and the limitations of Transformers inspire this work to propose a Global Representation (GR) based attention mechanism to alleviate the two heterophily and over-smoothing issues. First, this GR-based model integrates geometric information of the nodes of interest that conveys the structural properties of the graph. Unlike a typical Transformer where a node feature forms a Key, we propose to use GR to construct the Key, which discovers the relation between the nodes and the structural representation of the graph. Next, we present various compositions of GR emanating from nodes of interest and α-hop neighbors. Then, we explore this attention property with an extensive experimental test to assess the performance and the possible direction of improvements for future works. Additionally, we provide mathematical proof showing the efficient feature update in our proposed method. Finally, we verify and validate the performance of the model on eight benchmark datasets that show the effectiveness of the proposed method.
Lingping Kong 0001, Varun Ojha 0001, Ruobin Gao, Ponnuthurai N. Suganthan, Václav Snásel
Inf. Sci.5
2023 Enhancing Anchor Link Prediction in Information Networks through Integrated Embedding Techniques
Van-Vang Le, Phu Pham, Václav Snásel, Unil Yun, Bay Vo
Inf. Sci.3
2022 Improvement Graph Convolution Collaborative Filtering with Weighted Addition Input
Tin T. Tran, Václav Snásel
ACIIDS (1)2
2022 Predictive intelligence in evaluation of visual perception thresholds for visual pattern recognition and understanding
Urszula Ogiela, Václav Snásel
Inf. Process. Manag.2
2022 Improved spherical search with local distribution induced self-adaptation for hard non-convex optimization with and without constraints
Abhishek Kumar 0010, Swagatam Das, Václav Snásel
Inf. Sci.3
2022 Impact of chaotic dynamics on the performance of metaheuristic optimization algorithms: An experimental analysis
abstract
Random mechanisms including mutations are an internal part of evolutionary algorithms, which are based on the fundamental ideas of Darwin’s theory of evolution as well as Mendel’s theory of genetic heritage. In this paper, we debate whether pseudo-random processes are needed for evolutionary algorithms or whether deterministic chaos, which is not a random process, can be suitably used instead. Specifically, we compare the performance of 10 evolutionary algorithms driven by chaotic dynamics and pseudo-random number generators using chaotic processes as a comparative study. In this study, the logistic equation is employed for generating periodical sequences of different lengths, which are used in evolutionary algorithms instead of randomness. We suggest that, instead of pseudo-random number generators, a specific class of deterministic processes (based on deterministic chaos) can be used to improve the performance of evolutionary algorithms. Finally, based on our findings, we propose new research questions.
Ivan Zelinka, Quoc Bao Diep, Václav Snásel, Swagatam Das, Giacomo Innocenti, Alberto Tesi, Fabio Schoen, Nikolay V. Kuznetsov
Inf. Sci.3
2021 Quantum inspired meta-heuristic approaches for automatic clustering of colour images
abstract
In this article, quantum inspired incarnations of two swarm based meta-heuristic algorithms, namely, Crow Search Optimization Algorithm and Intelligent Crow Search Optimization Algorithm have been proposed for automatic clustering of colour images. The performance and effectiveness of the proposed algorithms have been judged by experimenting on 15 Berkeley images and five publicly available real life images of different sizes. The validity of the proposed algorithms has been justified with the help of four different cluster validity indices, namely, Pakhira Bandyopadhyay Maulik, I-index, Silhouette and CS-measure. Moreover, Sobol's sensitivity analysis has been performed to tune the parameters of the proposed algorithms. The experimental results prove the superiority of proposed algorithms with respect to optimal fitness, computational time, convergence rate, accuracy, robustness, t -test and Friedman test. Finally, the efficacy of the proposed algorithms has been proved with the help of quantitative evaluation of segmentation evaluation metrics.
Alokananda Dey, Sandip Dey, Siddhartha Bhattacharyya 0001, Jan Platos, Václav Snásel
Int. J. Intell. Syst.5
2019 Enhancement of dronogram aid to visual interpretation of target objects via intuitionistic fuzzy hesitant sets
abstract
In this paper, we address the hesitant information in enhancement task often caused by differences in image contrast. Enhancement approaches generally use certain filters which generate artifacts or are unable to recover all the objects details in images. Typically, the contrast of an image quantifies a unique ratio between the amounts of black and white through a single pixel. However, contrast is better represented by a group of pixels. We have proposed a novel image enhancement scheme based on intuitionistic hesitant fuzzy sets (IHFSs) for drone images (dronogram) to facilitate better interpretations of target objects. First, a given dronogram is divided into foreground and background areas based on an estimated threshold from which the proposed model measures the amount of black/white intensity levels. Next, we fuzzify both of them and determine the hesitant score indicated by the distance between the two areas for each point in the fuzzy plane. Finally, a hyperbolic operator is adopted for each membership grade to improve the photographic quality leading to enhanced results via defuzzification. The proposed method is tested on a large drone image database. Results demonstrate better contrast enhancement, improved visual quality, and better recognition compared to the state-of-the-art methods.
Biswajit Biswas, Siddhartha Bhattacharyya 0001, Jan Platos, Václav Snásel
Inf. Sci.4
2016 An Application of Neural Network in Method for Use Case Based Effort Estimation
abstract
Effort overruns is common problem in software development. Our main intention is to support estimation by method for classification of use cases. The goal of this paper is to evaluate usage of the feed-forward neural network for the Use Case classification purposes. Experimental results show that the feed-forward neural network classifier, using softmax activation function in the output layer and hyperbolic tangent activation function in the hidden layer, offers the best classification performance.
Radoslav Strba, Svatopluk Stolfa, Jakub Stolfa, Ivo Vondrák, Václav Snásel
EJC5
2014 Constructing ordinary sum differential equations using polynomial networks
Ladislav Zjavka, Václav Snásel
Inf. Sci.2
2012 Evolution of Author's Topic in Authorship Network
abstract
There may be several reasons why people publish together. Above all, the fact that the authors share common professional interests is the main reason. In our research we work with the DBLP dataset which contains the basic bibliographic information of publications from the computer science field. These data are freely available and contain highly relevant information about publication activity from the period of nearly fifty years, even though they are not complete. One of the goals of our research is to analyze and visualize the evolution of authors and co-authorship from the point of view of research topics. We present the results of our research in this paper. One of the results is also visualization in our online FORCOA.NET system.
Sarka Zehnalova, Zdenek Horak, Milos Kudelka, Václav Snásel
ASONAM4
2012 Swarm scheduling approaches for work-flow applications with security constraints in distributed data-intensive computing environments
Hongbo Liu 0001, Ajith Abraham, Václav Snásel, Seán F. McLoone
Inf. Sci.3
2012 Fast decoding algorithms for variable-lengths codes
Jirí Walder, Michal Krátký, Radim Baca, Jan Platos, Václav Snásel
Inf. Sci.5
2010 Finding Patterns of Students' Behavior in Synthetic Social Networks
abstract
Spectral clustering is a data mining method used for finding patterns in high dimensional datasets. It has been applied effectively to solve many problems in signal processing, bioinformatics, etc. In this paper spectral clustering was implemented to find students’ patterns of behavior in an elearning system, to explore the relationship between the similarity of students’behavior and their academic performance.
Gamila Obadi, Pavla Drázdilová, Jan Martinovic, Katerina Slaninová, Václav Snásel
ASONAM5
2009 Search Results Clustering Using Nonnegative Matrix Factorization (NMF)
abstract
There are many search engines in the Web and when asked, they return a long list of search results, ranked by their relevancies to the given query. Web users have to go through the list and examine the titles and (short) snippets sequentially to identify their required results. In this paper we present how usage of Nonnegative Matrix Factorization (NMF) can be good solution for the search results clustering.
Hussam M. Dahwa Abdulla, Martin Polovincak, Václav Snásel
ASONAM3
2009 Reducing Social Network Dimensions Using Matrix Factorization Methods
abstract
Since the availability of social networks data and the range of these data have significantly grown in recent years, new aspects have to be considered. In this paper we address computational complexity of social networks analysis and clarity of their visualization. Our approach uses combination of Formal Concept Analysis and well-known matrix factorization methods. The goal is to reduce the dimension of social network data and to measure the amount of information which is lost during the reduction.
Václav Snásel, Zdenek Horak, Jana Kocibova, Ajith Abraham
ASONAM1
2008 On the efficient search of an XML twig query in large DataGuide trees
abstract
XML (Extensible Mark-up Language) has been embraced as a new approach to data modeling. Nowadays, more and more information is formatted as semi-structured data, e.g., articles in a digital library, documents on the web, and so on. Implementation of an efficient system enabling storage and querying of XML documents requires development of new techniques. Many different techniques of XML indexing have been proposed in recent years.
Radim Baca, Michal Krátký, Václav Snásel
IDEAS3
2008 Compression of small text files
Jan Platos, Václav Snásel, Eyas El-Qawasmeh
Adv. Eng. Informatics2
2007 On the Efficient Processing Regular Path Expressions of an Enormous Volume of XML Data
Michal Krátký, Radim Baca, Václav Snásel
DEXA3
2006 Efficient Processing of Narrow Range Queries in Multi-dimensional Data Structures
abstract
Multi-dimensional data structures are applied in many real index applications, i.e. data mining, indexing multimedia data, indexing of text documents and so on. Many index structures and algorithms have been proposed. There are two major approaches to multi-dimensional indexing: data structures to indexing metric and vector spaces. R-trees, R*-trees and (B)UB-trees are representatives of the vector data structures. These data structures provide efficient processing of many types of queries, i.e. point queries, range queries and so on. As far as the vector data structures are concerned, the range query retrieves all points in defined hyper box in an n-dimensional space. The narrow range query is an important type of the range query. Its processing is inefficient in vector data structures. Moreover, the efficiency decreases as the dimension of the indexed space increases. We depict an application of the signature for more efficient processing of narrow range queries. The approach puts the signature into the multi-dimensional data structures like R-tree or UB-tree but original functionalities are preserved, i.e. the range query algorithm for general range query. The novel data structure is called the signature data structure, e.g., Signature R-tree or Signature UB-tree.
Michal Krátký, Václav Snásel, Jaroslav Pokorný, Pavel Zezula
IDEAS2
2006 Semantic Analysis of Web Pages Using Web Patterns
abstract
This paper introduces a novel method for semantic analysis of Web pages. Analysis is performed with regard to unwritten and empirically proven agreement between users and Web designers using Web patterns. This method is based on extraction of patterns which are characteristics for concrete domain. Patterns provide formalization of the agreement and allow assignment of semantics to parts of Web pages. Experimental results verify the effectives of the proposed method
Milos Kudelka, Václav Snásel, Ondrej Lehecka, Eyas El-Qawasmeh
Web Intelligence2
2006 A new range query algorithm for Universal B-trees
Tomás Skopal, Michal Krátký, Jaroslav Pokorný, Václav Snásel
Inf. Syst.4
2005 Nearest Neighbours Search Using the PM-Tree
Tomás Skopal, Jaroslav Pokorný, Václav Snásel
DASFAA3
2005 Efficient Searching in Large Inheritance Hierarchies
Michal Krátký, Svatopluk Stolfa, Václav Snásel, Ivo Vondrák
DEXA3
2005 Information Extraction from HTML Product Catalogues: From Source Code and Images to RDF
abstract
We describe an application of information extraction from company Web sites focusing on product offers. A statistical approach to text analysis is used in conjunction with different ways of image classification. Ontological knowledge is used to group the extracted items into structured objects. The results are stored in an RDF repository and made available for structured search.
Martin Labský, Vojtech Svátek, Ondrej Sváb-Zamazal, Pavel Praks, Michal Krátký, Václav Snásel
Web Intelligence6
2004 Metric Indexing for the Vector Model in Text Retrieval
Tomás Skopal, Pavel Moravec 0001, Jaroslav Pokorný, Václav Snásel
SPIRE4
2003 Revisiting M-Tree Building Principles
Tomás Skopal, Jaroslav Pokorný, Michal Krátký, Václav Snásel
ADBIS4
1999 Word-Based Compression Methods and Indexing for Text Retrieval Systems
Jiri Dvorský, Jaroslav Pokorný, Václav Snásel
ADBIS3
1999 Word-based Compression Methods for Large Text Documents
abstract
Summary form only given. We present a new compression method, called WLZW, which is a word-based modification of classic LZW. The algorithm is two-phase, it uses only one table for words and non-words (so called tokens), and a single data structure for the lexicon is usable as a text index. The length of words and non-words is restricted. This feature improves the compress ratio achieved. Tokens of unlimited length alternate, when they are read from the input stream. Because of restricted length of tokens alternating of tokens is corrupted, because some tokens are divided into several parts of same type. To save alternating of tokens two special tokens are created. They are empty word and empty non-word. They contain no character. Empty word is inserted between two non-words and empty non-word between two words. Alternating of tokens is saved for all sequences of tokens. The alternating of tokens is an important piece of information. With this knowledge the kind of the next token can be predicted. One selected (so-called victim) non-word can be deleted from input stream. An algorithm to search the victim is also presented. In the decompression phase, a deleted victim is recognized as an error in alternating of words and non-words in sequence. The algorithm was tested on many texts in different formats (ASCII, RTF). The Canterbury corpus, a large set, was used as a standard for publication results. The compression ratio achieved is fairly good, on average 25%-22%. Decompression is very fast. Moreover, the algorithm enables evaluation of database queries in given text. This supports the idea of leaving data in the compressed state as long as possible, and to decompress it when it is necessary.
Jiri Dvorský, Jaroslav Pokorný, Václav Snásel
Data Compression Conference3