Filip Zelezný

dblp:00/2873 · DBLP profile ↗
← Back
54ranked-venue papers
9as first author
3since 2021 · last 2021
0000-0001-9780-3376ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 8 first-author · 3 since 2021Theory of computation · 17 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-authorDatabases, data management, data science and information retrieval · 7Human-computer interaction and ubiquitous computing · 3Security and privacy · 1
YearPublicationVenuePosition
2021 Lossless Compression of Structured Convolutional Models via Lifting
Gustav Sír, Filip Zelezný, Ondrej Kuzelka
ICLR2
2021 Automatic Conjecturing of P-Recursions Using Lifted Inference
Jáchym Barvínek, Timothy van Bremen, Yuyi Wang 0001, Filip Zelezný, Ondrej Kuzelka
ILP4
2021 Beyond graph neural networks with lifted relational neural networks
Gustav Sír, Filip Zelezný, Ondrej Kuzelka
Mach. Learn.2
2020 Guest editors' introduction: special issue on Inductive Logic Programming (ILP 2019)
Dimitar Kazakov, Filip Zelezný
Mach. Learn.2
2019 Estimating sequence similarity from read sets for clustering next-generation sequencing data
Petr Rysavý, Filip Zelezný
Data Min. Knowl. Discov.2
2019 Learning to predict soccer results from relational data with gradient boosted trees
Ondrej Hubácek, Gustav Sír, Filip Zelezný
Mach. Learn.3
2019 Efficient Extraction of Network Event Types from NetFlows
abstract
To perform sophisticated traffic analysis, such as intrusion detection, network monitoring tools firstly need to extract higher-level information from lower-level data by reconstructing events and activities from as primitive information as individual network packets or traffic flows. Aggregating communication data into meaningful entities is an open problem and existing, typically clustering-based, solutions are often highly suboptimal, producing results that may misinterpret the extracted information and consequently miss many network events. We propose a novel method for the extraction of various predefined types of network events from raw network flow data. The new method is based on analysis of computational properties of the event types as prescribed by their attributes in a given descriptive language. The corresponding events are then extracted with a supreme recall as compared to a respective event extraction part of an in-production intrusion detection system Camnep.
Gustav Sír, Filip Zelezný
Secur. Commun. Networks2
2018 Lifted Relational Neural Networks: Efficient Learning of Latent Relational Structures
abstract
We propose a method to combine the interpretability and expressive power of firstorder logic with the effectiveness of neural network learning. In particular, we introduce a lifted framework in which first-order rules are used to describe the structure of a given problem setting. These rules are then used as a template for constructing a number of neural networks, one for each training and testing example. As the different networks corresponding to different examples share their weights, these weights can be efficiently learned using stochastic gradient descent. Our framework provides a flexible way for implementing and combining a wide variety of modelling constructs. In particular, the use of first-order logic allows for a declarative specification of latent relational structures, which can then be efficiently discovered in a given data set using neural network learning. Experiments on 78 relational learning benchmarks clearly demonstrate the effectiveness of the framework.
Gustav Sír, Vojtech Aschenbrenner, Filip Zelezný, Steven Schockaert, Ondrej Kuzelka
J. Artif. Intell. Res.3
2017 Estimating Sequence Similarity from Contig Sets
Petr Rysavý, Filip Zelezný
IDA2
2017 Stacked Structure Learning for Lifted Relational Neural Networks
Gustav Sír, Martin Svatos, Filip Zelezný, Steven Schockaert, Ondrej Kuzelka
ILP3
2017 Pruning Hypothesis Spaces Using Learned Domain Theories
Martin Svatos, Gustav Sír, Filip Zelezný, Steven Schockaert, Ondrej Kuzelka
ILP3
2016 Estimating Sequence Similarity from Read Sets for Clustering Sequencing Data
Petr Rysavý, Filip Zelezný
IDA2
2016 Learning Predictive Categories Using Lifted Relational Neural Networks
Gustav Sír, Suresh Manandhar, Filip Zelezný, Steven Schockaert, Ondrej Kuzelka
ILP3
2015 Novel gene sets improve set-level classification of prokaryotic gene expression data
abstract
BACKGROUND: Set-level classification of gene expression data has received significant attention recently. In this setting, high-dimensional vectors of features corresponding to genes are converted into lower-dimensional vectors of features corresponding to biologically interpretable gene sets. The dimensionality reduction brings the promise of a decreased risk of overfitting, potentially resulting in improved accuracy of the learned classifiers. However, recent empirical research has not confirmed this expectation. Here we hypothesize that the reported unfavorable classification results in the set-level framework were due to the adoption of unsuitable gene sets defined typically on the basis of the Gene ontology and the KEGG database of metabolic networks. We explore an alternative approach to defining gene sets, based on regulatory interactions, which we expect to collect genes with more correlated expression. We hypothesize that such more correlated gene sets will enable to learn more accurate classifiers. METHODS: We define two families of gene sets using information on regulatory interactions, and evaluate them on phenotype-classification tasks using public prokaryotic gene expression data sets. From each of the two gene-set families, we first select the best-performing subtype. The two selected subtypes are then evaluated on independent (testing) data sets against state-of-the-art gene sets and against the conventional gene-level approach. RESULTS: The novel gene sets are indeed more correlated than the conventional ones, and lead to significantly more accurate classifiers. The novel gene sets are indeed more correlated than the conventional ones, and lead to significantly more accurate classifiers. CONCLUSION: Novel gene sets defined on the basis of regulatory interactions improve set-level classification of gene expression data. The experimental scripts and other material needed to reproduce the experiments are available at http://ida.felk.cvut.cz/novelgenesets.tar.gz.
Matej Holec, Ondrej Kuzelka, Filip Zelezný
BMC Bioinform.3
2014 A method for reduction of examples in relational learning
Ondrej Kuzelka, Andrea Szabóová, Filip Zelezný
J. Intell. Inf. Syst.3
2014 Guest editors introduction: special issue on Inductive Logic Programming (ILP 2012)
Fabrizio Riguzzi, Filip Zelezný
Mach. Learn.2
2013 Guest editor's introduction: special issue of the ECML PKDD 2013 journal track
Hendrik Blockeel, Kristian Kersting, Siegfried Nijssen, Filip Zelezný
Data Min. Knowl. Discov.4
2013 Guest editor's introduction: special issue of the ECML PKDD 2013 journal track
Hendrik Blockeel, Kristian Kersting, Siegfried Nijssen, Filip Zelezný
Mach. Learn.4
2012 Extending the ball-histogram method with continuous distributions and an application to prediction of DNA-binding proteins
abstract
We introduce a novel method for prediction of DNA-binding propensity of proteins which extends our recently introduced ball-histogram method (Szabóova et al. 2012). Unlike the original ball-histogram method, it allows handling of continuous properties of protein regions. In experiments on four datasets of proteins, we show that the method improves upon the original ball-histogram method as well as other existing methods in terms of predictive accuracy.
Ondrej Kuzelka, Andrea Szabóová, Filip Zelezný
BIBM3
2012 Relational Learning with Polynomials
abstract
We describe a conceptually simple framework for transformation-based learning in hybrid relational domains. The proposed approach is related to hybrid Markov logic and to Gaussian logic framework. We evaluate the approach in three domains and show that it can achieve state-of-the-art performance while using only limited amount of information.
Ondrej Kuzelka, Andrea Szabóová, Filip Zelezný
ICTAI3
2012 Bounded Least General Generalization
Ondrej Kuzelka, Andrea Szabóová, Filip Zelezný
ILP3
2012 Comparative evaluation of set-level techniques in predictive classification of gene expression samples
abstract
BACKGROUND: Analysis of gene expression data in terms of a priori-defined gene sets has recently received significant attention as this approach typically yields more compact and interpretable results than those produced by traditional methods that rely on individual genes. The set-level strategy can also be adopted with similar benefits in predictive classification tasks accomplished with machine learning algorithms. Initial studies into the predictive performance of set-level classifiers have yielded rather controversial results. The goal of this study is to provide a more conclusive evaluation by testing various components of the set-level framework within a large collection of machine learning experiments. RESULTS: Genuine curated gene sets constitute better features for classification than sets assembled without biological relevance. For identifying the best gene sets for classification, the Global test outperforms the gene-set methods GSEA and SAM-GS as well as two generic feature selection methods. To aggregate expressions of genes into a feature value, the singular value decomposition (SVD) method as well as the SetSig technique improve on simple arithmetic averaging. Set-level classifiers learned with 10 features constituted by the Global test slightly outperform baseline gene-level classifiers learned with all original data features although they are slightly less accurate than gene-level classifiers learned with a prior feature-selection step. CONCLUSION: Set-level classifiers do not boost predictive accuracy, however, they do achieve competitive accuracy if learned with the right combination of ingredients. AVAILABILITY: Open-source, publicly available software was used for classifier learning and testing. The gene expression datasets and the gene set database used are also publicly available. The full tabulation of experimental results is available at http://ida.felk.cvut.cz/CESLT.
Matej Holec, Jirí Kléma, Filip Zelezný, Jakub Tolar
BMC Bioinform.3
2012 Prediction of DNA-binding propensity of proteins by the ball-histogram method using automatic template search
abstract
We contribute a novel, ball-histogram approach to DNA-binding propensity prediction of proteins. Unlike state-of-the-art methods based on constructing an ad-hoc set of features describing physicochemical properties of the proteins, the ball-histogram technique enables a systematic, Monte-Carlo exploration of the spatial distribution of amino acids complying with automatically selected properties. This exploration yields a model for the prediction of DNA binding propensity. We validate our method in prediction experiments, improving on state-of-the-art accuracies. Moreover, our method also provides interpretable features involving spatial distributions of selected amino acids.
Andrea Szabóová, Ondrej Kuzelka, Filip Zelezný, Jakub Tolar
BMC Bioinform.3
2011 Template-based semi-automatic workflow construction for gene expression data analysis
abstract
We propose a technique for semi-automatic construction of gene expression data analysis workflows by grammar-like inference based on predefined workflow templates. The templates represent routinely used sequences of procedures such as normalization, data transformation, classifier learning, etc. Variations of such workflows (such as different instantiations to specific algorithms) may entail significant variance in the quality of the analysis results and our formalism enables to automatically explore such variations. Adhering to proven templates helps preserve the sanity of explored workflows and prevents the combinatorial explosion encountered by fully automatic workflow planners. Here we propose the basic principles of template-based workflow construction and demonstrate their working in the publicly available tool XGENE.ORG for multi-platform gene expression analysis.
Jiri Belohradsky, David A. Monge, Filip Zelezný, Matej Holec, Carlos García Garino
CBMS3
2011 Subgroup Discovery Using Bump Hunting on Multi-relational Histograms
Radomír Cernoch, Filip Zelezný
ILP2
2011 Comparative Evaluation of Set-Level Techniques in Microarray Classification
Jirí Kléma, Matej Holec, Filip Zelezný, Jakub Tolar
ISBRA3
2011 Prediction of DNA-Binding Propensity of Proteins by the Ball-Histogram Method
Andrea Szabóová, Ondrej Kuzelka, Sergio Morales E., Filip Zelezný, Jakub Tolar
ISBRA4
2011 Gaussian Logic for Predictive Classification
Ondrej Kuzelka, Andrea Szabóová, Matej Holec, Filip Zelezný
ECML/PKDD (2)4
2011 Block-wise construction of tree-like relational features with monotone reducibility and redundancy
Ondrej Kuzelka, Filip Zelezný
Mach. Learn.2
2011 An experimental test of Occam's razor in classification
Jan Zahálka, Filip Zelezný
Mach. Learn.2
2011 Automating Knowledge Discovery Workflow Composition Through Ontology-Based Planning
abstract
The problem addressed in this paper is the challenge of automated construction of knowledge discovery workflows, given the types of inputs and the required outputs of the knowledge discovery process. Our methodology consists of two main ingredients. The first one is defining a formal conceptualization of knowledge types and data mining algorithms by means of knowledge discovery ontology. The second one is workflow composition formalized as a planning task using the ontology of domain and task descriptions. Two versions of a forward chaining planning algorithm were developed. The baseline version demonstrates suitability of the knowledge discovery ontology for planning and uses Planning Domain Definition Language (PDDL) descriptions of algorithms; to this end, a procedure for converting data mining algorithm descriptions into PDDL was developed. The second directly queries the ontology using a reasoner. The proposed approach was tested in two use cases, one from scientific discovery in genomics and another from advanced engineering. The results show the feasibility of automated workflow construction achieved by tight integration of planning and ontological reasoning.
Monika Záková, Petr Kremen, Filip Zelezný, Nada Lavrac
IEEE Trans Autom. Sci. Eng.3
2010 Speeding Up Planning through Minimal Generalizations of Partially Ordered Plans
Radomír Cernoch, Filip Zelezný
ILP2
2010 Seeing the World through Homomorphism: An Experimental Study on Reducibility of Examples
Ondrej Kuzelka, Filip Zelezný
ILP2
2010 Taming the Complexity of Inductive Logic Programming
Filip Zelezný, Ondrej Kuzelka
SOFSEM1
2009 Block-wise construction of acyclic relational features with monotone irreducibility and relevancy properties
abstract
We describe an algorithm for constructing a set of acyclic conjunctive relational features by combining smaller conjunctive blocks. Unlike traditional level-wise approaches which preserve the monotonicity of frequency, our block-wise approach preserves a form of monotonicity of the irreducibility and relevancy feature properties, which are important in propositionalization employed in the context of classification learning. With pruning based on these properties, our block-wise approach efficiently scales to features including tens of first-order literals, far beyond the reach of state-of-the art propositionalization or inductive logic programming systems.
Ondrej Kuzelka, Filip Zelezný
ICML2
2009 Integrating Multiple-Platform Expression Data through Gene Set Features
Matej Holec, Filip Zelezný, Jirí Kléma, Jakub Tolar
ISBRA2
2009 Guest editors' introduction: Special issue on Inductive Logic Programming (ILP-2008)
Filip Zelezný, Nada Lavrac
Mach. Learn.1
2008 Fast estimation of first-order clause coverage through randomization and maximum likelihood
abstract
In inductive logic programming, θ-subsumption is a widely used coverage test. Unfortunately, testing θ-subsumption is NP-complete, which represents a crucial efficiency bottleneck for many relational learners. In this paper, we present a probabilistic estimator of clause coverage, based on a randomized restarted search strategy. Under a distribution assumption, our algorithm can estimate clause coverage without having to decide subsumption for all examples. We implement this algorithm in program ReCovEr. On generated graph data and real-world datasets, we show that ReCovEr provides reasonably accurate estimates while achieving dramatic runtimes improvements compared to a state-of-the-art algorithm.
Ondrej Kuzelka, Filip Zelezný
ICML2
2008 A Restarted Strategy for Efficient Subsumption Testing
Ondrej Kuzelka, Filip Zelezný
Fundam. Informaticae2
2008 Sequential Data Mining: A Comparative Case Study in Development of Atherosclerosis Risk Factors
abstract
Sequential data represent an important source of potentially new medical knowledge. However, this type of data is rarely provided in a format suitable for immediate application of conventional mining algorithms. This paper summarizes and compares three different sequential mining approaches based, respectively, on windowing, episode rules, and inductive logic programming. Windowing is one of the essential methods of data preprocessing. Episode rules represent general sequential mining, while inductive logic programming extracts first-order features whose structure is determined by background knowledge. The three approaches are demonstrated and evaluated in terms of a case study STULONG. It is a longitudinal preventive study of atherosclerosis where the data consist of a series of long-term observations recording the development of risk factors and associated conditions. The intention is to identify frequent sequential/temporal patterns. Possible relations between the patterns and an onset of any of the observed cardiovascular diseases are also studied.
Jirí Kléma, Lenka Nováková, Filip Karel, Olga Stepánková, Filip Zelezný
IEEE Trans. Syst. Man Cybern. Part C5
2008 Learning Relational Descriptions of Differentially Expressed Gene Groups
abstract
This paper presents a method that uses gene ontologies (GOs), together with the paradigm of relational subgroup discovery, to find compactly described groups of genes differentially expressed in specific cancers. The groups are described by means of relational logic features, extracted from publicly available GO information, and are straightforwardly interpretable by medical experts. We applied the proposed method to three gene expression data sets with the following respective sets of sample classes: 1) acute lymphoblastic leukemia (ALL) versus acute myeloid leukemia (AML); 2) seven subtypes of ALL; and 3) 14 different types of cancers. Significant number of discovered groups of genes had a description that highlighted the underlying biological process responsible for distinguishing one class from the other classes. The quality of the discovered descriptions was also verified by cross validation. We believe that the presented approach will significantly contribute to the application of relational machine learning to gene expression analysis, given the expected increase in both the quality and quantity of gene/protein annotations in the near future.
Igor Trajkovski, Filip Zelezný, Nada Lavrac, Jakub Tolar
IEEE Trans. Syst. Man Cybern. Part C2
2007 Exploiting Term, Predicate, and Feature Taxonomies in Propositionalization and Propositional Rule Learning
Monika Záková, Filip Zelezný
ECML2
2006 ILP Through Propositionalization and Stochastic k-Term DNF Learning
Aline Paes, Filip Zelezný, Gerson Zaverucha, David Page, Ashwin Srinivasan 0001
ILP2
2006 Relational Data Mining Applied to Virtual Engineering of Product Designs
Monika Záková, Filip Zelezný, Javier A. García-Sedano, Cyril Masia Tissot, Nada Lavrac, Petr Kremen, Javier Molina
ILP2
2006 Propositionalization-based relational subgroup discovery with RSD
Filip Zelezný, Nada Lavrac
Mach. Learn.1
2006 Randomised restarted search in ILP
Filip Zelezný, Ashwin Srinivasan 0001, David Page
Mach. Learn.1
2005 Efficient construction of relational features
abstract
Devising algorithms for learning from multi-relational data is currently considered an important challenge. The wealth of traditional single-relational machine learning tools, on the other hand, calls for methods of 'propositionalization', i.e. conversion of multi-relational data into single-relational representations. A major stream of propositionalization algorithms is based on the construction of truth-valued features (first-order logic atom conjunctions), which capture relational properties of data and play the role of binary attributes in the resulting single-table representation. Such algorithms typically use backtrack depth first search for the syntactic construction of features complying to user's mode/type declarations. As such they incur a complexity factor exponential in the maximum allowed feature size. Here we present a polynomial-runtime alternative based on an efficient reduction between the feature construction problems on the propositional satisfiability (SAT) problem, such that the latter involves only Horn clauses and is therefore efficiently solvable.
Filip Zelezný
ICMLA1
2005 Efficient Sampling in Relational Feature Spaces
Filip Zelezný
ILP1
2004 A Monte Carlo Study of Randomised Restarted Search in ILP
Filip Zelezný, Ashwin Srinivasan 0001, David Page
ILP1
2004 Induction of comprehensible models for gene expression datasets by subgroup discovery methodology
Dragan Gamberger, Nada Lavrac, Filip Zelezný, Jakub Tolar
J. Biomed. Informatics3
2003 Comparative Evaluation of Approaches to Propositionalization
Mark-A. Krogel, Simon Alan Rawles, Filip Zelezný, Peter A. Flach, Nada Lavrac, Stefan Wrobel
ILP3
2002 RSD: Relational Subgroup Discovery through First-Order Feature Construction
Nada Lavrac, Filip Zelezný, Peter A. Flach
ILP2
2002 Lattice-Search Runtime Distributions May Be Heavy-Tailed
Filip Zelezný, Ashwin Srinivasan 0001, David Page
ILP1
2001 Learning Functions from Imperfect Positive Data
Filip Zelezný
ILP1