Tomás Kliegr

dblp:46/3161 · DBLP profile ↗
← Back
24ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-7261-0380ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 12 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Meaningless is better: Hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
abstract
This paper introduces a novel method, referred to as "hashing", which involves masking potentially bias-inducing words in large language models (LLMs) with hash-like meaningless identifiers to reduce cognitive biases and reliance on external knowledge. The method was tested across four sets of experiments involving a total of 640 prompts. Statistical analysis using chi-square tests showed significant improvements in all tested scenarios, which covered the Llama, ChatGPT, Copilot, Gemini and Mixtral models. In the first experiment, hashing decreased the conjunction fallacy rate in a modified version of the “Linda” problem aimed at evaluating susceptibility to cognitive biases. In the second experiment, it improved LLM results on the frequent itemset extraction task. In the third experiment, we found hashing is also effective when the Linda problem is presented in a tabular format rather than text, indicating that the technique works across various input representations. In the fourth experiment, consistent with psychological literature showing that step-by-step reasoning suppresses biases, we compared hashing's effectiveness with Chain of Thought (CoT) LLM models, finding that CoT also suppressed biases in LLMs. Overall, the proposed hashing method was shown to improve bias reduction and reduce the unwanted incorporation of external knowledge. Despite bias reduction, reduction in hallucination rates was inconsistent across LLM model types. These findings suggest that masking bias-inducing terms can improve LLM performance, although its effectiveness is model- and task-dependent.
Milena Chadimová, Eduard Jurásek, Tomás Kliegr
Inf. Process. Manag.3
2025 Traceable LLM-based validation of statements in knowledge graphs
abstract
This article presents a method for verifying RDF triples using LLMs , with an emphasis on providing traceable arguments. Because the LLMs cannot currently reliably identify the origin of the information used to construct the response to the user prompt, our approach is to avoid using internal LLM factual knowledge altogether. Instead, verified RDF statements are compared to chunks of external documents retrieved through a web search or Wikipedia. To assess the possible application of this retrieval augmented generation (RAG) workflow on biosciences content, we evaluated 1,719 positive statements from the BioRED dataset and the same number of newly generated negative statements. The resulting precision is 88 %, and recall is 44 %. This indicates that the method requires human oversight. We also evaluated the method on the SNLI dataset which allowed us to compare our approach with models specifically tuned for the natural language inference task. We demonstrate the method on Wikidata, where a SPARQL query is used to automatically retrieve statements needing verification. Overall, the results suggest that LLMs could be used for large-scale verification of statements in KGs, a task previously unfeasible due to human annotation costs.
Daniel Adam, Tomás Kliegr
Inf. Process. Manag.2
2025 LLM-based feature generation from text for interpretable machine learning
Vojtech Balek, Lukás Sýkora, Vilém Sklenák, Tomás Kliegr
Mach. Learn.4
2025 Explaining word embeddings with perfect fidelity: a case study in predicting research impact
Lucie Dvorackova, Marcin P. Joachimiak, Michal Cerny, Adriana Kubecova, Vilém Sklenák, Tomás Kliegr
Mach. Learn.6
2024 Explainable and interpretable machine learning and data mining
abstract
Abstract The growing number of applications of machine learning and data mining in many domains—from agriculture to business, education, industrial manufacturing, and medicine—gave rise to new requirements for how to inspect and control the learned models. The research domain of explainable artificial intelligence (XAI) has been newly established with a strong focus on methods being applied post-hoc on black-box models. As an alternative, the use of interpretable machine learning methods has been considered—where the learned models are white-box ones. Black-box models can be characterized as representing implicit knowledge—typically resulting from statistical and neural approaches of machine learning, while white-box models are explicit representations of knowledge—typically resulting from rule-learning approaches. In this introduction to the special issue on ‘Explainable and Interpretable Machine Learning and Data Mining’ we propose to bring together both perspectives, pointing out commonalities and discussing possibilities to integrate them.
Martin Atzmüller, Johannes Fürnkranz, Tomás Kliegr, Ute Schmid
Data Min. Knowl. Discov.3
2023 Apriori Modified for Action Rules Mining
abstract
Action rule mining is an extension of classification rule learning, in which an action rule provides information similar to that of a standard classification rule and also suggests a course of action. Action rules are used to obtain counterfactual explanations. Rule-based action rule mining involves two separate steps: mining classification rules, typically using the Apriori algorithm, and a user-set minimum support and confidence threshold. The second step involves forming action rules from the classification rules by utilizing a second set of user-set parameters including the desired and undesired values of the target attribute, a list of stable attributes, and a list of flexible attributes. This two-stage approach is inefficient as the first step tends to generate excessively many classification rules, most of which can never form an action rule meeting the second set of parameters. This paper describes a modified Apriori algorithm for action rule mining (Action-Apriori), which includes an enhanced version of downward closure that further reduces the space of possible candidates by checking whether the candidate itemsets also meet the second set of user-set parameters. The benchmarks show a consistent reduction in the learning time of the new algorithm compared to the state-of-the-art ARAS algorithm. This paper is supplemented by an open-source implementation.
Lukás Sýkora, Tomás Kliegr
K-CAP2
2023 QCBA: improving rule classifiers learned from quantitative data by recovering information lost by discretisation
abstract
Abstract A prediscretisation of numerical attributes which is required by some rule learning algorithms is a source of inefficiencies. This paper describes new rule tuning steps that aim to recover lost information in the discretisation and new pruning techniques that may further reduce the size of rule models and improve their accuracy. The proposed QCBA method was initially developed to postprocess quantitative attributes in models generated by Classification based on associations (CBA) algorithm, but it can also be applied to the results of other rule learning approaches. We demonstrate the effectiveness on the postprocessing of models generated by five association rule classification algorithms (CBA, CMAR, CPAR, IDS, SBRL) and two first-order logic rule learners (FOIL2 and PRM). Benchmarks on 22 datasets from the UCI repository show smaller size and the overall best predictive performance for FOIL2+QCBA compared to all seven baselines. Postoptimised CBA models have a better predictive performance compared to the state-of-the-art rule learner CORELS in this benchmark. The article contains an ablation study for the individual postprocessing steps and a scalability analysis on the KDD’99 Anomaly detection dataset.
Tomás Kliegr, Ebroul Izquierdo
Appl. Intell.1
2023 Introduction to the Special Issue on Logic Rules and Reasoning: Selected Papers from the 4th International Joint Conference on Rules and Reasoning (RuleML+RR 2020)
Tomás Kliegr, Víctor Gutiérrez-Basulto, Ahmet Soylu
Theory Pract. Log. Program.1
2021 A review of possible effects of cognitive biases on interpretation of rule-based machine learning models
abstract
While the interpretability of machine learning models is often equated with their mere syntactic comprehensibility, we think that interpretability goes beyond that, and that human interpretability should also be investigated from the point of view of cognitive science. The goal of this paper is to discuss to what extent cognitive biases may affect human understanding of interpretable machine learning models, in particular of logical rules discovered from data. Twenty cognitive biases are covered, as are possible debiasing techniques that can be adopted by designers of machine learning algorithms and software. Our review transfers results obtained in cognitive psychology to the domain of machine learning, aiming to bridge the current gap between these two areas. It needs to be followed by empirical studies specifically focused on the machine learning domain.
Tomás Kliegr, Stepán Bahník, Johannes Fürnkranz
Artif. Intell.1
2020 On cognitive preferences and the plausibility of rule-based models
abstract
Abstract It is conventional wisdom in machine learning and data mining that logical models such as rule sets are more interpretable than other models, and that among such rule-based models, simpler models are more interpretable than more complex ones. In this position paper, we question this latter assumption by focusing on one particular aspect of interpretability, namely the plausibility of models. Roughly speaking, we equate the plausibility of a model with the likeliness that a user accepts it as an explanation for a prediction. In particular, we argue that—all other things being equal—longer explanations may be more convincing than shorter ones, and that the predominant bias for shorter models, which is typically necessary for learning powerful discriminative models, may not be suitable when it comes to user acceptance of the learned models. To that end, we first recapitulate evidence for and against this postulate, and then report the results of an evaluation in a crowdsourcing study based on about 3000 judgments. The results do not reveal a strong preference for simple rules, whereas we can observe a weak preference for longer rules in some domains. We then relate these results to well-known cognitive biases such as the conjunction fallacy, the representative heuristic, or the recognition heuristic, and investigate their relation to rule length and plausibility.
Johannes Fürnkranz, Tomás Kliegr, Heiko Paulheim
Mach. Learn.2
2018 The Need for Interpretability Biases
Johannes Fürnkranz, Tomás Kliegr
IDA2
2018 Antonyms are similar: Towards paradigmatic association approach to rating similarity in SimLex-999 and WordSim-353
Tomás Kliegr, Ondrej Sváb-Zamazal
Data Knowl. Eng.1
2018 EasyMiner.eu: Web framework for interpretable machine learning based on rules and frequent itemsets
Stanislav Vojír, Vaclav Zeman, Jaroslav Kuchar, Tomás Kliegr
Knowl. Based Syst.4
2017 InBeat: JavaScript recommender system supporting sensor input and linked data
abstract
Interest Beat (inbeat.eu) is an open source recommender framework that fulfills some of the demands raised by emerging applications that infer ratings from sensor input or use linked open data cloud for feature expansion. As a recommender algorithm, InBeat uses association rules, which allow to explain why a specific recommendation was made. Due to modular architecture, other algorithms can be easily plugged in. InBeat has a pure JavaScript version, which allows to confine processing to a client-side device. There is a performance optimized server-side bundle, which succesfully participated in two recent recommender competitions involving large volumes of streaming data. InBeat works on a number of platforms and is also available for Docker.
Jaroslav Kuchar, Tomás Kliegr
Knowl. Based Syst.2
2016 Crowdsourced Corpus with Entity Salience Annotations
Milan Dojchinovski, Dinesh Reddy, Tomás Kliegr, Tomas Vitvar, Harald Sack
LREC3
2016 LHD 2.0: A text mining approach to typing entities in knowledge graphs
Tomás Kliegr, Ondrej Sváb-Zamazal
J. Web Semant.1
2015 Linked hypernyms: Enriching DBpedia with Targeted Hypernym Discovery
abstract
The Linked Hypernyms Dataset (LHD) provides entities described by Dutch, English and German Wikipedia articles with types in the DBpedia namespace. The types are extracted from the first sentences of Wikipedia articles using Hearst pattern matching over part-of-speech annotated text and disambiguated to DBpedia concepts. The dataset covers 1.3 million RDF type triples from English Wikipedia, out of which 1 million RDF type triples were found not to overlap with DBpedia, and 0.4 million with YAGO2s. There are about 770 thousand German and 650 thousand Dutch Wikipedia entities assigned a novel type, which exceeds the number of entities in the localized DBpedia for the respective language. RDF type triples from the German dataset have been incorporated to the German DBpedia. Quality assessment was performed altogether based on 16.500 human ratings and annotations. For the English dataset, the average accuracy is 0.86, for German 0.77 and for Dutch 0.88. The accuracy of raw plain text hypernyms exceeds 0.90 for all languages. The LHD release described and evaluated in this article targets DBpedia 3.8, LHD version for the DBpedia 3.9 containing approximately 4.5 million RDF type triples is also available.
Tomás Kliegr
J. Web Semant.1
2014 Orwellian Eye: Video Recommendation with Microsoft Kinect
abstract
This paper demonstrates Interest Beat (InBeat.eu) as a recommender system for online videos, which determines user interest in the content based on gaze tracking with Microsoft Kinect in addition to explicit user feedback. Content of the videos is represented using a semantic wikifier. User profile is constructed from preference rules, which are discovered with an association rule learner.
Tomás Kliegr, Jaroslav Kuchar
ECAI1
2014 Towards Linked Hypernyms Dataset 2.0: complementing DBpedia with hypernym discovery
Tomás Kliegr, Ondrej Sváb-Zamazal
LREC1
2013 Entityclassifier.eu: Real-Time Classification of Entities in Text with Wikipedia
Milan Dojchinovski, Tomás Kliegr
ECML/PKDD (3)2
2013 GAIN: web service for user tracking and preference learning - a smart TV use case
abstract
GAIN (inbeat.eu) is a web application and service for capturing and preprocessing user interactions with semantically described content. GAIN outputs a set of instances in tabular form suitable for further processing with generic machine-learning algorithms. GAIN is demoed as a component of a "SMART-TV" recommender system. Content is automatically described with DBpedia types using a Named Entity Recognition (NER) system. Interest is determined based on explicit user actions and user's attention computed by 3D head pose estimation. Preference rules are learnt with an association rule mining algorithm. These can be e.g. deployed to a business rules system, acting as a recommender.
Jaroslav Kuchar, Tomás Kliegr
RecSys2
2012 Association Rule Mining Following the Web Search Paradigm
Radek Skrabal, Milan Simunek, Stanislav Vojír, Andrej Hazucha, Tomás Marek, David Chudán, Tomás Kliegr
ECML/PKDD (2)7
2011 SEWEBAR-CMS: semantic analytical report authoring for data mining results
Tomás Kliegr, Vojtech Svátek, Martin Ralbovský, Milan Simunek
J. Intell. Inf. Syst.1
2009 Semantic Analytical Reports: A Framework for Post-processing Data Mining Results
Tomás Kliegr, Martin Ralbovský, Vojtech Svátek, Milan Simunek, Vojtech Jirkovský, Jan Nemrava, Jan Zemánek
ISMIS1