Jakub Klímek

dblp:35/7915 · DBLP profile ↗
← Back
18ranked-venue papers in the field
6as first author
6since 2021 · last 2025
0000-0001-7234-3051ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 9 (2 first)Database Systems & Data Management · 5 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (2 first)Data Mining & Knowledge Discovery · 1Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2025 Data Specification Vocabulary (DSV): Representation of Application Profiles of Semantic Data Specifications
Jakub Klímek, Stepán Stenchlák, Petr Skoda 0001
iiWAS1
2024 Enhancing Domain Modeling with Pre-trained Large Language Models: An Automated Assistant for Domain Modelers
Dominik Prokop, Stepán Stenchlák, Petr Skoda 0001, Jakub Klímek, Martin Necaský
ER4
2023 Semantic, Technical and Legal Interoperability of European Company Open Data in Practice: The STIRData Approach
Jakub Klímek, Alexandros Chortaras, Jakub Mísek, Jim J. Yang, Steinar Skagemo, Vassilis Tzouvaras
DATA1
2022 Open dataset discovery using context-enhanced similarity search
David Bernhauer, Martin Necaský, Petr Skoda 0001, Jakub Klímek, Tomás Skopal
Knowl. Inf. Syst.4
2021 Simplod: Simple SPARQL Query Builder for Rapid Export of Linked Open Data in the Form of CSV Files
abstract
In the last decade, linked open data (LOD) became the de-facto highest standard of publishing data on the Web, a.k.a. 5-star open data. One of the advantages of LOD is better data discovery thanks to the linkable nature of the data. Unfortunately, in the wider IT expert community, RDF and SPARQL is considered unnecessarily complex, hard to understand and hard to process, especially when transferring open data in bulk. However, it is this wider IT expert community which is the one supposed to both publish data produced in the public administration as open data, and to build consumer-facing applications on top of the data. In this paper, we propose an approach to making LOD easily accessible to the community used to CSV files. We propose a tool called Simplod focused on rapid, straightforward and customizable formulation of a SPARQL SELECT query intended for customizable bulk transformation of LOD published in SPARQL endpoints into CSV files. In addition, Simplod configurations can be stored in and shared via Solid pods.
Antonín Jares, Jakub Klímek
iiWAS2
2021 Similarity vs. Relevance: From Simple Searches to Complex Discovery
Tomás Skopal, David Bernhauer, Petr Skoda 0001, Jakub Klímek, Martin Necaský
SISAP4
2020 Evaluation Framework for Search Methods Focused on Dataset Findability in Open Data Catalogs
abstract
Many institutions publish datasets as Open Data in catalogs, however, their retrieval remains problematic issue due to the absence of dataset search benchmarking. We propose a framework for evaluating findability of datasets, regardless of retrieval models used. As task-agnostic labeling of datasets by ground truth turns out to be infeasible in the general domain of open data datasets, the proposed framework is based on evaluation of entire retrieval scenarios that mimic complex retrieval tasks. In addition to the framework we present a proof of concept specification and evaluation on several similarity-based retrieval models and several dataset discovery scenarios within a catalog, using our experimental evaluation tool. Instead of traditional matching of query with metadata of all the datasets, in similarity-based retrieval the query is formulated using a set of datasets (query by example) and the most similar datasets to the query set are retrieved from the catalog as a result.
Petr Skoda 0001, David Bernhauer, Martin Necaský, Jakub Klímek, Tomás Skopal
iiWAS4
2019 Improving Findability of Open Data Beyond Data Catalogs
abstract
There is a vast amount of datasets available as Open Data on the Web. However, it is challenging for consumers to find datasets relevant to their goals. This is because the available metadata in catalogs is not descriptive enough. Nevertheless, datasets exist in various types of contexts not expressed in the metadata. These may include information about the data publisher, the legislation related to dataset publication, etc. In this paper we describe an idea of a data model that enables consumers to better understand the data. We propose to define a formal model for representation of the datasets and their contexts, and we propose to apply existing similarity techniques, adjust them to fit each identified dataset context type and combine them together to measure similarity of datasets in new ways, improving their findability.
Tomás Skopal, Jakub Klímek, Martin Necaský
iiWAS2
2019 Explainable Similarity of Datasets Using Knowledge Graph
Petr Skoda 0001, Jakub Klímek, Martin Necaský, Tomás Skopal
SISAP2
2019 DCAT-AP representation of Czech National Open Data Catalog and its impact
Jakub Klímek
J. Web Semant.1
2018 Publication and usage of official Czech pension statistics Linked Open Data
Jakub Klímek, Jan Kucera 0002, Martin Necaský, Dusan Chlapek
J. Web Semant.1
2017 LinkedPipes ETL in use: practical publication and consumption of linked data
abstract
Companies and institutions now realize the potential of Linked Open Data (LOD) and they start publishing their own data as LOD. However, publishing LOD is still a challenging task. One of the main reasons is a lack of user friendly tooling which would properly support the whole LOD publishing process. The process typically consists of source data extraction, transformation to RDF, alignment with commonly used vocabularies, linking to other datasets, computing metadata, publishing on the web as a dump, loading into a triplestore and recording the dataset in a data catalog such as CKAN. In this paper we present LinkedPipes ETL, a tool for ETL-like LOD publishing, which mainly focuses on supporting such LOD publishing workflows in a user friendly way. In addition, the tool also eases consumption of already existing LOD data sources as it addresses some of the practical issues associated with it. Finally, the tool itself uses Linked Data technologies for representation of the ETL processes. We describe LinkedPipes ETL and its main distinguishing features in context of the use cases in which the tool has already been deployed. They include an institution of public administration, a municipality, a university, a software company and an open data initiative.
Jakub Klímek, Petr Skoda 0001
iiWAS1
2017 Platform for automated previews of linked data
abstract
While the number of Linked Data (LD) datasets grows, the support for their consumption is still quite limited. Data publishers expect high reuse of their LD in different applications. At the same time, their users expect that the applications will be reusable for different LD datasets. However, the reality is different because existing LD often use proprietary vocabularies and unique combinations of nonproprietary ones. In this paper, we introduce a backend of a platform which interconnects LD datasets on one side with applications for previewing the datasets on the other. The platform backend is automatically discovers so called application pipelines for a given set of datasets. Each application pipeline specifies a sequence of transformation steps which transform the original data to a form required by an application. The platform is also able to execute a chosen pipeline and provide the result to the application in a prepared SPARQL endpoint.
Martin Necaský, Jirí Helmich, Jakub Klímek
iiWAS3
2015 Efficient Exploration of Linked Data Cloud
abstract
As the size of semantic data available as Linked Open Data (LOD) increases, the demand for methods for automated exploration of data sets grows as well. A data consumer needs to search for data sets meeting his interest and look into them using suitable visualization techniques to check whether the data sets are useful or not. In the recent years, particular advances have been made in the field, e.g., automated ontology matching techniques or LOD visualization platforms. However, an integrated approach to LOD exploration is still missing. On the scale of the whole web, the current approaches allow a user to discover data sets using keywords or manually through large data catalogs. Existing visualization techniques presume that a data set is of an expected type and structure. The aim of this position paper is to show the need for time and space efficient techniques for discovery of previously unknown LOD data sets on the base of a consumer’s interest and their automated visualization which we address in our ongoing work
Jakub Klímek, Martin Necaský, Bogdan Kostov, Miroslav Blasko, Petr Kremen
DATA1
2013 Formal Linked Data Visualization Model
abstract
Recently, the amount of semantic data available in the Web has increased dramatically. The potential of this vast amount of data is enormous but in most cases it is difficult for users to explore and use this data, especially for those without experience with Semantic Web technologies. Applying information visualization techniques to the Semantic Web helps users to easily explore large amounts of data and interact with them. In this article we devise a formal Linked Data Visualization Model (LDVM), which allows to dynamically connect data with visualizations. We report about our implementation of the LDVM comprising a library of generic visualizations that enable both users and data analysts to get an overview on, visualize and explore the Data Web and perform detailed analyzes on Linked Data.
Josep Maria Brunetti, Sören Auer, Roberto García 0001, Jakub Klímek, Martin Necaský
iiWAS4
2013 Linked Open Data for Healthcare Professionals
abstract
Physicians are overwhelmed with many different drugs and the need to know a lot of information about all of them. That is, however, almost impossible in the fast evolving area of pharmaceutical industry. Although many data sources about drugs are published on the Web, structured or unstructured, it is very time consuming to search through them. In this paper we identify these data sources according to information needs of physicians. We show that they can be relatively easily integrated using the Linked Data principles and, in case of unstructured data, NLP methods. An application on the top of the integrated data sets is presented as a possible tool for clinical decision support.
Jakub Kozák, Martin Necaský, Jan Dedek, Jakub Klímek, Jaroslav Pokorný
iiWAS4
2013 Methodology for Design and Evolution of XML Schemas using Conceptual Modeling
abstract
XML has achieved the leading role among languages for data representation and, thus, the amount of related technologies and applications exploiting them grows fast. However, only a small percentage of applications is static and remains unchanged since its first deployment. Most of the applications change with newly coming user requirements and changing environment. In this paper we describe a framework and a methodology for management of evolution and change propagation throughout XML applications. We also introduce its proof-of-concept implementation called eXolutio, which has been developed and improved in our research group during last few years. We also provide an evaluation of the methodology in the domain of electronic health.
Martin Necaský, Jakub Klímek, Jakub Malý, Irena Holubová
iiWAS2
2012 When conceptual model meets grammar: A dual approach to XML data modeling
Martin Necaský, Irena Holubová, Jakub Klímek, Jakub Malý
Data Knowl. Eng.3