Tomasz Miksa

dblp:150/0536 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0002-4929-7875ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
YearPublicationVenuePosition
2026 You Shall Not Pass (Without Consent): Enforcing Data Sovereignty with Solid Pods
abstract
Privacy-preserving data analysis must carefully balance the need for secure, meaningful computation on sensitive personal data with the fundamental rights of individuals to retain control over their information. Solid (Social Linked Data) presents an open protocol where users store and manage their data in personal, access-controlled pods. However, its potential for integration as a decentralized data store into existing infrastructures for privacy-preserving computations remains underexplored. We address how Solid can be effectively integrated into such platforms to support decentralized data sharing while meeting the technical requirements of privacy-aware research. To address this, we propose the Solid Gateway , a mediator that facilitates consent-driven access to Solid Pods within existing analysis environments. The Solid Gateway introduces request-specific authentication and authorization, manages access permissions, and orchestrates the retrieval of only the data necessary to fulfill individual data requests. Central to this approach is a novel granular data-sharing strategy, which restructures user data into minimal request-specific subsets, thus reducing unnecessary data transfers and limiting the exposure of irrelevant information. This ensures that contributors retain sovereignty over their data, while allowing privacy-preserving analysis to operate on decentralized sources. Our experimental evaluation, conducted on controlled artificial datasets, confirms the feasibility of our integration. The results demonstrate a significant reduction in data exposure while achieving improved data retrieval performance compared to existing approaches. Also, we compare our proposed solution against the WellFort architecture and demonstrate that our approach offers competitive fetch performance and significantly improves processing efficiency. Although the controlled nature of the evaluation limits comparability with existing platforms, it provides a reproducible foundation for future studies and practical deployments. This work contributes a concrete, extensible design for combining Solid with privacy-preserving computation, identifies key tradeoffs between privacy, performance, and system complexity, and opens pathways for future research into SPARQL integration, validation with established datasets, and the application of FAIR principles within Solid .
Tobias Hajszan, Moritz Staudinger, Tomasz Miksa
ACM Trans. Web3
2024 Mission Reproducibility: An Investigation on Reproducibility Issues in Machine Learning and Information Retrieval Research
abstract
This paper analyzes the most common problems limiting reproducibility of Information Retrieval research and provides researchers with insights and guidelines to improve the reproducibility of experiments and to allow the verification of obtained results. We conducted a study on 45 reproduction reports off 17 different papers, which have been published at renowned IR conferences. We analyzed the reports qualitatively and quantitatively and looked into the different insights from different groups. Occurring problems are classified into three problem families and 13 categories and afre then analyzed with respect to their influence on the reproduction process as well as on their frequency of appearance over time and per conference. Of these 17 different papers, 14 papers were reproducible to a certain degree without significant differences to the original results, but in many cases not the whole experiment was reproducible due to missing code, information or data. Also, we look at assumptions that were made when reproducing the different papers, as some experiment workflows were incomplete and information was missing. In addition, we propose recommendations to make machine learning research more reproducible and FAIR.
Moritz Staudinger, Bettina M. J. Kern, Tomasz Miksa, Lukas Arnhold, Peter Knees, Andreas Rauber, Allan Hanbury
e-Science3
2023 Reproducible Query Processing and Data Citation of in Situ Soil Moisture Data
abstract
Data in today's dynamic world undergoes constant change and evolution, spanning various formats such as text, websites, tweets, and sensor readings. Storing and referencing these diverse data types pose significant challenges due to data movement, changes in content or structure, and limited availability. Efficient data identification is crucial for speeding up scientific discovery and result validation, especially when data accessibility is guaranteed. Recent years have witnessed progress in data citation practices, with conferences mandating the inclusion of utilized and generated data. However, existing solutions primarily cater to static datasets, rendering them ineffective for dynamically evolving ones. This paper addresses this gap by providing a tailored dynamic data citation prototype for the International Soil Moisture Network, one of the largest scientific in situ soil moisture databases. Our work encompasses the implementation and evaluation of different data versioning strategies and a query store architecture that enables the citation, reproducibility, and verification of large sets of SQL queries to recreate data requests by users. By applying the RDA Dynamic Data Citation Guidelines, we assess the necessary needs for such a system and further measure the performance and storage impact of our proposed approaches.
Moritz Staudinger, Tobias Hajszan, Tomasz Miksa, Irene Himmelbauer, Daniel Aberer, Andreas Rauber, Wouter Dorigo
e-Science3
2023 Combining Semantic Web and Machine Learning for Auditable Legal Key Element Extraction
Anna Breit, Laura Waltersdorfer, Fajar J. Ekaputra, Sotirios Karampatakis, Tomasz Miksa, Gregor Käfer
ESWC5
2019 Data Identification and Process Monitoring for Reproducible Earth Observation Research
abstract
Earth observation researchers use specialised computing services for satellite image processing offered by various data backends. The source of data is often the same, for example Sentinel-2 satellites operated by Copernicus, but the way how data is pre-processed, corrected, updated, and later analysed may differ among the backends. Backends often lack mechanisms for data versioning, for example, data corrections are not tracked. Furthermore, an evolving software stack used for data processing remains a black box to researchers. Researchers have no means to identify why executions of the same code deliver different results. This hinders reproducibility of earth observation experiments. In this paper, we present how infrastructure of existing earth observation data backends can be modified to support reproducibility. The proposed extensions are based on recommendations of the Research Data Alliance regarding data identification and the VFramework for automated process provenance documentation. We implemented these extensions at the Earth Observation Data Centre, a partner in the openEO consortium. We evaluated the solution on a variety of usage scenarios, providing also performance and storage measures to evaluate the impact of the modifications. The results indicate reproducibility can be supported with minimal performance and storage overhead.
Bernhard Gößwein, Tomasz Miksa, Andreas Rauber, Wolfgang Wagner 0001
eScience2
2019 Ten principles for machine-actionable data management plans
abstract
Data management plans (DMPs) are documents accompanying research proposals and project outputs. DMPs are created as free-form text and describe the data and tools employed in scientific investigations. They are often seen as an administrative exercise and not as an integral part of research practice. There is now widespread recognition that the DMP can have more thematic, machine-actionable richness with added value for all stakeholders: researchers, funders, repository managers, research administrators, data librarians, and others. The research community is moving toward a shared goal of making DMPs machine-actionable to improve the experience for all involved by exchanging information across research tools and systems and embedding DMPs in existing workflows. This will enable parts of the DMP to be automatically generated and shared, thus reducing administrative burdens and improving the quality of information within a DMP. This paper presents 10 principles to put machine-actionable DMPs (maDMPs) into practice and realize their benefits. The principles contain specific actions that various stakeholders are already undertaking or should undertake in order to work together across research communities to achieve the larger aims of the principles themselves. We describe existing initiatives to highlight how much progress has already been made toward achieving the goals of maDMPs as well as a call to action for those who wish to get involved.
Tomasz Miksa, Stephanie Renee Simms, Daniel Mietchen, Sarah Jones
PLoS Comput. Biol.1
2018 Debunking Active Data Management Plans
abstract
This poster focuses on the topic of active or machine-actionable data management plans. It aims at clarifying the concept and present an overview on its state of the art. With particular focus on the work by the RDA DMP Common Standards working group and the 10 rules for maDMP (Miksa et al. 2018).
João Cardoso 0002, Tomasz Miksa, José Borbinha
IEEE BigData2
2018 Framing the scope of the common data model for machine-actionable Data Management Plans
abstract
Currently, research requires processing data at a large scale. Data is not anymore a collection of static documents, but often a continuous stream of information flowing into information systems. Researchers need to manage their data efficiently not only to keep it safe, but also to ensure that it can be later correctly interpreted and reused. Existing solutions are not sufficient. Traditional Data Management Plans are manually created text documents that describe how research data will be handled. Yet, researchers must implement all actions by themselves. Machine-actionable Data Management Plans are a new approach that allows systems to act on behalf of researchers and other stakeholders involved in data management, to help them manage data in an efficient and scalable way. This paper summarises the results of work performed by the Research Data Alliance working group on Data Management Plan Common Standards to realise this vision. The paper describes results of consultations and proof of concept tools that help in: identifying needs for information of stakeholders involved in data management; defining the scope of the common data model for Machine-actionable Data Management Plans to allow for exchange of information between systems; identifying necessary services and components of infrastructure that support automation of data management tasks.
Tomasz Miksa, João Cardoso 0002, José Borbinha
IEEE BigData1
2018 Research Data Preservation Using Process Engines and Machine-Actionable Data Management Plans
Asztrik Bakos, Tomasz Miksa, Andreas Rauber
TPDL2
2017 Using ontologies for verification and validation of workflow-based experiments
Tomasz Miksa, Andreas Rauber
J. Web Semant.1
2016 Identifying impact of software dependencies on replicability of biomedical workflows
Tomasz Miksa, Andreas Rauber, Eleni Mina
J. Biomed. Informatics1
2014 Ontologies for Describing the Context of Scientific Experiment Processes
abstract
The re-usability and repeatability of e-Science experiments is widely understood as a requirement of validating and reusing previous work in data-intensive domains. Experiments are, however, often complex chains of processing, involving a number of data sources, computing infrastructure, software tools, or external and third-party services, rendering repeatability a challenging task. Another important aspect of many experiments is in the social and organisational dimension - very often, knowledge on how experiments are performed is tacit and remains with the researcher, and the collaborative and distributed aspects especially of larger collaborative experiments adds to this challenge. Therefore, a number of approaches have tackled this issue from various angles -- initiatives for data sharing, code versioning and publishing as open source, the use of workflow engines to formalise the steps taken in an experiment, to ways to describe the complex environment an experiment is executed in, e.g. via Research Objects. In this paper, we present a model that has a specific focus on the technical infrastructure that is the basis of the research experiment. We demonstrate how this model can be applied to describe e-Science experiments, and align and compare it to Research Objects.
Rudolf Mayer, Tomasz Miksa, Andreas Rauber
eScience2
2014 Resilient Web Services for Timeless Business Processes
abstract
Many business and scientific processes make extensive use of service-oriented architectures, using distributed services. These are often provided by third parties and are thus not under direct control of process owners. In this paper we discuss the issues of ensuring continuous and faithful execution of processes in distributed environments, focusing specifically on Web Services. Recently, we introduced a specification of Resilient Web Services, that makes current Web Services more robust, and a framework for the monitoring of Web Services, that allows detecting anomalies. In this paper, we describe alternative implementations of the framework for monitoring of Web Services. We also present possible approaches easing the deployment of Resilient Web Services: a framework consisting of tools deployable at the Web Service operator site enabling easy transformation of a regular Web Service into a Resilient Web Service, and a registry with notifications that decorates existing Web Services with resilient methods.
Tomasz Miksa, Rudolf Mayer, Marco Unterberger, Andreas Rauber
iiWAS1