Achim Streit

dblp:s/AchimStreit · DBLP profile ↗
← Back
9ranked-venue papers in the field
0as first author
3since 2021 · last 2024
0000-0002-5065-469XORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6Data Mining & Knowledge Discovery · 2Database Systems & Data Management · 1
YearPublicationVenuePosition
2024 FAIR Digital Objects for the Realization of Globally Aligned Data Spaces
abstract
The FAIR principles are globally accepted guidelines for improved data management practices with the potential to align data spaces on a global scale. In practice, this is only marginally achieved through the different ways in which organizations interpret and implement these principles. The concept of FAIR Digital Objects provides a way to realize a domain-independent abstraction layer that could solve this problem, but its specifications are currently diverse, contradictory, and restricted to semantic models. In this work, we introduce a rigorously formalized data model with a set of assertions using formal expressions to provide a common baseline for the implementation of FAIR Digital Objects. The model defines how these objects enable machine-actionable decisions based on the principles of abstraction, encapsulation, and entity relationship to fulfill FAIR criteria for the digital resources they represent. We provide implementation examples in the context of two use cases and explain how our model can facilitate the (re)use of data across domains. We also compare how our model assertions are met by FAIR Digital Objects as they have been described in other projects. Finally, we discuss our results’ adoption criteria, limitations, and perspectives in the big data context. Overall, our work represents an important milestone for various communities working towards globally aligned data spaces through FAIRification.
Nicolas Blumenröhr, Philipp-Joachim Ost, Felix Kraus, Achim Streit
IEEE Big Data4
2024 Model Fusion via Neuron Transplantation
Muhammed Öz, Nicholas Kiefer, Charlotte Debus, Jasmin Hörter, Achim Streit, Markus Götz
ECML/PKDD (4)5
2022 Training Parameterized Quantum Circuits with Triplet Loss
Christof Wendenius, Eileen Kuehn, Achim Streit
ECML/PKDD (5)3
2020 HeAT - a Distributed and GPU-accelerated Tensor Framework for Data Analytics
abstract
To cope with the rapid growth in available data, the efficiency of data analysis and machine learning libraries has recently received increased attention. Although great advancements have been made in traditional array-based computations, most are limited by the resources available on a single computation node. Consequently, novel approaches must be made to exploit distributed resources, e.g. distributed memory architectures. To this end, we introduce HeAT, an array-based numerical programming framework for large-scale parallel processing with an easy-to-use NumPy-like API. HeAT utilizes PyTorch as a node-local eager execution engine and distributes the workload on arbitrarily large high-performance computing systems via MPI. It provides both low-level array computations, as well as assorted higher-level algorithms. With HeAT, it is possible for a NumPy user to take full advantage of their available resources, significantly lowering the barrier to distributed data analysis. When compared to similar frameworks, HeAT achieves speedups of up to two orders of magnitude.
Markus Götz, Charlotte Debus, Daniel Coquelin, Kai Krajsek, Claudia Comito, Philipp Knechtges, Björn Hagemeier, Michael Tarnawa, Simon Hanselmann, Martin Siggel, Achim Basermann, Achim Streit
IEEE BigData12
2018 Concept and Analysis of Information Spaces to improve Prediction-Based Compression
abstract
One of the scientific communities that generate the largest amounts of data today are the climate sciences. New climate models enable model integration at unprecedented resolution, simulating decades and centuries of climate change, including many complex interactions in the Earth system, under different scenarios. Previously, the CPU intensive numerical integration’s used to be the bottleneck. Nowadays, limited storage space and ever increasing model output is the bigger challenge. The number of variables stored for post-processing analysis has to be limited to keep the data amounts small. For this reason, we look at lossless compression of climate data to make better use of available storage space. More specifically, we investigate prediction-based data compression. In prediction-based compression, data is processed in a predefined sequence. A prediction is provided for each data point based on prior data in the sequence. We show that there is a significant dependence of the compression ratio on the chosen traversal method and the underlying spatiotemporal data model. We examine the influence of this structural dependency on compression algorithms and explore possibilities to retrieve this information to improve compression ratios. To do this, we introduce the concept of Information Spaces (IS), which helps improve the predictions made by individual predictors by nearly 10% on average. More importantly, the standard deviation of the compression results is decreased by over 20% on average. The use of IS provides better predictions and more consistent compression ratios. Furthermore, it allows options for consolidation and fine-granular tuning of predictions, which are not possible with many common approaches used today.
Ugur Çayoglu, Frank Tristram, Jörg Meyer 0001, Tobias Kerzenmacher, Peter Braesicke, Achim Streit
IEEE BigData6
2018 The Challenge of a Strong Speed-Up of a Bio-Medical Big Data Application
abstract
Digital data of patients can aid a pathologist along the diagnostic process [1]. Medical devices generate data sets that are processed by specialized computing applications, which often run on a single computer. The resolution power of the devices is increasing steadily and, consequently, the volumes of the data sets are also growing and can no longer be analyzed in a reasonable amount of time. Big Data tools like Apache Spark [2] provide methods for analyzing data, however, they are not directly applicable and need considerable implementation efforts, in general. Usually, well established analysis tools for medical data are designed to run on single workstations. These tools are not designed to meet current and future challenges. Migrating processing tools from single nodes to distributed environments is nontrivial. Moreover, partitioning data sets for a parallel processing is a further challenge [3].In this work, we continue our efforts for improving the speedup of a bio-medical big data application further by partitioning the images of a Whole Slide Image (WSI) [4] into sub–tiles and by analyzing these sub–tiles on a cluster of computer nodes. The idea is to benefit from the divide and conquer strategy. However, it is shown that the score parameter is determined incorrectly, when the software package is applied to each sub–tile and the score parameters of all sub–tiles are combined in an apparently natural manner. The cause of this anomaly is determined and a solution suggested. The original software is based on implicit assumptions. For example, the size of the tiles is assumed to be 1024 × 1024 px2. The anomaly shows up when this constraint is reduced.
Marco Strutz, Bjoern Lindequist, Hermann Heßling, Achim Streit
IEEE BigData4
2018 A modular software framework for compression of structured climate data
abstract
Through the introduction of next-generation models the climate sciences have experienced a breakthrough in high-resolution simulations. In the past, the bottleneck was the numerical complexity of the models, nowadays it is the required storage space for the model output. One way to tackle the data storage challenge is through data compression.
Ugur Çayoglu, Jennifer Schröter, Jörg Meyer 0001, Achim Streit, Peter Braesicke
SIGSPATIAL/GIS4
2015 On a new approach to the index selection problem using mining algorithms
abstract
Considering the wide usage of databases and their ever growing size, it is crucial to improve the query processing performance. Selection of an appropriate set of indexes for the workload processed by the database system is an important part of physical design and performance tuning. This selection is a non-trivial tasks, especially considering possible number of native indexes in modern databases. We introduce a new approach to the index selection problem using data mining. The method recommends the creation of indexes as well as the type of each index. This results in more precise index recommendations that allows not only to create ascending and descending indexes, but also special indexes supported by the database system. Mining of queries results in candidate indexes for which virtual indexes get created. As the approach does not require modifications of the database system, it is generically applicable. Evaluations of the scalability are given for different workloads for the document-based NoSQL database MongoDB.
Parinaz Ameri, Jörg Meyer 0001, Achim Streit
IEEE BigData3
2014 Evaluating the performance and scalability of the Ceph distributed storage system
abstract
As the data needs in every field continue to grow, storage systems have to grow and therefore need to adapt to the increasing demands of performance, reliability and fault tolerance. This also increases their complexity and costs. Improving the performance and scalability of storage systems while maintaining low costs is thus crucial. The evaluated open source storage system Ceph promises to reliably store data distributed across many nodes. Ceph is targeted at commodity hardware. This study investigates how Ceph performs in different setups and compares this with the theoretical maximum performance of the hardware. We used a bottom-up approach to benchmark Ceph at different architectural levels. We varied the amount of storage nodes and clients to test the scalability of the system. Our experiments revealed that Ceph delivers the promised scalability, and uncovered several points with improvement potential. We observed a significant increase of the write throughput by moving the Ceph journal to a faster location (in memory). Moreover, while the system scaled with the increasing number of clients operating the cluster, we noticed a slight performance degradation after the saturation point. We tested two optimisation strategies - increasing the available RAM or the object size - and noted a write throughput increase of up to 9% and 27%, respectively. Our findings improve the understanding of Ceph and should benefit future users through the presented strategies for tackling various performance limitations.
Diana Gudu, Marcus Hardt, Achim Streit
IEEE BigData3