EDBT 2026 Demo / reviewers in the wild / expert
Tobias Hahn
dblp:31/4193
· DBLP profile ↗
8ranked-venue papers
6as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 first-author · 5 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ABACUS: ASIP-Based Avro Schema-Customizable Parser Acceleration on FPGAsabstractBig Data applications frequently process data streams encoded in semi-structured data formats such as JSON, Protobuf, or Avro. Parsing these data formats into a representation to then be processed by a CPU frequently takes up a major share of the processing time. As a remedy, JSON and Avro FPGA accelerators have been introduced that can parse the data directly in the data path, offloading this workload from the CPU without requiring any additional data movement. However, these accelerators are schema-specific circuits that require time-consuming resynthesis processes for schema adaptations. This is particularly critical in Big Data applications, where multiple schemas may be in use simultaneously. As a remedy, we present an application-specific instruction set processor (ASIP) architecture for parsing Avro data on FPGAs. An instruction program controls the ASIP to parse a specific schema. Any schema change therefore only requires the loading of a new instruction sequence into an instruction memory. It is also shown that this approach is more resource-efficient than related work, as functional units only need to be instantiated once for each Avro data type. Our experimental evaluation shows that we can achieve a throughput of 707–818 MB/s per kLUT which is about 7 to 14 times higher than the throughput per LUT achieved in related work. Tobias Hahn, Daniel Schüll, Stefan Wildermann, Jürgen Teich |
DDECS | 1 |
| 2024 | JSON-CooP: A JSON Decompression/Parsing Co-Design for FPGAsabstractBig Data applications frequently involve the processing of data streams encoded in semi-structured data formats such as JSON. A major challenge here is that the parsing of such data formats is usually highly complex. Accelerating JSON parsing on FPGAs has therefore become a focus of recent research. However, as JSON data is highly sparse, compression is frequently applied before writing records to storage or transmitting them over the network. Consequently, the data must be decompressed before parsing it into a suitable format for further processing.While existing work has addressed decompression and parsing of semi-structured data separately, we propose a co-design for JSON decompression/parsing. This co-design includes a compression scheme tailored specifically for JSON, along with an FPGA parser architecture capable of operating directly on compressed input data. Additionally, our design adopts a lazy decompression approach to only decompress projected attributes, thereby significantly reducing the workload on the decompressor. Our experimental evaluation shows that this co-design exploiting several synergies results in higher resource efficiency compared to existing approaches for JSON parsing on FPGAs, while also enabling the decompression of data. For high compression factors, we observe a 1.8 x increase in parsed JSON tuples per second and LUT compared to the most efficient related approach. Moreover, our presented compression scheme offers higher compression factors for JSON data than related work while providing similar performance. Tobias Hahn, Stefan Wildermann, Jürgen Teich |
FPL | 1 |
| 2023 | SPEAR-JSON: Selective Parsing of JSON to Enable Accelerated Stream Processing on FPGAsabstractBig Data applications frequently involve the processing of data streams encoded in semi-structured data formats such as JSON. A major challenge here is that the parsing of such data formats is usually highly complex. Accelerating JSON parsing on FPGAs has therefore become a focus of recent research. FPGA accelerators were presented which serve as a co-processor for a CPU to convert JSON into a format that is easier for the CPU to process, e.g., Apache Arrow. However, in case the parsed data should be further processed on the FPGA, such solutions are insufficient as the format created is unsuitable for further processing on FPGAs and, above all, because the accelerators have an immense resource requirement. In this paper, we present a novel FPGA parser architecture that is able to interpret JSON data to selectively extract attributes based on a query expression into a format suitable for stream processing on FPGAs. Furthermore, it is shown how the sparsity of JSON can be used to implement a resource-efficient design, only requiring few FPGA resources. This leaves the major share of resources free for accelerating subsequent processing steps of a given application. Our experimental evaluation shows that we can achieve a throughput of 51.1 MB/s per kLUT which is about 3.8 times higher than the throughput per LUT achievable on the most efficient related approach. Tobias Hahn, Stefan Wildermann, Jürgen Teich |
FPL | 1 |
| 2022 | Raw Filtering of JSON Data on FPGAsabstractMany Big Data applications include the processing of data streams on semi-structured data formats such as JSON. A disadvantage of such formats is that an application may spend a significant amount of processing time just on unselectively parsing all data. To relax this issue, the concept of raw filtering is proposed with the idea to remove data from a stream prior to the costly parsing stage. However, as accurate filtering of raw data is often only possible after the data has been parsed, raw filters are designed to be approximate in the sense of allowing false-positives in order to be implemented efficiently. Contrary to previously proposed CPU-based raw filtering techniques that are restricted to string matching, we present FPGA-based primitives for filtering strings, numbers and also number ranges. In addition, a primitive respecting the basic structure of JSON data is proposed that can be used to further increase the accuracy of introduced raw filters. The proposed raw filter primitives are designed to allow for their composition according to a given filter expression of a query. Thus, complex raw filters can be created for FPGAs which enable a drastical decrease in the amount of generated false-positives, particularly for IoT workload. As there exists a trade-off between accuracy and resource consumption, we evaluate primitives as well as composed raw filters using different queries from the RiotBench benchmark. Our results show that up to 94.3% of the raw data can be filtered without producing any observed false-positives using only a few hundred LUTs. Tobias Hahn, Andreas Becher, Stefan Wildermann, Jürgen Teich |
DATE | 1 |
| 2022 | Auto-Tuning of Raw Filters for FPGAsabstractMany Big Data applications include the processing of data streams on semi-structured data formats such as JSON. A disadvantage of these formats, however, is that applications may require a significant portion of their processing time to unselectively parse all data. As a remedy, so-called raw filters have been introduced in the past, aiming to reduce the data load before the costly parsing stage. Since filtering unparsed data can also become very costly, raw filters can be designed to filter data approximately, in the sense that they allow false positives to occur, in order to be implemented efficiently. While previously proposed CPU-based solutions are restricted to just string filtering, FPGA approaches have recently been proposed with much more expressive raw filters, allowing also to capture numbers and structural relationships. Yet, as a consequence of the variety of filter possibilities as well as the limited amount of resources available on FPGAs, the selection of optimal filters before their deployment has been identified as a complex problem resulting in the potential need to select less expressive filters in order to consume fewer resources. Many Big Data applications (e.g., stream processing) operate on incoming real-time data over long, potentially unlimited time periods. As a consequence, the conditions for which such a filter is optimized can change over time after its deployment. In this realm, this paper presents a new methodology which automatically adapts the hardware accelerator for raw filtering by means of dynamic hardware reconfiguration. Data is sampled on-the-fly during operation and used by an optimizer-in-the-loop to select and generate a raw filter with optimized selectivity for these data samples. As the optimizer has to take into account the resource costs of the hardware accelerator, we introduce models to estimate the resource costs in order to avoid performing a full synthesis. The filter selection problem can thus be solved within a few minutes with results close to the accurate resource cost estimation. If the selectivity of a query changes over time, such as seasonal differences in the analysis of IoT data, the system can auto-tune its filter to adapt to the situation. Depending on the query and the variability of inherent data changes, significant improvements in the amount of filtered data are presented, resulting in a significant parsing speedup in comparison to a state-of-the-art non-adaptive approach. Tobias Hahn, Stefan Wildermann, Jürgen Teich |
FPL | 1 |
| 2019 | Scheduling Self-Suspending Tasks: New and Old ResultsabstractIn computing systems, a job may suspend itself (before it finishes its execution) when it has to wait for certain results from other (usually external) activities. For real-time systems, such self-suspension behavior has been shown to induce performance degradation. Hence, the researchers in the real-time systems community have devoted themselves to the design and analysis of scheduling algorithms that can alleviate the performance penalty due to self-suspension behavior. As self-suspension and delegation of parts of a job to non-bottleneck resources is pretty natural in many applications, researchers in the operations research (OR) community have also explored scheduling algorithms for systems with such suspension behavior, called the master-slave problem in the OR community. This paper first reviews the results for the master-slave problem in the OR literature and explains their impact on several long-standing problems for scheduling self-suspending real-time tasks. For frame-based periodic real-time tasks, in which the periods of all tasks are identical and all jobs related to one frame are released synchronously, we explore different approximation metrics with respect to resource augmentation factors under different scenarios for both uniprocessor and multiprocessor systems, and demonstrate that different approximation metrics can create different levels of difficulty for the approximation. Our experimental results show that such more carefully designed schedules can significantly outperform the state-of-the-art. Jian-Jia Chen, Tobias Hahn, Ruben Hoeksma, Nicole Megow, Georg von der Brüggen |
ECRTS | 2 |
| 2012 | Vulnerabilities through Usability Pitfalls in Cloud Services: Security Problems due to Unverified Email AddressesabstractCloud storage services become increasingly interesting for users to easily backup or synchronize their data. On top of this basic functionality, these services offer functions for collaboration that allow users to share their files with selected other persons in a user-friendly way. We have identified that several cloud storage services do not verify whether the registrating customer is the real owner of the email address entered during the registration. Cloud providers omit the verification for reasons of usability. Here, user-friendliness goes too far at the cost of security. This vulnerability combined with collaboration functions allows attacks on cloud customers. In this paper, we explain which attacks are possible. Missing email verification and collaboration functions allow espionage and malware distribution attacks. Execution is very easy, i.e., they can be done without coding expertise or special tools. Tobias Hahn, Thomas Kunz, Markus Schneider 0002, Sven Vowe |
TrustCom | 1 |
| 2007 | All-Optoelectronic Terahertz Imaging Systems and Examples of Their ApplicationabstractWe give an overview over several all-optoelectronic measurement systems which we have developed for transmittive and reflective imaging in the terahertz (THz) frequency range. The systems employ either pulsed or continuous-wave THz radiation. In both cases, they work on the basis of single-pixel scanning. Addressing the potential for imaging in the medical and dental field, and the application of THz radiation for industrial surface and interface characterization, we explore dark-field imaging where the imaging contrast originates from diffraction and scattering effects coming from topography or refractive-index variations. Torsten Loffler, Karsten J. Siebert, Noboru Hasegawa, Tobias Hahn, Hartmut G. Roskos |
Proc. IEEE | 4 |