EDBT 2026 Demo / reviewers in the wild / expert
Thilo Pionteck
dblp:52/851
· DBLP profile ↗
34ranked-venue papers
12as first author
11since 2021 · last 2024
0000-0001-6518-1226ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 11 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring the Versal AI Engines for Signal Processing in Radio AstronomyabstractNowadays, heterogeneous architectures are widely used to overcome the ongoing demand for increased computing performance at the edge, such as the pre-processing of raw antenna data in radio telescope systems. The Versal Adaptive SoC is a novel heterogeneous architecture that includes Programmable Logic (PL), a Processing System, and Artificial Intelligence Engines (AIEs) interconnected by a programmable Network-on-Chip. In this work, we explore the AIEs to evaluate their capabilities for real-time signal processing in radio telescope systems. We focus on the implementation of a Polyphase Filter Bank (PFB), which is a representative signal processing operation that consists of Finite Impulse Response (FIR) filters and the Fast-Fourier Transform (FFT) algorithm. We analyzed the performance of the AIEs with regard to the requirements of LOFAR, the world’s largest low-frequency radio telescope. By means of the roofline model, we reveal that the AIEs theoretically meet the LOFAR computational requirements, but the FIR vendor-provided library implementation did not reach the required performance. Therefore, we explore several optimization strategies for the FIR implementation on the AIEs and analyze the communication options between the PL and the AIE array. Finally, we have developed an efficient PFB implementation that requires only 12 AIEs. A prototype on a VC1902 device achieves a throughput of 437 MSPS, which is more than sufficient for a single antenna polarization of the LOFAR system. This work is open source and publicly available at: https://git.astron.nl/rd/acap. Victor Van Wijhe, Vincent Sprave, Daniele Passaretti, Nikolaos Alachiotis 0001, Gerrit Grutzeck, Thilo Pionteck, Steven van der Vlugt |
FPL | 6 |
| 2023 | What Happens When Two Multi-Query Optimization Paradigms Combine? - A Hybrid Shared Sub-Expression (SSE) and Materialized View Reuse (MVR) Study
Bala Gurumurthy, Vasudev Raghavendra Bidarkar, David Broneske, Thilo Pionteck, Gunter Saake |
ADBIS | 4 |
| 2023 | A Flexible and Scalable Reconfigurable FPGA Overlay Architecture for Data-Flow ProcessingabstractWe present a flexible and scalable FPGA overlay architecture for data-flow applications. The overlay consists of a 2D grid of tiles consisting of compute units that can be exchanged at runtime. The overlay can deal with unbalanced data-flow graphs and is based on AXI-Stream to facilitate extensibility using IP and High-Level Synthesis. To support compute units that also require random access to the system memory, AXI4-ports can be enabled per tile at design-time. We implement a prototype of the overlay architecture tailored to the application domain of analytical query processing on a Xilinx Alveo U280 board. In this system, an overlay grid of$11\times 4$compute unit tiles occupies one-third of the available resources. with a design I/O throughput of$11\times 3.75\ \text{GB}/\mathrm{s}$. The overlay and HLS-based SIMD compute units provide full-throughput data processing, but are limited by the memory subsystem implemented with vendor IPs. Anna Drewes, Vitalii Burtsev, Bala Gurumurthy, Martin Wilhelm, David Broneske, Gunter Saake, Thilo Pionteck |
FCCM | 7 |
| 2023 | ADAMANT: A Query Executor with Plug-In Interfaces for Easy Co-processor IntegrationabstractToday’s processor landscape is increasingly heterogeneous with the availability of co-processors. This landscape impacts query engines, as they need to be reworked to keep competitive performance by leveraging the underlying architectures. Such a rework might be costly if, for each external processor or SDK, peripheral components needed to be developed as well; resulting in redundant effort and adoption difficulties. In this paper, we propose an approach to overcome these shortcomings through ADAMANT – a query executor equipped with interfaces to plug-in new co-processors without reworking other components of a query engine. ADAMANT consists of 1) pluggable interfaces that allow interaction with co-processors, encapsulating operator implementations, and 2) a unified runtime that handles the execution on arbitrary co-processors, with a chunked execution model for scalable query processing. To evaluate ADAMANT’s versatility, we plug different implementations of a CPU/GPU-based system (using OpenCL, OpenMP, & CUDA) and analyze their performance on TPC-H queries. We identify a 4x performance difference between an arbitrary chunked execution vs. a more architecturally conscious pipelined execution. Furthermore, our comparisons with HeavyDB show complex performance variations from speed-ups up to a factor of 2x from our hardware-conscious execution. We envision initiatives like ADAMANT to ease the study of complex optimizations required in co-processor systems, paving the way for efficient and portable data management tools without cutbacks. Bala Gurumurthy, David Broneske, Gabriel Campero Durand, Thilo Pionteck, Gunter Saake |
ICDE | 4 |
| 2023 | A comprehensive modeling approach for the task mapping problem in heterogeneous systems with dataflow processing unitsabstractSummary We introduce a new model for the task mapping problem to aid in the systematic design of algorithms for heterogeneous systems including, but not limited to, CPUs, GPUs, and FPGAs. A special focus is set on the communication between the devices, its influence on parallel execution, as well as on device‐specific differences regarding parallelizability and streamability. We give a comprehensive description on how a given task mapping can be abstractly evaluated including mappings to dataflow‐based hardware accelerators. We show how this model can be utilized in different system design phases and present two novel mixed‐integer linear programs to demonstrate the usage of the model, showing significant improvements compared to pure CPU mapping for randomly generated task graphs. To the best of our knowledge, we present the first ILP for task mapping that considers pipelining effects when streaming tasks on an FPGA. Martin Wilhelm, Hanna Geppert, Anna Drewes, Thilo Pionteck |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Novel insights on atomic synchronization for sort-based group-by on GPUsabstractAbstract Using heterogeneous processing devices, like GPUs, to accelerate relational database operations is a well-known strategy. In this context, the operation is highly interesting for two reasons. Firstly, it incurs large processing costs. Secondly, its results (i.e., aggregates) are usually small, reducing data movement costs whose compensation is a major challenge for heterogeneous computing. Generally, for computation on GPUs, one relies either on sorting or hashing. Today, empirical results suggest that hash-based approaches are superior. However, by concept, hashing induces an unpredictable memory access pattern conflicting with the architecture of GPUs. This motivates studying why current sort-based approaches are generally inferior. Our results indicate that current sorting solutions cannot exploit the full parallel power of modern GPUs. Experimentally, we show that the issue arises from the need to synchronize parallel threads that access the shared memory location containing the aggregates via . Our quantification of the optimal performance motivates us to investigate how to minimize the overhead of atomics. This results in different variants using atomics, where the best variants almost mitigate the atomics overhead entirely. The results of a large-scale evaluation reveal that our approach achieves a 3x speed-up over existing sort-based approaches and up to 2x speed-up over hash-based approaches. Bala Gurumurthy, David Broneske, Martin Schäler, Thilo Pionteck, Gunter Saake |
Distributed Parallel Databases | 4 |
| 2023 | Hybrid CPU/GPU/APU accelerated query, insert, update and erase operations in hash tables with string keysabstractAbstract Modern computer systems can use different types of hardware acceleration to achieve massive performance improvements. Some accelerators like FPGA and dedicated GPU (dGPU) need optimized data structures for the best performance and often use dedicated memory. In contrast, APUs, which are a combination of a CPU and an integrated GPU (iGPU), support shared memory and allow the iGPU to work together with the CPU on pointer-based data structures. First, we develop an approach for dGPU to accelerate queries in libcuckoo and robin-map and when looking at accelerating insert, updates and erase operations in the original libcuckoo using OneAPI on an APU. We evaluate the dGPU against the CPU variants and our dGPU approach adapted for the CPU and also in a hybrid context by using longer keys on the CPU and shorter keys on the dGPU. In comparison with the original libcuckoo algorithm, our dGPU approach achieves a speed-up of 2.1, and in comparison with the robin-map a speed-up of 1.5. For hybrid workloads, our approach is efficient if long keys are processed on the CPU and short keys are processed on the dGPU. By processing a mixture of 20% long keys on the CPU and 80% short keys on dGPU, our hybrid approach has a 40% higher throughput than the CPU only approach. In addition, we develop a hybrid APU approach for insert, update and erase operations in the original libcuckoo structure focusing on shared memory with iGPU accelerated look-ups of the positions for insert, update and erase operations. Tobias Groth, Sven Groppe, Thilo Pionteck, Franz Valdiek, Martin Koppehel |
Knowl. Inf. Syst. | 3 |
| 2022 | Accelerated Parallel Hybrid GPU/CPU Hash Table Queries with String Keys
Tobias Groth, Sven Groppe, Thilo Pionteck, Franz Valdiek, Martin Koppehel |
DEXA (2) | 3 |
| 2021 | Bridging the Frequency Gap in Heterogeneous 3D SoCs through Technology-Specific NoC Router ArchitecturesabstractIn heterogeneous 3D System-on-Chips (SoCs), NoCs with uniform properties suffer one major limitation; the clock frequency of routers varies due to different manufacturing technologies. For example, digital nodes allow for a higher clock frequency of routers than mixed-signal nodes. This large frequency gap is commonly tackled by complex and expensive pseudo-mesochronous or asynchronous router architectures. Here, a more efficient approach is chosen to bridge the frequency gap. We propose to use a heterogeneous network architecture. We show that reducing the number of VCs allows to bridge a frequency gap of up to 2x. We achieve a system-level latency improvement of up to 47% for uniform random traffic and up to 59% for PARSEC benchmarks, a maximum throughput increase of 50%, up to 68% reduced area and 38% reduced power in an exemplary setting combining 15-nm digital and 30-nm mixed-signal nodes and comparing against a homogeneous synchronous network architecture. Versus asynchronous and pseudo-mesochronous router architectures, the proposed optimization consistently performs better in area, in power and the average flit latency improvement can be larger than 51%. Jan Moritz Joseph, Lennart Bamberg, Geonhwa Jeong, Ruei-Ting Chien, Rainer Leupers, Alberto García Ortiz, Tushar Krishna, Thilo Pionteck |
ASP-DAC | 8 |
| 2021 | Configurable Pipelined Datapath for Data Acquisition in Interventional Computed TomographyabstractIn this paper, we present a novel Configurable Pipelined Datapath for custom Data Acquisition Systems (DASs) in the context of Interventional Computed Tomography (iCT). The introduction of new medical techniques during surgery results in a multitude of challenging requirements on the hardware architecture (e.g., real-time data processing, high data-rate, hard deadline on data acquisition and storage). Therefore, we propose a Pipelined Datapath for System-on-Chip FPGAs, that can be configured at design time and run time. It supports all data processing steps from the detector sensors' data acquisition to the image-processing reconstruction system. It can be configured, depending on the CT physical limitations and the medical scenarios. Besides, the Pipelined Datapath avoids any buffering on off-chip memory, differently from other architectures in literature. Daniele Passaretti, Thilo Pionteck |
FCCM | 2 |
| 2021 | CuART - a CUDA-based, scalable Radix-Tree lookup and update engineabstractIn this work we present an optimized version of the Adaptive Radix Tree (ART) index structure for GPUs. We analyze an existing GPU implementation of ART (GRT), identify bottlenecks and present an optimized data structure and layout to improve the lookup and update performance. We show that our implementation outperforms the existing approach by a factor up to 2 times for lookups and up to 10 times for updates using the same GPU. We also show that the sequential memory layout presented here is beneficial for lookup-intensive workloads on the CPU, outperforming the ART by up to 10 times. We analyze the impact of the memory architecture of the GPU, where it becomes visible that traditional GDDR6(X) is beneficial for the index lookups due to the faster clock rates compared to High Bandwidth Memory (HBM). Martin Koppehel, Tobias Groth, Sven Groppe, Thilo Pionteck |
ICPP | 4 |
| 2020 | Hardware-aided update acceleration in a hybrid Semantic Web database system
Dennis Heinrich, Stefan Werner 0002, Christopher Blochwitz, Thilo Pionteck, Sven Groppe |
J. Supercomput. | 4 |
| 2019 | System-Level Optimization of Network-on-Chips for Heterogeneous 3D System-on-ChipsabstractFor a system-level design of Networks-on-Chip for 3D heterogeneous System-on-Chip (SoC), the locations of components, routers and vertical links are determined from an application model and technology parameters. In conventional methods, the two inputs are accounted for separately; here, we define an integrated problem that considers both application model and technology parameters. We show that this problem does not allow for exact solution in reasonable time, as common for many design problems. Therefore, we contribute a heuristic by proposing design steps, which are based on separation of intralayer and interlayer communication. The advantage is that this new problem can be solved with well-known methods. We use 3D Vision SoC case studies to quantify the advantages and the practical usability of the proposed optimization approach. We achieve up to 18.8% reduced white space and up to 12.4% better network performance in comparison to conventional approaches. Jan Moritz Joseph, Dominik Ermel, Lennart Bamberg, Alberto García Ortiz, Thilo Pionteck |
ICCD | 5 |
| 2019 | Crosstalk optimization for through-silicon vias by exploiting temporal signal misalignment
Lennart Bamberg, Jan Moritz Joseph, Thilo Pionteck, Alberto García Ortiz |
Integr. | 3 |
| 2019 | Simulation environment for link energy estimation in networks-on-chip with virtual channels
Jan Moritz Joseph, Lennart Bamberg, Imad Hajjar, Robert Schmidt 0003, Thilo Pionteck, Alberto García Ortiz |
Integr. | 5 |
| 2018 | Hardware-Accelerated Index Construction for Semantic WebabstractIn this paper, an optimized data structure for managing triples used in a Semantic Web Database and a hardwareengine for index construction are presented. We propose anFPGA-centric design, which we call Hardware-Triplestore. Aspart of the design, a scalable and parallel architecture forTriplestore construction is introduced. We propose a hybrid datastructure consisting of three layers, one for every element ofthe semantic triple. The data structure is optimized for ourhardware-centric design and is stored on an external DDR4-Memory. The Hardware-Triplestore is evaluated separately fromthe rest of the database system and achieves an insertion rateof 1.24 million triples per second, which is 17 times faster thanone of the fastest software Triplestore-RDF-3X-. Christopher Blochwitz, Julian Wolff, Mladen Berekovic, Dennis Heinrich, Sven Groppe, Jan Moritz Joseph, Thilo Pionteck |
FPT | 7 |
| 2018 | Efficient Inter-Kernel Communication for OpenCL Database Operators on FPGAsabstractMany modern database engines use OpenCL to target heterogeneous hardware. Queries are evaluated by execution of chains of low-level operators. The common paradigm for OpenCL workloads facilitates communication between kernels using buffers in off-chip memory. This poses a severe performance limitation due to weak memory systems of FPGAs in contrast to the memory hierarchy available in CPUs and GPUs. To overcome this bottleneck, we propose the use of structural optimizations of kernel code. On-chip pipelining and code fusion are analyzed as alternatives to buffer-based inter-kernel communication. We assess the impact on resource utilization and system throughput and thereby demonstrate that properly structured code achieves a speedup of more than 4x over the default paradigm. This shows that it is essential for chains of kernels to consider not only optimization techniques for individual kernels, but also optimization of inter-kernel communication. Tobias Drewes, Jan Moritz Joseph, Bala Gurumurthy, David Broneske, Gunter Saake, Thilo Pionteck |
FPT | 6 |
| 2016 | Accelerated join evaluation in Semantic Web databases by using FPGAsabstractSummary While the amount of information steadily increases, the requirements on the response time to query these information become more strict. Under those conditions, conventional database systems reach their limits and cannot meet these performance requirements anymore. In recent years, systems with many processing cores are considered to satisfy these demands. Furthermore, these systems include more and more heterogeneous cores tailor‐made to solve one specific task in an efficient manner. However, dedicated hardware accelerators are inflexible and cannot be adapted to the requirements of a dedicated query. Thus, the challenge is orchestrating the diversity of the functionality of all the cores to be optimized for performance/energy efficiency. In this paper, a concept is introduced on how to develop a flexible Field‐Programmable Gate Arrays (FPGA)‐based hardware accelerator to improve the performance of query evaluation in a Semantic Web database. As a first step to the hardware/software system, several joint algorithms are implemented on an FPGA and evaluated against a well‐developed software solution (implemented in C). The comparison shows a significant speedup of up to 10 times. Because of the complexity of the join operator, it is promising that the overall performance of query evaluation can be further enhanced by processing whole queries on an FPGA. Copyright © 2015 John Wiley & Sons, Ltd. Stefan Werner 0001, Dennis Heinrich, Marc Stelzner, Volker Linnemann, Thilo Pionteck, Sven Groppe |
Concurr. Comput. Pract. Exp. | 5 |
| 2016 | RAW 2014: Random Number Generators on FPGAsabstractRandom numbers are important ingredients in a number of applications. Especially in a security context, they must be well distributed and unpredictable. We investigate the practical use of random number generators (RNGs) that are built from digital elements found in FPGAs. For this, we implement different types of ring oscillators (ROs) and memory collision-based circuits on FPGAs from major vendors. Implementing RNGs on the same device as the rest of the system benefits an overall reduction of vulnerability to attacks and wire tapping. Nevertheless, we investigate different attacks by tampering with power supply, chip temperature, and by exposition to strong magnetic fields and X-radiation. We also consider their usability as massively deployed components, whose functionality cannot be tested individually anymore, by conducting a technology invariance experiment. Our experiments show that BlockRAM-based RNGs cannot be considered as a suitable entropy source. We further show that RO-based RNGs work reliably under a wide range of operating conditions. While magnetic fields and X-rays did not induce any notable change, voltage and temperature variations caused an increase in propagation delays within the circuits. We show how reliable RNGs can be constructed and deployed on FPGAs. Michael Raitza, Markus Vogt, Christian Hochberger, Thilo Pionteck |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2010 | Optimizing Runtime Reconfiguration DecisionsabstractPartially reconfigurable hardware accelerators enable the offloading of computative intensive tasks from software to hardware at runtime. Beside handling the technical aspects, finding a proper reconfiguration point in time is of great importance for the overall system performance. Determination of a suitable point of reconfiguration demands the evaluation of performance degradation during runtime reconfiguration and expected performance benefit after reconfiguration. Three different approaches to determine a proper point of reconfiguration are discussed. Delays and weighted transitions are used to reduce the number of reconfigurations while keeping system performance at a maximum. Evaluation is done with a simulation model of a runtime reconfigurable network coprocessor. Results show that the number of reconfigurations can be reduced by about 35% for a given application scenario. By optimizing runtime reconfiguration decisions, the overall system performance is even higher than compared to pure threshold based reconfiguration decision schemes. Thilo Pionteck, Steffen Sammann, Carsten Albrecht |
EUC | 1 |
| 2008 | On the design parameters of runtime reconfigurable systemsabstractThis paper explores the design space for runtime reconfigurable systems. A broad range of systems is surveyed and a set of parameters applicable for characterizing runtime reconfigurable systems is proposed. Compared to other surveys the focus is set on the system architecture, not on the underlying hardware structure. This allows a discussion that primarily considers the actual motivation for utilising runtime reconfiguration in system designs instead of discussing the limitations of actual hardware platforms. Thilo Pionteck, Carsten Albrecht, Roman Koch, Erik Maehle |
FPL | 1 |
| 2008 | Network processorsabstractTraditional design of network processors is complicated by two conflicting demands, flexibility and performance. On the one side, network processors should be flexible enough to adapt to changing protocols and varying traffic profiles, on the other side they have to cope with increasing data rates of network links. This demonstrator shows that runtime reconfigurable systems have the potential to optimise both criteria without affecting each other negatively. The demonstrator addresses edge router applications and consists of two independently developed subsystems, the FlexPath NP architecture designed at the TU Munchen and the Dyna-CORE architecture designed at the University of Lubeck. Thilo Pionteck, Roman Koch, Carsten Albrecht, Erik Maehle, Michael Meitinger, Rainer Ohlendorf, Thomas Wild, Andreas Herkersdorf |
FPL | 1 |
| 2008 | Performance Analysis of Bus-Based Interconnects for a Run-Time Reconfigurable Co-Processor PlatformabstractGrowing bandwidth of network connections as well as strong progress in network protocols and new applications require efficient and flexible network hardware. Network processors are applied for packet processing in routers and gateways. Unfortunately, deep-packet processing tasks lack the support of dedicated co-processors. Because of numerous time-consuming algorithms required, a dynamically re- configurable co-processor for network processors backing payload processing was proposed. It performs computationally intensive tasks without loss of flexibility. A crucial issue is the interconnect of such a system. Bus-based interconnects are explored utilising a software model of this co-processor to determine the performance impact of the on-chip interconnection on the overall performance of the co-processor. A single bus as well as a multiple bus system are evaluated. With regard to reconfiguration overhead, the simulation results show strength and weakness of both systems by latency, throughput, and packet buffer requirements. Carsten Albrecht, Philipp Roß, Roman Koch, Thilo Pionteck, Erik Maehle |
PDP | 4 |
| 2007 | Communication Architectures for Dynamically Reconfigurable FPGA DesignsabstractThis paper gives a survey of communication architectures which allow for dynamically exchangeable hardware modules. Four different architectures are compared in terms of reconfiguration capabilities, performance, flexibility and hardware requirements. A set of parameters for the classification of the different communication architectures is presented and the pro and cons of each architecture are elaborated. The analysis takes a minimal communication system for connecting four hardware modules as a common basis for the comparison of the diverse data given in the papers on the different architectures. Thilo Pionteck, Carsten Albrecht, Roman Koch, Erik Maehle, Michael Hübner 0001, Jürgen Becker 0001 |
IPDPS | 1 |
| 2006 | A dynamically reconfigurable packet-switched network-on-chipabstractThis paper presents the design of an adaptable NoC for FPGA based dynamically reconfigurable SoCs. At runtime, switches can be added or removed from the network, allowing to adapt the NoC to the number, size and location of currently configured hardware modules. By using dynamic routing tables, reconfiguration can be done without stopping or stalling the NoC. The proposed architecture avoids the limitations of bus-based interconnection schemes which are often applied in partially dynamically reconfigurable FPGA designs Thilo Pionteck, Carsten Albrecht, Roman Koch |
DATE | 1 |
| 2006 | Applying Partial Reconfiguration to Networks-On-ChipsabstractThis paper presents CoNoChi, an adaptable network-on-chip for dynamically reconfigurable hardware designs. CoNoChi is designed for taking advantage of the partial dynamic reconfiguration capabilities of modern FPGAs and applies this feature to adapt the network structure to the location, number and size of currently configured hardware modules. The network consists of the minimal number of switches required. Switches can be added or removed from the network by a global control instance at runtime. Compared to common fixed network-on-chip structures, the CoNoChi architecture reduces the area requirements and latency of the network and eases the online placement of hardware modules. Two variants of CoNoChi are presented: one is based on a homogeneous hardware structure that is dynamically reconfigurable on logic block level, and the other one is adapted to the limited partial reconfiguration capabilities of Xilinx Virtex-II (Pro) FPGAs Thilo Pionteck, Roman Koch, Carsten Albrecht |
FPL | 1 |
| 2006 | An adaptive system-on-chip for network applicationsabstractThis paper presents the hardware architecture of DynaCORE, a dynamically reconfigurable system-on-chip for network applications. DynaCORE is an application specific coprocessor for offloading computationally intensive tasks from a network processor. The system-on-chip architecture is based on an adaptable network-on-chip which allows the dynamic replacement of hardware modules as well as the adaptation of the on-chip communication structure. The coprocessor leverages the active partial reconfiguration feature of modern FPGAs in order to adapt to shifting demand patterns. An embedded general-purpose processor core within the coprocessor runs software which manages the configurations of the device. With reference to a prototypical implementation targeting a Xilinx Virtex-II Pro FPGA, this paper focuses on on-chip communication issues. Topics include the integration of PowerPC processor cores into the configurable logic as well as the mode of operation of the network-on-chip Roman Koch, Thilo Pionteck, Carsten Albrecht, Erik Maehle |
IPDPS | 2 |
| 2005 | On The Design of A Dynamically Reconfigurable Function-Unit for Error Detection and Correction
Thilo Pionteck, Thomas Stiefmeier, Thorsten Staake, Manfred Glesner |
VLSI-SoC | 1 |
| 2004 | On the design of a function-specific reconfigurable: hardware accelerator for the MAC-layer in WLANsabstractThis work presents the hardware design of a dynamically reconfigurable function unit (RFU) to accelerate computation-intensive tasks in Medium Access Control (MAC) layers of WLANs. The function unit is integrated in a pipelined 32 bit RISC processor and provides full hardware support for the Advanced Encryption Standard (AES) as specified in upcoming WLAN standards such as IEEE 802.11i. Dynamic reconfiguration allows the processor to use arithmetic components and memory elements of the RFU not only for AES, but also for additional tasks common in the MAC-layer. With our approach it is possible to accelerate Reed-Solomon-Code generation, Cyclic Redundancy Checks as well as other encryption standards like SQUARE, Magenta and Twofish by supporting Galois Field multiplication and table look-ups. The integration of the reconfigurable unit in the processor core results in an architecture that can simultaneously support control-flow and data-flow oriented tasks. This architecture was prototyped onto a Virtex2 FPGA. Thilo Pionteck, Thorsten Staake, Thomas Stiefmeier, Lukusa D. Kabulepa, Manfred Glesner |
FPGA | 1 |
| 2004 | A Dynamically Reconfigurable Function-Unit for Error Detection and Correction in Mobile Terminals
Thilo Pionteck, Thomas Stiefmeier, Thorsten Staake, Manfred Glesner |
FPL | 1 |
| 2003 | Reconfiguration requirements for high speed wireless communication systemsabstractThis paper focuses on the reconfiguration requirements of hardware platforms for high speed wireless communication systems. Due to the underlying trade-off between flexibility and efficiency, many reconfigurable hardware solutions and FPGA implementations are prone to significant energy and performance penalties in comparison to application specific hardware designs. These penalties can only be alleviated by designing reconfigurable architectures for selected applications fields, since each application field has only a limited set of flexibility requirements. In this paper the analysis is conducted for emerging wireless communication standards based on the OFDM (Orthogonal Frequency Division Multiplexing) or CDMA (Code Division Multiple Access) transmission technique. The requirements for the physical and the medium access layers are analyzed separately. Focus is also set on different market segments. Thilo Pionteck, Lukusa D. Kabulepa, Clemens Schlachta, Manfred Glesner |
FPT | 1 |
| 2003 | Exploring the Capabilities of Reconfigurable Hardware for OFDM-based WLANs
Thilo Pionteck, Lukusa D. Kabulepa, Manfred Glesner |
VLSI-SOC | 1 |
| 2002 | A Framework for Teaching (Re)Configurable Architectures in Student Projects
Thilo Pionteck, Peter Zipf, Lukusa D. Kabulepa, Manfred Glesner |
FPL | 1 |
| 2001 | Efficient Mapping of Pre-synthesized IP-Cores onto Dynamically Reconfigurable Array Architectures
Jürgen Becker 0001, Nicolas Liebau, Thilo Pionteck, Manfred Glesner |
FPL | 3 |