Beat Weiss

dblp:31/5707 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-4069-1286ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Magnetic Tape Storage Technology
abstract
Magnetic tape provides a cost-effective way to retain the exponentially increasing volumes of data being created in recent years. The low cost per terabyte combined with tape’s low energy consumption make it an appealing option for storing infrequently accessed data and has resulted in a resurgence in use of the technology. Magnetic tape as a digital data storage technology was first commercialized in the early 1950’s and has evolved continuously since then. Despite its long history, tape has significant potential for continued capacity and data rate scaling. This article strives to provide an overview of linear magnetic tape technology, usage, history, and future outlook. After a short introduction, the article delves into the details of how modern tape drives and media operate, including the basic mechanism and physics of magnetic recording, current tape media technology, state-of-the-art tape head technology, tape layout and encoding, data retrieval, timing-based servo and mechatronics of a tape drive, and the capabilities of current drives. This is followed by a discussion of tape libraries, an overview of tape library performance modeling research, operating system-level and application-level tape support, tape use cases, and the future scaling potential and outlook of tape. The article concludes with a history of tape hardware, media, usage and software.
Mark A. Lantz, Simeon Furrer, Martin Petermann, Hugo E. Rothuizen, Stella Brach, Luzius Kronig, Ilias Iliadis, Beat Weiss, Edwin R. Childers, David Pease
ACM Trans. Storage8
2024 A System Development Kit for Big Data Applications on FPGA-based Clusters: The EVEREST Approach
abstract
Modern big data workflows are characterized by computationally intensive kernels. The simulated results are often combined with knowledge extracted from AI models to ultimately support decision-making. These energy-hungry workflows are increasingly executed in data centers with energy-efficient hard-ware accelerators since FPG As are well-suited for this task due to their inherent parallelism. We present the H2020 project EVEREST, which has developed a system development kit (SDK) to simplify the creation of FPGA-accelerated kernels and manage the execution at runtime through a virtualization environment. This paper describes the main components of the EVEREST SDK and the benefits that can be achieved in our use cases.
Christian Pilato, Subhadeep Banik, Jakub Beránek, Fabien Brocheton, Jerónimo Castrillón, Riccardo Cevasco, Radim Cmar, Serena Curzel, Fabrizio Ferrandi, Karl F. A. Friebel, Antonella Galizia, Matteo Grasso, Paulo Silva 0002, Jan Martinovic, Gianluca Palermo, Michele Paolino, Andrea Parodi, Antonio Parodi, Fabio Pintus, Raphael Polig, David Poulet, Francesco Regazzoni 0001, Burkhard Ringlein, Roberto Rocco, Katerina Slaninová, Tom Slooff, Stephanie Soldavini, Felix Suchert, Mattia Tibaldi, Beat Weiss, Christoph Hagleitner
DATE30
2024 Automated parallel execution of distributed task graphs with FPGA clusters
abstract
Over the years, Field Programmable Gate Arrays (FPGA) have been gaining popularity in the High Performance Computing (HPC) field, because their reconfigurability enables very fine-grained optimizations with low energy cost. However, the different characteristics, architectures, and network topologies of the clusters have hindered the use of FPGAs at a large scale. In this work, we present an evolution of OmpSs@FPGA, a high-level task-based programming model and extension to OmpSs-2, that aims at unifying all FPGA clusters by using a message-passing interface that is compatible with FPGA accelerators. These accelerators are programmed with C/C++ pragmas, and synthesized with High-Level Synthesis tools. The new framework includes a custom protocol to exchange messages between FPGAs, agnostic of the architecture and network type. On top of that, we present a new communication paradigm called Implicit Message Passing (IMP), where the user does not need to call any message-passing API. Instead, the runtime automatically infers data movement between nodes. We test classic message passing and IMP with three benchmarks on two different FPGA clusters. One is cloudFPGA, a disaggregated platform with AMD FPGAs that are only connected to the network through UDP/TCP/IP. The other is ESSPER, composed of CPU-attached Intel FPGAs that have a private network at the ethernet level. In both cases, we demonstrate that IMP with OmpSs@FPGA can increase the productivity of FPGA programmers at a large scale thanks to simplifying communication between nodes, without limiting the scalability of applications. We implement the N-body, Heat simulation and Cholesky decomposition benchmarks, and show that FPGA clusters get 2.6x and 2.4x better performance per watt than a CPU-only supercomputer for N-body and Heat.
Juan Miguel De Haro Ruiz, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Tomohiro Ueno, Kentaro Sano, Burkhard Ringlein, François Abel, Beat Weiss
Future Gener. Comput. Syst.9
2023 Composability of Cloud Accelerators in Virtual World Simulations
abstract
Immersive and interactive experiences offered by virtual world simulations (VWS) are becoming increasingly indispensable in today's digital era, particularly in domains such as gaming, engineering, and education. Cloud computing provides a compelling solution for VWS, with its scalability, costeffectiveness, and accessibility, combined with efficient processing capabilities of hardware accelerators such as FPGAs and GPUs. This paper presents APEIRON, an in-progress distributed library built on the Ray framework, that aims to facilitate the effective composition of hardware accelerators on public clouds and emerging disaggregated platforms. Additionally, this work introduces a novel composability metric, analyzing the intricate relationship between resource utilization and workload performance. This study focuses on theoretical composability analyses, laying groundwork for optimized utilization of hardware resources to enhance efficiency, thereby potentially reducing operational costs.
Dionysios Diamantopoulos, Burkhard Ringlein, Beat Weiss, Mark A. Lantz, François Abel
CLOUD3
2022 OmpSs@cloudFPGA: An FPGA Task-Based Programming Model with Message Passing
abstract
Nowadays, a new parallel paradigm for energy-efficient heterogeneous hardware infrastructures is required to achieve better performance at a reasonable cost on high-performance computing applications. Under this new paradigm, some application parts are offloaded to specialized accelerators that run faster or are more energy-efficient than CPUs. Field-Programmable Gate Arrays (FPGA) are one of those types of accelerators that are becoming widely available in data centers. This paper proposes OmpSs@cloudFPGA, which includes novel extensions to parallel task-based programming models that enable easy and efficient programming of heterogeneous clusters with FPGAs. The programmer only needs to annotate, with OpenMP-like pragmas, the tasks of the application that should be accelerated in the cluster of FPGAs. Next, the proposed programming model framework automatically extracts parts annotated with High-Level Synthesis (HLS) pragmas and synthesizes them into hardware accelerator cores for FPGAs. Additionally, our extensions include and support two novel features: 1) FPGA-to-FPGA direct communication using a Message Passing Interface (MPI) similar Application Programming Interface (API) with one-to-one and collective communications to alleviate host communication channel bottleneck, and 2) creating and spawning work from inside the FPGAs to their own accelerator cores based on an MPI rank-like identification. These features break the classical host-accelerator model, where the host (typically the CPU) generates all the work and distributes it to each accelerator. We also present an evaluation of OmpSs@cloudFPGA for different parallel strategies of the N-Body application on the IBM cloudFPGA research platform. Results show that for cluster sizes up to 56 FPGAs, the performance scales linearly. To the best of our knowledge, this is the best performance obtained for N-body over FPGA platforms, reaching 344 Gpairs/s with 56 FPGAs. Finally, we compare the performance and power consumption of the proposed approach with the ones obtained by a classical execution on the MareNostrum 4 supercomputer, demonstrating that our FPGA approach reduces power consumption by an order of magnitude.
Juan Miguel De Haro Ruiz, Rubén Cano, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, François Abel, Burkhard Ringlein, Beat Weiss
IPDPS10
2021 Acceleration-as-a-µService: A Cloud-native Monte-Carlo Option Pricing Engine on CPUs, GPUs and Disaggregated FPGAs
abstract
The evolution of cloud applications into loosely-coupled microservices opens new opportunities for hardware accelerators to improve workload performance. Existing accelerator techniques for cloud sacrifice the consolidation benefits of microservices. This paper presents CloudiFi, a framework to deploy and compare accelerators as a cloud service. We evaluate our framework in the context of a financial workload and present early results indicating up to 485 x gains in microservice response time.
Dionysios Diamantopoulos, Raphael Polig, Burkhard Ringlein, Mitra Purandare, Beat Weiss, Christoph Hagleitner, Mark A. Lantz, François Abel
CLOUD5
2021 A Case for Function-as-a-Service with Disaggregated FPGAs
abstract
The slowdown of Moore's law and the end of Dennard scaling created a demand for specialized accelerators, including Field Programmable Gate Arrays (FPGAs), in cloud data centers. At the same time, compute resources are increasingly consumed via public and private clouds and traditional applications are modernized using scalable microservices and Function-as-a-Service (FaaS) offerings. Nonetheless, true FaaS based on FPGAs or other accelerators is virtually absent from the offering catalogs of all major cloud providers. In addition, FPGA applications are typically coded in a monolithic fashion, due to device and vendor specific dependencies, which reduces the portability and usability of FPGA cloud offerings further. However, FPGA-based FaaS can improve execution efficiency and minimize (tail-) latencies while decreasing costs. We propose a novel system architecture, called Mantle, that uses disaggregated FPGAs to enable scalable, usable, portable and efficient FaaS offerings for FPGAs. Our experimental results demonstrate a significant reduction of end-to-end service provisioning time to below 7 seconds and an increase in execution efficiency by a factor of 4 with negligible overhead.
Burkhard Ringlein, François Abel, Dionysios Diamantopoulos, Beat Weiss, Christoph Hagleitner, Marc Reichenbach, Dietmar Fey
CLOUD4
2020 ZRLMPI: A Unified Programming Model for Reconfigurable Heterogeneous Computing Clusters
abstract
Over the past two decades, the Message Passing Interface (MPI) has evolved as the de-facto standard for programming High-Performance Computing (HPC) clusters. Its widespread utilization led to the rapid development of applications and high reusability. Meanwhile, energy- and compute-efficient devices such as Field-Programmable Gate Arrays (FPGAs) are stepping into modern data centers and HPC clusters to address the nearing end of technology scaling. This combination of traditional CPU servers and FPGA nodes leads to Reconfigurable Heterogeneous HPC (ReH2PC) systems that are particularly cumbersome to program because of the absence of a standard programming model. This work advocates the use of MPI to program such ReH2PC clusters and presents a proof of concept based on a cross-compiler, a High-Level Synthesis library, C++ library, an FPGA- and a CPU-runtime environment. The result is a one-click solution, which compiles a standard MPI application for a ReH2PC cluster.
Burkhard Ringlein, François Abel, Alexander Ditter, Beat Weiss, Christoph Hagleitner, Dietmar Fey
FCCM4
2019 System Architecture for Network-Attached FPGAs in the Cloud using Partial Reconfiguration
abstract
Emerging applications such as deep neural networks, bioinformatics or video encoding impose a high computing pressure on the Cloud. Reconfigurable technologies like Field-Programmable Gate Arrays (FPGAs) can handle such compute-intensive workloads in an efficient and performant way. To seamlessly incorporate FPGAs into existing Cloud environments and leverage their full power efficiency, FPGAs should be directly attached to the data center network and operate independent of power-hungry CPUs. This raises new questions about resource management, application deployment and network integrity. We present a system architecture for managing a large number of network-attached FPGAs in an efficient, flexible and scalable way. To ensure the integrity of the infrastructure, we use partial reconfiguration to separate the non-privileged user logic from the privileged system logic. To create a really scalable and agile cloud service, the management of all resources builds on the Representational State Transfer (REST) concept.
Burkhard Ringlein, François Abel, Alexander Ditter, Beat Weiss, Christoph Hagleitner, Dietmar Fey
FPL4
2013 A case for centrally controlled wireless sensor networks
Urs Hunkeler, Clemens Lombriser, Hong Linh Truong 0002, Beat Weiss
Comput. Networks4
2011 A power-efficient wireless sensor network for continuously monitoring seismic vibrations
abstract
We present a novel power-efficient wireless sensor network for continuously monitoring and analyzing seismic vibrations with sensor nodes and forwarding the retrieved information with low-cost relay nodes to backend applications. The applied vibration sensing algorithms are derived from the DIN 4150-3 standard. All nodes in the network are battery-powered and equipped with an IEEE 802.15.4 compatible radio transceiver. The nodes communicate with each other by executing a novel power-efficient protocol stack, which provides all network functions required by the vibration-sensing application and uses a publish/subscribe messaging protocol for communicating between the network nodes and the backend applications. Results obtained in certification and field tests show that the proposed vibration-sensing solution is standard-compliant, and that the wireless vibration sensor network (WVSN) exhibits excellent performance in terms of packet delivery rate, latency, and power efficiency.
Beat Weiss, Hong Linh Truong 0002, Wolfgang Schott, Andrea Munari, Clemens Lombriser, Urs Hunkeler, Pierre R. Chevillat
SECON1