EDBT 2026 Demo / reviewers in the wild / expert
Manuel Renz
dblp:172/9966
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0003-8786-891XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Practicality of Memory Compression for GPUs with High-Throughput Simplifications
Manuel Renz, Sohan Lal |
CF | 1 |
| 2023 | Beyond Compression Ratio: A Throughput Analysis of Memory Compression Techniques for GPUsabstractMemory compression is increasingly used as a technique to synthetically increase the off-chip memory bandwidth of GPUs by transferring data in a compressed format between on-chip and off-chip memory. The increased memory bandwidth results in a speedup for the bandwidth-limited applications. State-of-the-art memory compression techniques often target a high compression ratio, however, a high compression ratio alone is not sufficient for full integration into throughput-oriented GPUs. To deploy a memory compression technique, the throughput of the compression technique has to keep up with the bandwidth of the off-chip memory of GPUs. Unfortunately, the throughput of the state-of-the-art memory compression techniques is often not discussed in detail and mostly the emphasis is only placed on the achieved compression ratio. In this work, we present a throughput analysis of several state-of-the art memory compression techniques and study their capability to match modern GPUs memory bandwidth. We implement several memory compression techniques in hardware and synthesize designs using Synopsys Design Compiler with 14 nm ASIC libraries to analyze the throughput, area, and power consumption. Our analysis shows that simple compression techniques that have moderate compression ratios but higher throughput are more suitable for practical implementation in GPUs considering the area (up to 11.8× lower), power consumption (up to 7× lower), and the complexity of implementation. Manuel Renz, Sohan Lal |
ICCD | 1 |
| 2022 | Memory Access Granularity Aware Lossless Compression for GPUsabstractHigh-bandwidth off-chip memory has played a key role in the success of Graphics Processing Units (GPUs) as an accelerator. However, as memory bandwidth scaling continues to lag behind the computational power, it remains a key bottleneck in computing systems. While memory compression has shown immense potential to increase the effective memory bandwidth by compressed data transfers between on-chip and off-chip memory, the large memory access granularity (MAG) of off-chip memory limits compression techniques from achieving a high effective compression ratio. Unfortunately, state-of-the-art lossless memory compression techniques do not take the large MAG of off-chip memory into account. A recent study has used MAG-aware approximation to increase the effective compression ratio, however, not all applications can tolerate errors, which limits its applicability. We propose extensions and GPU-specific optimizations to adapt a lossless memory compression technique to a MAG size to increase the effective compression ratio and performance gain. Our technique is based on the well-known Base-Delta-Immediate (BDI) compression technique that compresses a memory block to a common base and multiple deltas. We leverage the key observation that deltas often contain enough leading zeros to compress a block to a multiple of MAG without any loss of information. We show that MAG-aware BDI provides, on average, 48 % higher effective compression ratio, 10% (up to 27%) higher speedup, and 16% bandwidth reduction compared to normal BDI. While BDI, FPC, and CPACK have a similar compression ratio, MAG-aware BDI outperforms FPC, CPACK, and SLC by 56%, 47%, and 33%, respectively. Sohan Lal, Manuel Renz, Julian Hartmer, Ben H. H. Juurlink |
IPDPS | 2 |
| 2019 | Analyzing Efficient Stream Processing on Modern HardwareabstractModern Stream Processing Engines (SPEs) process large data volumes under tight latency constraints. Many SPEs execute processing pipelines using message passing on shared-nothing architectures and apply a partition-based scale-out strategy to handle high-velocity input streams. Furthermore, many state-of-the-art SPEs rely on a Java Virtual Machine to achieve platform independence and speed up system development by abstracting from the underlying hardware. In this paper, we show that taking the underlying hardware into account is essential to exploit modern hardware efficiently. To this end, we conduct an extensive experimental analysis of current SPEs and SPE design alternatives optimized for modern hardware. Our analysis highlights potential bottlenecks and reveals that state-of-the-art SPEs are not capable of fully exploiting current and emerging hardware trends, such as multi-core processors and high-speed networks. Based on our analysis, we describe a set of design changes to the common architecture of SPEs to scale-up on modern hardware. We show that the single-node throughput can be increased by up to two orders of magnitude compared to state-of-the-art SPEs by applying specialized code generation, fusing operators, batch-style parallelization strategies, and optimized windowing. This speedup allows for deploying typical streaming applications on a single or a few nodes instead of large clusters. Steffen Zeuch, Sebastian Breß, Tilmann Rabl, Bonaventura Del Monte, Jeyhun Karimov, Clemens Lutz, Manuel Renz, Jonas Traub, Volker Markl |
Proc. VLDB Endow. | 7 |