EDBT 2026 Demo / reviewers in the wild / expert
Sohan Lal
dblp:133/9924
· DBLP profile ↗
12ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-2325-1705ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Portable Compiler-Runtime Approach for Scalability PredictionabstractHighly scalable parallel applications can efficiently solve expensive computational problems when run on a large number of compute nodes. However, selecting the optimal number of nodes for a compute job of a given size is non-trivial, and allocating too few or too many nodes may not yield the expected performance. Knowing the scaling behavior of an application in advance enables us, for example, to make optimal use of the available hardware resources. We introduce a novel, portable approach to predict the scalability of parallel applications written in modern high-level programming models. We propose a predictive compiler-runtime framework based on Celerity, a task-based distributed runtime system that enables executing SYCL codes on clusters. The framework targets a broad range of computing systems, from CPU to GPU clusters, and proposes a model that combines machine learning, communication modeling and DAG heuristics. Experimental results on two large-scale clusters, JUWELS and Marconi-100, show accurate scalability prediction of unseen single and multi-task applications. Nicolai Stawinoga, Sohan Lal, Biagio Cosenza, Philip Salzmann, Peter Thoman, Thomas Fahringer |
Future Gener. Comput. Syst. | 2 |
| 2025 | Enhancing Practicality of Memory Compression for GPUs with High-Throughput Simplifications
Manuel Renz, Sohan Lal |
CF | 2 |
| 2025 | Robust LFSR-based Scrambling to Mitigate Stencil Attack on Main MemoryabstractMain memory plays a pivotal role in the storage of computational data in a wide range of applications, including highly sensitive assets such as banking transactions, cryptographic keys, and user credentials. However, memory systems remain vulnerable to advanced physical and side-channel attacks, including cold boot attacks that exploit residual data after power-down. To mitigate such risks, Intel’s DDR3 memory scrambler uses a Linear Feedback Shift Register (LFSR)-based stream cipher to obscure memory contents. Nevertheless, this mechanism has been shown to be susceptible to stencil attack, a cold boot technique that reconstructs the scrambling key by leveraging the linear and periodic nature of the keystream. This article proposes a novel, lightweight, and secure scrambling architecture based on a generic LFSR designed to enhance the security of DDR3 memory against cold boot attacks. The proposed generic LFSR-based mechanism eliminates differential keystream periodicity by introducing an address- and seed-dependent LFSR structure, thereby rendering differential key recovery techniques computationally infeasible. Furthermore, unlike traditional AES-based memory encryption that incurs high latency and area overhead, the proposed approach achieves comparable security guarantees with low hardware complexity and zero access latency. The hardware implementation results on the Xilinx VCU118 FPGA show that the proposed scheme consumes only 252 LUTs, 256 registers and 104 slices, comparable to the Intel DDR3 scrambler, while offering superior resilience against the cold boot, warm boot, and probing attacks. These results demonstrate the practicality of the proposed scheme for secure memory systems in resource-constrained environments. Gaurav Kumar 0001, Kushal Pravin Nanote, Sohan Lal, Yamuna Prasad, Satyadev Ahlawat |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2023 | Beyond Compression Ratio: A Throughput Analysis of Memory Compression Techniques for GPUsabstractMemory compression is increasingly used as a technique to synthetically increase the off-chip memory bandwidth of GPUs by transferring data in a compressed format between on-chip and off-chip memory. The increased memory bandwidth results in a speedup for the bandwidth-limited applications. State-of-the-art memory compression techniques often target a high compression ratio, however, a high compression ratio alone is not sufficient for full integration into throughput-oriented GPUs. To deploy a memory compression technique, the throughput of the compression technique has to keep up with the bandwidth of the off-chip memory of GPUs. Unfortunately, the throughput of the state-of-the-art memory compression techniques is often not discussed in detail and mostly the emphasis is only placed on the achieved compression ratio. In this work, we present a throughput analysis of several state-of-the art memory compression techniques and study their capability to match modern GPUs memory bandwidth. We implement several memory compression techniques in hardware and synthesize designs using Synopsys Design Compiler with 14 nm ASIC libraries to analyze the throughput, area, and power consumption. Our analysis shows that simple compression techniques that have moderate compression ratios but higher throughput are more suitable for practical implementation in GPUs considering the area (up to 11.8× lower), power consumption (up to 7× lower), and the complexity of implementation. Manuel Renz, Sohan Lal |
ICCD | 2 |
| 2022 | Memory Access Granularity Aware Lossless Compression for GPUsabstractHigh-bandwidth off-chip memory has played a key role in the success of Graphics Processing Units (GPUs) as an accelerator. However, as memory bandwidth scaling continues to lag behind the computational power, it remains a key bottleneck in computing systems. While memory compression has shown immense potential to increase the effective memory bandwidth by compressed data transfers between on-chip and off-chip memory, the large memory access granularity (MAG) of off-chip memory limits compression techniques from achieving a high effective compression ratio. Unfortunately, state-of-the-art lossless memory compression techniques do not take the large MAG of off-chip memory into account. A recent study has used MAG-aware approximation to increase the effective compression ratio, however, not all applications can tolerate errors, which limits its applicability. We propose extensions and GPU-specific optimizations to adapt a lossless memory compression technique to a MAG size to increase the effective compression ratio and performance gain. Our technique is based on the well-known Base-Delta-Immediate (BDI) compression technique that compresses a memory block to a common base and multiple deltas. We leverage the key observation that deltas often contain enough leading zeros to compress a block to a multiple of MAG without any loss of information. We show that MAG-aware BDI provides, on average, 48 % higher effective compression ratio, 10% (up to 27%) higher speedup, and 16% bandwidth reduction compared to normal BDI. While BDI, FPC, and CPACK have a similar compression ratio, MAG-aware BDI outperforms FPC, CPACK, and SLC by 56%, 47%, and 33%, respectively. Sohan Lal, Manuel Renz, Julian Hartmer, Ben H. H. Juurlink |
IPDPS | 1 |
| 2021 | QSLC: Quantization-Based, Low-Error Selective Approximation for GPUsabstractGPUs use a large memory access granularity (MAG) that often results in a low effective compression ratio for memory compression techniques. The low effective compression ratio is caused by a significant fraction of compressed blocks that have a few bytes above a multiple of MAG. While MAG-aware selective approximation, based on a tree structure, has been used to increase the effective compression ratio and the performance gain, approximation results in a high error that is reduced by using complex optimizations. We propose a simple quantization-based approximation technique (QSLC) that can also selectively approximate a few bytes above MAG. While the quantization-based approximation technique has a similar performance to the state-of-the-art tree-based selective approximation, the average error for the quantization-based technique is 5× lower. We further trade-off the two techniques and show that the area and power overhead of the quantization-based technique is 12.1× and 7.6× lower than the state-of-the-art, respectively. Our sensitivity analysis to different block sizes further shows the opportunities and the significance of MAG-aware selective approximation. Sohan Lal, Jan Lucas, Ben H. H. Juurlink |
DATE | 1 |
| 2020 | SYCL-Bench: A Versatile Cross-Platform Benchmark Suite for Heterogeneous Computing
Sohan Lal, Aksel Alpay, Philip Salzmann, Biagio Cosenza, Alexander Hirsch, Nicolai Stawinoga, Peter Thoman, Thomas Fahringer, Vincent Heuveline |
Euro-Par | 1 |
| 2019 | SLC: Memory Access Granularity Aware Selective Lossy Compression for GPUsabstractMemory compression is a promising approach for reducing memory bandwidth requirements and increasing performance, however, memory compression techniques often result in a low effective compression ratio due to large memory access granularity (MAG) exhibited by GPUs. Our analysis of the distribution of compressed blocks shows that a significant percentage of blocks are compressed to a size that is only a few bytes above a multiple of MAG, but a whole burst is fetched from memory. These few extra bytes significantly reduce the compression ratio and the performance gain that otherwise could result from a higher raw compression ratio. To increase the effective compression ratio, we propose a novel MAG aware Selective Lossy Compression (SLC) technique for GPUs. The key idea of SLC is that when lossless compression yields a compressed size with few bytes above a multiple of MAG, we approximate these extra bytes such that the compressed size is a multiple of MAG. This way, SLC mostly retains the quality of a lossless compression and occasionally trades small accuracy for higher performance. We show a speedup of up to 35% normalized to a state-of-the-art lossless compression technique with a low loss in accuracy. Furthermore, average energy consumption and energy-delay-product are reduced by 8.3% and 17.5%, respectively. Sohan Lal, Jan Lucas, Ben H. H. Juurlink |
DATE | 1 |
| 2018 | Optimal DC/AC data bus inversion codingabstractGDDR5 and DDR4 memories use data bus inversion (DBI) coding to reduce termination power and decrease the number of output transitions. Two main strategies exist for encoding data using DBI: DBI DC minimizes the number of outputs transmitting a zero, while DBI AC minimizes the number of signal transitions. We show that neither of these strategies is optimal and reduction of interface power of up to 6% can be achieved by taking both the number of zeros and the number of signal transitions into account when encoding the data. We then demonstrate that a hardware implementation of optimal DBI coding is feasible, results in a reduction of system power and requires only an insignificant additional die area. Jan Lucas, Sohan Lal, Ben H. H. Juurlink |
DATE | 2 |
| 2017 | E^2MC: Entropy Encoding Based Memory Compression for GPUsabstractModern Graphics Processing Units (GPUs) provide much higher off-chip memory bandwidth than CPUs, but many GPU applications are still limited by memory bandwidth.Unfortunately, off-chip memory bandwidth is growing slower than the number of cores and has become a performance bottleneck.Thus, optimizations of effective memory bandwidth play a significant role for scaling the performance of GPUs.Memory compression is a promising approach for improving memory bandwidth which can translate into higher performance and energy efficiency.However, compression is not free and its challenges need to be addressed, otherwise the benefits of compression may be offset by its overhead.We propose an entropy encoding based memory compression (E 2 MC) technique for GPUs, which is based on the well-known Huffman encoding.We study the feasibility of entropy encoding for GPUs and show that it achieves higher compression ratios than state-of-the-art GPU compression techniques.Furthermore, we address the key challenges of probability estimation, choosing an appropriate symbol length for encoding, and decompression with low latency.The average compression ratio of E 2 MC is 53% higher than the state of the art.This translates into an average speedup of 20% compared to no compression and 8% higher compared to the state of the art.Energy consumption and energy-delayproduct are reduced by 13% and 27%, respectively.Moreover, the compression ratio achieved by E 2 MC is close to the optimal compression ratio given by Shannon's source coding theorem. Sohan Lal, Jan Lucas, Ben H. H. Juurlink |
IPDPS | 1 |
| 2016 | The neuro vector engine: Flexibility to improve convolutional net efficiency for wearable vision
Maurice Peemen, Runbin Shi, Sohan Lal, Ben H. H. Juurlink, Bart Mesman, Henk Corporaal |
DATE | 3 |
| 2013 | How a single chip causes massive power bills GPUSimPow: A GPGPU power simulatorabstractModern GPUs are true power houses in every meaning of the word: While they offer general-purpose (GPGPU) compute performance an order of magnitude higher than that of conventional CPUs, they have also been rapidly approaching the infamous “power wall”, as a single chip sometimes consumes more than 300W. Thus, the design space of GPGPU microarchitecture has been extended by another dimension: power. While GPU researchers have previously relied on cycle-accurate simulators for estimating performance during design cycles, there are no simulation tools that include power as well. To mitigate this issue, we introduce the GPUSimPow power estimation framework for GPGPUs consisting of both analytical and empirical models for regular and irregular hardware components. To validate this framework, we build a custom measurement setup to obtain power numbers from real graphics cards. An evaluation on a set of well-known benchmarks reveals an average relative error of 11.7% between simulated and hardware power for GT240 and an average relative error of 10.8% for GTX580. The simulator has been made available to the public [1]. Jan Lucas, Sohan Lal, Michael Andersch, Mauricio Alvarez-Mesa, Ben H. H. Juurlink |
ISPASS | 2 |