EDBT 2026 Demo / reviewers in the wild / expert
Mohamed Assem Ibrahim
dblp:252/8057
· DBLP profile ↗
8ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0002-4129-0310ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 51% Memory systems · 21% Hardware accelerators and domain-specific architectures · 16% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.5 | 1 | 2021 | Analyzing and Leveraging Decoupled L1 Caches in GPUs · HPCA 2021 |
GPUs and heterogeneous computing › GPU cache
GPU cache hierarchy |
0.5 | 1 | 2021 | Analyzing and Leveraging Decoupled L1 Caches in GPUs · HPCA 2021 |
Hardware accelerators and domain-specific architectures › pattern matching accelerator
automata processor |
0.3 | 1 | 2018 | Architectural Support for Efficient Large-Scale Automata Processing · MICRO 2018 |
GPUs and heterogeneous computing
GPU resource management |
0.3 | 1 | 2018 | Efficient and Fair Multi-programming in GPUs via Effective Bandwidth Management · HPCA 2018 |
Memory systems
memory bandwidth management |
0.3 | 1 | 2018 | Efficient and Fair Multi-programming in GPUs via Effective Bandwidth Management · HPCA 2018 |
Hardware accelerators and domain-specific architectures
spatial architecture |
0.3 | 1 | 2018 | Architectural Support for Efficient Large-Scale Automata Processing · MICRO 2018 |
GPUs and heterogeneous computing › GPU programming
dynamic parallelism |
0.3 | 1 | 2017 | Controlled Kernel Launch for Dynamic Parallelism in GPUs · HPCA 2017 |
GPUs and heterogeneous computing
GPU runtime systems |
0.3 | 1 | 2017 | Controlled Kernel Launch for Dynamic Parallelism in GPUs · HPCA 2017 |
GPUs and heterogeneous computing
GPU scheduling |
0.3 | 1 | 2017 | Controlled Kernel Launch for Dynamic Parallelism in GPUs · HPCA 2017 |
Parallel and multicore computing › task scheduling
kernel scheduling |
0.3 | 1 | 2017 | Controlled Kernel Launch for Dynamic Parallelism in GPUs · HPCA 2017 |
Reconfigurable computing and FPGAs
automata processing |
0.1 | 1 | 2018 | Architectural Support for Efficient Large-Scale Automata Processing · MICRO 2018 |
Methods — techniques the papers use, named apart from their topics
architectural simulation · 0.5profiling-based state prediction · 0.3pattern-based TLP management · 0.3effective bandwidth metric · 0.3runtime framework · 0.3kernel mixing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FinGraV: Methodology for Fine-Grain GPU Power Visibility and InsightsabstractUbiquity of AI makes optimizing GPU power a priority as large GPU-based clusters are often employed to train and serve AI models. An important first step in optimizing GPU power consumption is high-fidelity and fine-grain power measurement of key AI computations on GPUs. To this end, we observe that as GPUs get more powerful, the resulting sub-millisecond to millisecond executions make fine-grain power analysis challenging. In this work, we first carefully identify the challenges in obtaining fine-grain GPU power profiles. To address these challenges, we devise FinGraV methodology where we employ execution time binning, careful CPU-GPU time synchronization, and power profile differentiation to collect finegrain GPU power profiles across prominent AI computations and across a spectrum of scenarios. Using the said FinGraV power profiles, we provide both, guidance on accurate power measurement and, in-depth view of power consumption on state-of-the-art AMD Instinct™ MI300X. For the former, we highlight a methodology for power differentiation across executions. For the latter, we make several observations pertaining to GPU subcomponent power consumption and GPU power proportionality across different scenarios. We believe that FinGraV unlocks both an accurate and a deeper view of power consumption of GPUs and opens up avenues for power optimization of these ubiquitous accelerators. Varsha Singhania, Shaizeen Aga, Mohamed Assem Ibrahim |
ISPASS | 3 |
| 2021 | Analyzing and Leveraging Decoupled L1 Caches in GPUsabstractGraphics Processing Units (GPUs) use caches to provide on-chip bandwidth as a way to address the memory wall. However, they are not always efficiently utilized for optimal GPU performance. We find that the main source of this inefficiency stems from the tightly-coupled design of cores with L1 caches. First, such a design assumes a per-core private local L1 cache in which each core independently caches the required data. This allows the same cache line to get replicated across cores, which wastes precious cache capacity. Second, due to the many-to-few traffic pattern, the tightly-coupled design leads to low per-core L1 bandwidth utilization while L2/memory is heavily utilized.To address these inefficiencies, we renovate the conventional GPU cache hierarchy by proposing a new DC-L1 (DeCoupled-L1) cache - an L1 cache separated from the GPU core. We show how decoupling the L1 cache from the GPU core provides opportunities to reduce data replication across the L1s and increase their bandwidth utilization. Specifically, we investigate how to aggregate the DC-L1s; how to manage data placement across the aggregated DC-L1s; and how to efficiently connect the DC-L1s to the GPU cores and the L2/memory partitions. Our evaluation shows that our new cache design boosts the useful L1 cache bandwidth and achieves significant improvement in performance and energy efficiency across a wide set of GPGPU applications while reducing the overall NoC area footprint. Mohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H. Loh, Adwait Jog |
HPCA | 1 |
| 2020 | Analyzing and Leveraging Shared L1 Caches in GPUsabstractGraphics Processing Units (GPUs) concurrently execute thousands of threads, which makes them effective for achieving high throughput for a wide range of applications. However, the memory wall often limits peak throughput. GPUs use caches to address this limitation, and hence several prior works have focused on improving cache hit rates, which in turn can improve throughput for memory-intensive applications. However, almost all of the prior works assume a conventional cache hierarchy where each GPU core has a private local L1 cache and all cores share the L2 cache. Our analysis shows that this canonical organization does not allow optimal utilization of caches because the private nature of L1 caches allows multiple copies of the same cache line to get replicated across cores. Mohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H. Loh, Adwait Jog |
PACT | 1 |
| 2019 | Analyzing and Leveraging Remote-Core Bandwidth for Enhanced Performance in GPUsabstractBandwidth achieved from local/shared caches and memory is a major performance determinant in Graphics Processing Units (GPUs). These existing sources of bandwidth are often not enough for optimal GPU performance. Therefore, to enhance the performance further, we focus on efficiently unlocking an additional potential source of bandwidth, which we call as remote-core bandwidth. The source of this bandwidth is based on the observation that a fraction of data (i.e., L1 read misses) required by one GPU core can also be found in the local (L1) caches of other GPU cores. In this paper, we propose to efficiently coordinate the data movement across cores in GPUs to exploit this remote-core bandwidth. However, we find that its efficient detection and utilization presents several challenges. To this end, we specifically address: a) which data is shared across cores, b) which cores have the shared data, and c) how we can get the data as soon as possible. Our extensive evaluation across a wide set of GPGPU applications shows that significant performance improvement can be achieved at a modest hardware cost on account of the additional bandwidth received from the remote cores. Mohamed Assem Ibrahim, Hongyuan Liu 0002, Onur Kayiran, Adwait Jog |
PACT | 1 |
| 2019 | Address-stride assisted approximate load value prediction in GPUsabstractValue prediction holds the promise of significantly improving the performance and energy efficiency. However, if the values are predicted incorrectly, significant performance overheads are observed due to execution rollbacks. To address these overheads, value approximation is introduced, which leverages the observation that the rollbacks are not necessary as long as the application-level loss in quality due to value misprediction is acceptable to the user. However, in the context of Graphics Processing Units (GPUs), our evaluations show that the existing approximate value predictors are not optimal in improving the prediction accuracy as they do not consider memory request order, a key characteristic in determining the accuracy of value prediction. As a result, the overall data movement reduction benefits are capped as it is necessary to limit the percentage of predicted values (i.e., prediction coverage) for an acceptable value of application-level error. Mohamed Assem Ibrahim, Sparsh Mittal, Adwait Jog |
ICS | 2 |
| 2018 | Efficient and Fair Multi-programming in GPUs via Effective Bandwidth ManagementabstractManaging the thread-level parallelism (TLP) of GPGPU applications by limiting it to a certain degree is known to be effective in improving the overall performance. However, we find that such prior techniques can lead to sub-optimal system throughput and fairness when two or more applications are co-scheduled on the same GPU. It is because they attempt to maximize the performance of individual applications in isolation, ultimately allowing each application to take a disproportionate amount of shared resources. This leads to high contention in shared cache and memory. To address this problem, we propose new application-aware TLP management techniques for a multi-application execution environment such that all co-scheduled applications can make good and judicious use of all the shared resources. For measuring such use, we propose an application-level utility metric, called effective bandwidth, which accounts for two runtime metrics: attained DRAM bandwidth and cache miss rates. We find that maximizing the total effective bandwidth and doing so in a balanced fashion across all co-located applications can significantly improve the system throughput and fairness. Instead of exhaustively searching across all the different combinations of TLP configurations that achieve these goals, we find that a significant amount of overhead can be reduced by taking advantage of the trends, which we call patterns, in the way application's effective bandwidth changes with different TLP combinations. Our proposed pattern-based TLP management mechanisms improve the system throughput and fairness by 20% and 2x, respectively, over a baseline where each application executes with a TLP configuration that provides the best performance when it executes alone. Fan Luo 0002, Mohamed Assem Ibrahim, Onur Kayiran, Adwait Jog |
HPCA | 3 |
| 2018 | Architectural Support for Efficient Large-Scale Automata ProcessingabstractThe Automata Processor (AP) accelerates applications from domains ranging from machine learning to genomics. However, as a spatial architecture, it is unable to handle larger automata programs without repeated reconfiguration and re-execution. To achieve high throughput, this paper proposes for the first time architectural support for AP to efficiently execute large-scale applications. We find that a large number of existing and new Non-deterministic Finite Automata (NFA) based applications have states that are never enabled but are still configured on the AP chips leading to their underutilization. With the help of careful characterization and profiling-based mechanisms, we predict which states are never enabled and hence need not be configured on AP. Furthermore, we develop SparseAP, a new execution mode for AP to efficiently handle the mis-predicted NFA states. Our detailed simulations across 26 applications from various domains show that our newly proposed execution model for AP can obtain 2.1x geometric mean speedup (up to 47x) over the baseline AP execution. Hongyuan Liu 0002, Mohamed Assem Ibrahim, Onur Kayiran, Sreepathi Pai, Adwait Jog |
MICRO | 2 |
| 2017 | Controlled Kernel Launch for Dynamic Parallelism in GPUsabstractDynamic parallelism (DP) is a promising feature for GPUs, which allows on-demand spawning of kernels on the GPU without any CPU intervention. However, this feature has two major drawbacks. First, the launching of GPU kernels can incur significant performance penalties. Second, dynamically-generated kernels are not always able to efficiently utilize the GPU cores due to hardware-limits. To address these two concerns cohesively, we propose SPAWN, a runtime framework that controls the dynamically-generated kernels, thereby directly reducing the associated launch overheads and queuing latency. Moreover, it allows a better mix of dynamically-generated and original (parent) kernels for the scheduler to effectively hide the remaining overheads and improve the utilization of the GPU resources. Our results show that, across 13 benchmarks, SPAWN achieves 69% and 57% speedup over the flat (non-DP) implementation and baseline DP, respectively. Xulong Tang, Ashutosh Pattnaik, Huaipan Jiang, Onur Kayiran, Adwait Jog, Sreepathi Pai, Mohamed Assem Ibrahim, Mahmut T. Kandemir, Chita R. Das |
HPCA | 7 |