Ahmed Abousamra

dblp:45/7357 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
2since 2021 · last 2023
0000-0003-1768-9425ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 7 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Interconnection networks and networks-on-chip · 35% Memory systems · 33% Processor architecture and microarchitecture · 18%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 50% Program analysis · 50%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
chip multiprocessor
0.222012
Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Codesign of NoC and Cache Organization for Reducing Access Latency in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Interconnection networks and networks-on-chip › switching
circuit switching
0.212013
Proactive circuit allocation in multiplane NoCs · DAC 2013
Memory systems › cache
cache organization
0.112012
Codesign of NoC and Cache Organization for Reducing Access Latency in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Interconnection networks and networks-on-chip
communication locality
0.112012
Codesign of NoC and Cache Organization for Reducing Access Latency in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Storage systems
data placement
0.112012
Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Memory systems › cache design
non-uniform cache architecture
0.112012
Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Compilers and program optimization › compiler optimization
compiler-directed optimization
0.012012
Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Program analysis › memory analysis
memory access pattern analysis
0.012012
Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Memory systems
access latency reduction
0.012012
Codesign of NoC and Cache Organization for Reducing Access Latency in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Interconnection networks and networks-on-chip › network reconfiguration
network configuration
0.012012
Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012

Methods — techniques the papers use, named apart from their topics

data partitioning algorithm · 0.3compiler analysis · 0.3benchmark simulation · 0.2simulation · 0.1hybrid circuit/packet switching · 0.1
YearPublicationVenuePosition
2023 Fast Performance Analysis for NoCs With Weighted Round-Robin Arbitration and Finite Buffers
abstract
Weighted round-robin (WRR) arbitration provides global fairness in networks-on-chip (NoCs) as opposed to the commonly used round-robin and priority-based arbitration techniques. However, the large number of weights explodes the design space and exacerbates performance (latency-throughput) tuning. Therefore, fast and accurate performance analysis techniques for NoCs are crucial for accelerating design space exploration and accurate pre-silicon evaluation. This article presents the first comprehensive performance analysis technique for NoCs with WRR arbitration and finite buffers. It can handle bursty traffic and is scalable to large NoC sizes. The proposed technique first estimates the probability that a queue is full and uses this result to compute the modified service time and queuing delay. Thorough experimental evaluations with synthetic traffic and real applications show that the proposed analytical model is always more than 10% accurate compared to cycle-accurate simulations. Moreover, the proposed performance analysis technique is five orders of magnitude faster than cycle-accurate simulations for a$16\times16$mesh NoC.
Sumit K. Mandal, Shruti Yadav Narayana, Raid Ayoub, Michael Kishinevsky, Ahmed Abousamra, Ümit Y. Ogras
IEEE Trans. Very Large Scale Integr. Syst.5
2021 Theoretical Analysis and Evaluation of NoCs with Weighted Round-Robin Arbitration
abstract
Fast and accurate performance analysis techniques are essential in early design space exploration and pre-silicon evaluations, including software eco-system development. In particular, on-chip communication continues to play an increasingly important role as the many-core processors scale up. This paper presents the first performance analysis technique that targets networks-on-chip (NoCs) that employ weighted round-robin (WRR) arbitration. Besides fairness, WRR arbitration provides flexibility in allocating bandwidth proportionally to the importance of the traffic classes, unlike basic round-robin and priority-based arbitration. The proposed approach first estimates the effective service time of the packets in the queue due to WRR arbitration. Then, it uses the effective service time to compute the average waiting time of the packets. Next, we incorporate a decomposition technique to extend the analytical model to handle NoC of any size. The proposed approach achieves less than 5% error while executing real applications and 10% error under challenging synthetic traffic with different burstiness levels.
Sumit K. Mandal, Jie Tong, Raid Ayoub, Michael Kishinevsky, Ahmed Abousamra, Ümit Y. Ogras
ICCAD5
2017 GPU Performance Estimation using Software Rasterization and Machine Learning
abstract
This paper introduces a predictive modeling framework to estimate the performance of GPUs during pre-silicon design. Early-stage performance prediction is useful when simulation times impede development by rendering driver performance validation, API conformance testing and design space explorations infeasible. Our approach builds a Random Forest regression model to analyze DirectX 3D workload behavior when executed by a software rasterizer, which we have extended with a workload characterizer to collect further performance information via program counters. In addition to regression models, this work produces detailed feature rankings which can provide valuable architectural insight, and accurate performance estimates for an Intel integrated Skylake generation GPU. Our models achieve reasonable out-of-sample-error rates of 14%, with an average simulation speedup of 327x.
Kenneth O'Neal, Philip Brisk, Ahmed Abousamra, Zack Waters, Emily Shriver
ACM Trans. Embed. Comput. Syst.3
2013 Proactive circuit allocation in multiplane NoCs
abstract
This work explores a method for efficient pre-allocation of circuits in network-on-chip (NoC) to reduce communication latency and improve performance. Circuit pre-allocation eliminates the time cost of circuit establishment by using request messages to reserve the circuits for their anticipated reply messages. Requests reserve circuits in a priority order rather than for a particular time slot, avoiding delays or blocking even if the newly requested circuits conflict with previously reserved ones. Benchmark simulations show speedup in execution time of up to 16%, with an average of 8% for communication sensitive benchmarks, over a leading proposal in pre-configuring circuits.
Ahmed Abousamra, Alex K. Jones, Rami G. Melhem
DAC1
2013 Ordering circuit establishment in multiplane NoCs
abstract
Segregating networks-on-chips (NoCs) into data and control planes yields several opportunities for improving power and performance in chip-multiprocessor systems (CMPs). This article describes a hybrid packet/circuit switched multiplane network optimized to reduce latency in order to improve system performance and/or reduce system energy. Unlike traditional circuit preallocation techniques which require timestamps to reserve circuit resources, this article proposes an order-based preallocation scheme . By enforcing the order in which resources are scheduled and utilized rather than a fixed time, the NoC can take advantage of messages that arrive early while naturally tolerating message delays due to contention. Ordered circuit establishment is presented using two techniques. First, Déjà Vu switching preestablishes circuits for data messages once a cache hit is detected and prior to the requested data becoming available. Second, using Red Carpet Routing , circuits are proactively reserved for a return data message as a request message traverses the NoC. The reduced communication latency over configured circuits enable system performance improvement or saving NoC energy by reducing voltage and frequency without sacrificing performance. In simulations of 16 and 64 core CMPs, Déjà Vu switching enabled average NoC energy savings of 43% and 53% respectively. On the other hand, simulations of communication sensitive benchmarks using Red Carpet Routing show speedup in execution time of up to 16%, with an average of 10% over a purely packet switched NoC and an average of 8% over preconfiguring circuits using Déjà Vu switching .
Ahmed Abousamra, Alex K. Jones, Rami G. Melhem
ACM Trans. Design Autom. Electr. Syst.1
2012 Déjà Vu Switching for Multiplane NoCs
abstract
In chip-multiprocessors (CMPs) the network-on-chip (NoC) carries cache coherence and data messages. These messages may be classified into critical and non-critical messages. Hence, instead of having one interconnect plane to serve all traffic, power can be saved if the NoC is split into two planes: a fast plane dedicated to the critical messages and a slower, more power-efficient plane dedicated only to the non-critical messages. This split, however, can be beneficial to save energy only if system performance is not significantly degraded by the slower plane. In this work we first motivate the need for a timely delivery of the "non-critical" messages. Second, we propose Déjà Vu switching, a simple algorithm that enables reducing the voltage and frequency of one plane while reducing communication latency through circuit switching and support of advance, possibly conflicting, circuit reservations. Finally, we study the constraints that govern how slow the power-efficient plane can operate without negatively impacting system performance. We evaluate our design through simulations of 16 and 64 core CMPs. The results show that we can achieve an average NoC energy savings of 43% and 53%, respectively.
Ahmed Abousamra, Rami G. Melhem, Alex K. Jones
NOCS1
2012 Codesign of NoC and Cache Organization for Reducing Access Latency in Chip Multiprocessors
abstract
Reducing data access latency is vital to achieving performance improvements in computing. For chip multiprocessors (CMPs), data access latency depends on the organization of the memory hierarchy, the on-chip interconnect, and the running workload. Several network-on-chip (NoC) designs exploit communication locality to reduce communication latency by configuring special fast paths or circuits on which communication is faster than the rest of the NoC. However, communication patterns are directly affected by the cache organization and many cache organizations are designed in isolation of the underlying NoC or assume a simple NoC design, thus possibly missing optimization opportunities. In this work, we take a codesign approach of the NoC and cache organization. First, we propose a hybrid circuit/packet-switched NoC that exploits communication locality through periodic configuration of the most beneficial circuits. Second, we design a Unique Private (UP) caching scheme targeting the class of interconnects which exploit communication locality to improve communication latency. The Unique Private cache stores the data that are mostly accessed by each processor core in the core's locally accessible cache bank, while leveraging dedicated high-speed circuits in the interconnect to provide remote cores with fast access to shared data. Simulations of a suite of scientific and commercial workloads show that our proposed design achieves a speedup of 15.2 and 14 percent on a 16-core and a 64-core CMP, respectively, over the state-of-the-art NoC-Cache codesigned system that also exploits communication locality in multithreaded applications.
Ahmed Abousamra, Alex K. Jones, Rami G. Melhem
IEEE Trans. Parallel Distributed Syst.1
2012 Compiler-Assisted Data Distribution and Network Configuration for Chip Multiprocessors
abstract
Data access latency, a limiting factor in the performance of chip multiprocessors, grows significantly with the number of cores in nonuniform cache architectures with distributed cache banks. To mitigate this effect, we use a compiler-based approach to leverage data access locality, choose an optimized data placement and efficiently configure the on-chip network. The proposed experimental compiler framework employs novel compilation techniques to discover and represent multithreaded memory access patterns (MMAPs). At runtime, symbolic MMAPs are resolved and used by a partitioning algorithm to choose a partition of allocated memory blocks among the forked threads in the analyzed application. This partition is used to enforce data ownership by associating the data with the core that executes the thread owning the data. Based on the partition, the communication pattern of the application can be extracted. We demonstrate how this information can be used in an experimental architecture to accelerate applications. In particular, our compiler assisted data partitioning approach shows a 20 percent speedup over shared caching and 5 percent speedup over the closest runtime approximation, first touch. By leveraging the communication pattern we can achieve a comparable performance to a system that uses a complex centralized network configuration system at runtime. Thus, our final system saves significant runtime complexity and achieves an 5.1 percent additional speedup through the addition of the reconfigurable network.
Yong Li 0009, Ahmed Abousamra, Rami G. Melhem, Alex K. Jones
IEEE Trans. Parallel Distributed Syst.2
2011 NoC-aware cache design for multithreaded execution on tiled chip multiprocessors
abstract
In chip multiprocessors (CMPs), data access latency depends on the memory hierarchy organization, the on-chip interconnect (NoC), and the running workload. Reducing data access latency is vital to achieving performance improvements and scalability of threaded applications. Multithreaded applications generally exhibit sharing of data among the program threads, which generates coherence and data traffic on the NoC.
Ahmed Abousamra, Alex K. Jones, Rami G. Melhem
HiPEAC1
2011 Two-hop Free-space based optical interconnects for chip multiprocessors
abstract
Many resources are shared among the cores of chip-multiprocessors (CMPs), in particular on-chip caches and memory systems. Efficient intra-chip communication is necessary for efficient resource sharing and the performance of such systems, especially in future CMPs with hundreds or thousands of cores. Current Free-space optical networks-on-chip (NoCs) provide the potential to avoid the reduced wire performance and degraded signal integrity facing electronic networks. However, current proposals utilize fixed direction lasers and mirrors to realize one-hop all-to-all connectivity, which results in difficulties scaling to larger numbers of processors. In this paper we present two-hop optical strategies that provide better performance over the one-hop strategy while improving on both the required resources and scalability for future large scale CMPs.
Ahmed Abousamra, Rami G. Melhem, Alex K. Jones
NOCS1
2010 NoC-aware cache design for chip multiprocessors
abstract
The performance of chip multiprocessors (CMPs) is dependent on the data access latency, which is highly dependent on the design of the on-chip interconnect (NoC) and the organization of the memory caches. However, prior research attempts to optimize the performance of the NoC and cache mostly in isolation of each other. In this work we present a NoC-aware cache design that focuses on communication locality; a property both the cache and NoC affect and can exploit.
Ahmed Abousamra, Rami G. Melhem, Alex K. Jones
PACT1
2010 Compiler-assisted data distribution for chip multiprocessors
abstract
Data access latency, a limiting factor in the performance of chip multiprocessors, grows significantly with the number of cores in non-uniform cache architectures with distributed cache banks. To mitigate this effect, it is necessary to leverage the data access locality and choose an optimum data placement. Achieving this is especially challenging when other constraints such as cache capacity, coherence messages and runtime overhead need to be considered. This paper presents a compiler-based approach used for analyzing data access behavior in multi-threaded applications. The proposed experimental compiler framework employs novel compilation techniques to discover and represent multi-threaded memory access patterns (MMAPs). At run time, symbolic MMAPs are resolved and used by a partitioning algorithm to choose a partition of allocated memory blocks among the forked threads in the analyzed application. This partition is used to enforce data ownership by associating the data with the core that executes the thread owning the data. We demonstrate how this information can be used in an experimental architecture to accelerate applications. In particular, our compiler assisted approach shows a 20% speedup over shared caching and 5% speedup over the closest runtime approximation, "first touch".
Yong Li 0009, Ahmed Abousamra, Rami G. Melhem, Alex K. Jones
PACT2