Mahdi Hamzeh

dblp:08/6881 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.532014
Branch-Aware Loop Mapping on CGRAs · DAC 2014
REGIMap: register-aware application mapping on coarse-grained reconfigurable architectures (CGRAs) · DAC 2013
EPIMap: using epimorphism to map applications on CGRAs · DAC 2012
Compilers and program optimization
accelerator compilation
0.212014
Branch-Aware Loop Mapping on CGRAs · DAC 2014
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
loop mapping
0.212013
REGIMap: register-aware application mapping on coarse-grained reconfigurable architectures (CGRAs) · DAC 2013
Reconfigurable computing and FPGAs
application mapping
0.112012
EPIMap: using epimorphism to map applications on CGRAs · DAC 2012
Reconfigurable computing and FPGAs › FPGA compilation
compiler mapping
0.112012
EPIMap: using epimorphism to map applications on CGRAs · DAC 2012

Methods — techniques the papers use, named apart from their topics

predication · 0.4maximal weighted clique · 0.2heuristic · 0.2heuristic search · 0.1epimorphism · 0.1
YearPublicationVenuePosition
2015 Path selection based acceleration of conditionals in CGRAs
ShriHari RajendranRadhika, Aviral Shrivastava, Mahdi Hamzeh
DATE3
2014 Branch-Aware Loop Mapping on CGRAs
abstract
One of the challenges that all accelerators face, is to execute loops that have if-then-else constructs. There are three ways to accelerate loops with an if-then-else construct on a Coarse-grained reconfigurable architecture (CGRA): full predication, partial predication, and dual-issue scheme. In comparison with the other schemes, dual-issue scheme may achieve the best performance, but it requires compiler support -- which does not exist. In this paper, we develop compiler techniques to map loops with conditionals on CGRA for the dual-issue scheme. Our experiments show: i) 40% of loops that can be accelerated on CGRA have conditionals, ii) The proposed dual-issue scheme enables our compiler to accelerate loops 40% faster than full predication scheme proposed in [12], and iii) Our compiler assisted dual issue scheme can exploit richer interconnects, if present.
Mahdi Hamzeh, Aviral Shrivastava, Sarma B. K. Vrudhula
DAC1
2013 REGIMap: register-aware application mapping on coarse-grained reconfigurable architectures (CGRAs)
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) are an extremely attractive platform when both performance and power efficiency are paramount. Although the power-efficiency of CGRAs can be very high, their performance critically hinges upon the capabilities of the compiler. This is because a CGRA compiler has to perform explicit pipelining, scheduling, placement, and routing of operations. Existing CGRA compilers struggle with two main problems: 1) effectively utilizing the local register files in the PEs, and 2) high compilation times. This paper significantly improves the state-of-the-art in CGRA compilers by first creating a precise and general formulation of the problem of loop mapping on CGRAs, considering the local registers, and from the insights gained from the problem formulation, distilling an efficient and constructive heuristic solution. We show that the mapping problem, once characterized, can be reduced to the problem of finding maximal weighted clique in the product graph of the time-extended CGRA and the data dependence graph of the kernel. The heuristic we've developed results in average of 1.89 X better performance than the state-of-the-art methods when applied to several kernels from multimedia and SPEC2006 benchmarks. A unique feature of our heuristic is that it learns from failed attempts and constructively changes the schedule to achieve better mappings at lower compilation times.
Mahdi Hamzeh, Aviral Shrivastava, Sarma B. K. Vrudhula
DAC1
2012 EPIMap: using epimorphism to map applications on CGRAs
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) are an attractive platform that promise simultaneous high-performance and high power-efficiency. One of the primary challenges in using CGRAs is to develop efficient compilers that can automatically and efficiently map applications to the CGRA. To this end, this paper makes several contributions: i) Using Re-computation for Resource Limitations: For the first time in CGRA compilers, we propose the use of re-computation as a solution for resource limitation problem. This extends the solutions space, and enables better mappings, ii) General Problem Formulation: A precise and general formulation of the application mapping problem on a CGRA is presented, and its computational complexity is established. iii) Extracting an Efficient Heuristic: Using the insights from the problem formulation, we design an effective global heuristic called EPIMap. EPIMap transforms the input specification (a directed graph) to an Epimorphic equivalent graph that satisfies the necessary conditions for mapping on to a CGRA, reducing the search space. Experimental results on 14 important kernels extracted from well known benchmark programs show that using EPIMap can improve the performance of the kernels on CGRA by more than 2.8X on average, as compared to one of the best existing mapping algorithm, EMS. EPIMap was able to achieve the theoretical best performance for 9 out of 14 benchmarks, while EMS could not achieve the theoretical best performance for any of the benchmarks. EPIMap achieves better mappings at acceptable increase in the compilation time.
Mahdi Hamzeh, Aviral Shrivastava, Sarma B. K. Vrudhula
DAC1
2011 Enabling Multithreading on CGRAs
abstract
Coarse-Grained Reconfigurable Arrays or CGRAs are programmable fabrics that promise both high performance and high power efficiency. Traditionally, CGRAs were used to accelerate extremely-embedded systems, and were typically manually programmed. However, as CGRAs are conceived to be used as more general-purpose accelerators, there is a need to develop software tools and capabilities. Much work has been done on developing compiler techniques for CGRAs, making programming them easier, however, there is no support for multithreading. As an accelerator to a multithreaded processor, CGRAs now are restricted to accelerating only one kernel of one thread running on the processor at any point in time. Supporting multithreading is difficult, since the start times and end times of threads are dynamic in nature, while CGRAs are statically scheduled. In this paper, we propose a strategy to do multithreading on a CGRA. The chief capability that we develop is a scheme to quickly transform an existing application mapping using the entire CGRA to one using only a fraction of it. Our experimental results on kernels from multimedia applications demonstrate that multithreading support can improve the total throughput of a CGRA by over 30%, 75%, and 150% on 4×4, 6×6, and 8×8 CGRAs, respectively, compared to single-threaded methods.
Aviral Shrivastava, Jared Pager, Reiley Jeyapaul, Mahdi Hamzeh, Sarma B. K. Vrudhula
ICPP4
2009 Computationally efficient active rule detection method: Algorithm and architecture
Mahdi Hamzeh, Hamid Reza Mahdiani, Ahmad Saghafi, Sied Mehdi Fakhraie, Caro Lucas
Fuzzy Sets Syst.1