Omar Ragheb

dblp:226/3781 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0009-6517-9586ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 7 since 2021
YearPublicationVenuePosition
2026 ARCS: Architecture-Responsive CGRA Scheduling
abstract
Scheduling is a key aspect in mapping applications to coarse-grained reconfigurable architectures (CGRA). During scheduling, the number of pipeline registers on each path is determined to ensure that the data for an operation arrives at the correct cycle. Traditional scheduling methods, such as as soon as possible (ASAP) and as late as possible (ALAP), determine the pipeline registers without considering the placement of operations. This can restrict the mapping algorithm, forcing placement to accommodate scheduling and limiting routing options to satisfy both scheduling and placement. To overcome the limitations of traditional schedulers, we propose a method that adaptively adjusts the schedule during the initial routing phase. In this approach, the mapping algorithm begins with an ASAP schedule to establish initial schedule constraints. We then utilize simulated annealing for placement, and we employ a architecture-responsive scheduling algorithm post-placement to update the schedule of each edge based on the placement and generate an initial routing solution with overlaps. Afterwards, the PathFinder algorithm is applied to the initial routing solution, along with the generated schedule, to find a valid routing with no overlaps. Our results demonstrate that the architecture-responsive scheduling approach maintains a quality comparable to that of conventional ASAP scheduling. Furthermore, architecture-responsive scheduling enables generic mapping of applications onto restricted architectures that do not allow routes to bypass pipeline registers, a challenge that traditional schedulers do not address.
Omar Ragheb, Jason Helge Anderson
ASP-DAC1
2024 MLIR-to-CGRA: A Versatile MLIR-Based Compiler Framework for CGRAs
abstract
Coarse-grained reconfigurable architectures (CGRAs) are programmable hardware platforms with coarse-grained programmable logic blocks and word-wide configurable interconnect. In this paper, we describe a high-level compilation framework tailored for CGRAs. The input to the framework is a C-language program and description of the available CG RA architectural features. This work primarily focuses on automatically handling sequential rectangular loop access patterns. It achieves this by running automated design space exploration (DSE) written within the MLIR compiler representation [7] to leverage both spatial and temporal parallelism inherent in CG RAs. Furthermore, the compiler is engineered to be architecture-agnostic. It employs standard loop (and other) compiler optimization passes alongside CG RA-specific passes to generate a data flow graph (DFG) tailored for a CG RA implementation. In an experimental study, we optimize kernels to increase CG RA mappability and show the impact of automated DSE on the kernel performance. Moreover, the framework eases the task of compilation from software to RISC-V+CGRA hybrid systems.
Omar Ragheb, Stephen Wicklund, Jason Helge Anderson
ASAP2
2024 CLUMAP: Clustered Mapper for CGRAs with Predication
abstract
Coarse-grained reconfigurable architectures (CGRAs) have gained popularity as accelerators for compute-intensive kernels. Complex CGRA architectures that support key features such as multi-context and predication are being developed to support a wider range of kernels. However, mapping applications on these complex architectures poses significant challenges. In this paper, we provide an architecture-agnostic clustered mapping technique and a new cost function tailored for simulated-annealing placement. The mapper simplifies placement and routing phases, demonstrating significant speedup for popular CGRA architectures: HyCUBE and ADRES. Additionally, our method demonstrates an increase in mapping success for the ADRES architecture.
Omar Ragheb, Jason Helge Anderson
DAC1
2022 Modeling and Exploration of Elastic CGRAs
abstract
Elastic design concepts have the potential to bring multiple benefits to coarse-grained reconfigurable arrays (CGRAs) architecture, including the ability to interface with memories, having unknown latencies, incorporate run-time variable-latency processing elements, and ease the CGRA mapping challenges of scheduling, placement and routing. However, there are overheads in terms of power, performance and area (PPA) associated with the design and implementation of elastic circuits. In this paper, we quantify these overheads in the CGRA context by first extending an open-source CGRA modelling and exploration framework (CGRA-ME) [4] to allow elastic circuit primitives (e.g. fork, join, merge, diverge, etc.) to be used when composing/modelling a CGRA architecture. We then use this new capability to “elasticize” two widely studied CGRA architectures, ADRES [11] and HyCUBE [8]. The PPA of the elastic versions of the CGRAs are compared with their traditional statically scheduled counterparts. We also evaluate the PPA “cost” of several elastic-circuit design points, such as elastic buffer length and inclusion of merge and diverge components.
Omar Ragheb, David Ma, Jason Helge Anderson
FPL1
2021 CGRA-ME: An Open-Source Framework for CGRA Architecture and CAD Research : (Invited Paper)
abstract
Coarse-grained reconfigurable arrays (CGRAs) are programmable hardware platforms that can be used to realize application-specific accelerators for higher performance and energy efficiency. A CGRA is a 2D array of configurable logic blocks & interconnect, where the logic blocks are typically large & ALU-like, and the interconnect is word-wide. CGRA-ME is a software framework that enables the modelling and exploration of CGRA architectures, as well as research on CGRA CAD algorithms. With CGRA-ME, an architect can specify a CGRA architecture at a high level of abstraction. A set of applications can be mapped onto the architecture to assess the mappability, power, performance and cost. CGRA-ME also allows one to generate synthesizable Verilog RTL for the modelled CGRA, permitting its implementation as an ASIC or FPGA overlay. In this paper, we describe the CGRA-ME framework [5] and overview its capabilities and current limitations. We discuss ongoing and prior research conducted with the framework, as well as outline future plans. We believe CGRA-ME will be a valuable contribution to the community, enabling new research on CGRA CAD & architectures.
Jason Helge Anderson, Rami Beidas, Vimal Chacko, Hsuan Hsiao, Xiaoyi Ling, Omar Ragheb, Xinyuan Wang 0003
ASAP6
2021 High-Level Synthesis of Transactional Memory
Omar Ragheb, Jason Helge Anderson
ASP-DAC1
2021 Profiling-Based Control-Flow Reduction in High-Level Synthesis
abstract
Control flow in a program can be represented in a directed graph, called the control flow graph (CFG). Nodes in the graph represent straight-line segments of code, basic blocks, and directed edges between nodes correspond to transfers of control. We present a methodology to selectively reduce control flow by collapsing basic blocks into their parent blocks, revealing increased instruction-level parallelism to a high-level synthesis (HLS) scheduler, thereby raising circuit performance.We evaluate our approach within an HLS tool that allows a C-language software program to be automatically synthesized into a hardware circuit, using the CHStone benchmark suite [1], targeting an Intel Cyclone V FPGA. For individual benchmark circuits we observe cycle count reductions up to 20.7% and wall-clock time reductions up to 22.6%, and 6% on average.
Austin Liolli, Omar Ragheb, Jason Helge Anderson
FPT2
2018 High-Level Synthesis of FPGA Circuits with Multiple Clock Domains
abstract
We consider the high-level synthesis of circuits with multiple clock domains in a bid to raise circuit performance. A profiling-based approach is used to select time-intensive sub-circuits within a larger circuit to operate on separate clock domains. This isolates the critical paths of the sub-circuits from the larger circuit, allowing the sub-circuits to be clocked at the highest-possible speed. The open-source LegUp high-level synthesis tool (HLS) is modified to automatically insert clock-domain-crossing circuitry for signals crossing between two domains. The scheduling and binding phases of HLS were changed to reflect the impact of multiple clock domains on memory. Namely, the block RAMs in FPGAs are dual-port, where each port can operate on a different domain, implying that sub-circuits on different domains can access shared memory provided the domains of the memory ports are consistent with the sub-circuit domains. In an experimental study, we apply multi-clock domain HLS to the CHStone benchmark suite and demonstrate average wall-clock time improvements of 33%.
Omar Ragheb, Jason Helge Anderson
FCCM1
2018 A High-Level Synthesis Case Study on Light Propagation Simulation in Turbid Media
abstract
In this work, we look into the benefit of using High-Level Synthesis (HLS) in building and accelerating complex systems with floating-point operations. We present a highly-optimized Monte-Carlo (MC) simulator for light propagation in 3D voxel-based biological tissue representations using HLS. We show how to utilize HLS in creating efficient structures that help achieve the desired throughput. We use Vivado to implement the design on a Xilinx Kintex Ultrascale FPGA running at 150 MHz. With a design time of 1.5 months, experimental results show a 3x speedup against the fastest software simulator published to date.
Abdul-Amir Yassine, Yasmin Afsharnejad, Omar Ragheb, Vaughn Betz, Paul Chow
FCCM3