Boma Anantasatya Adhi

dblp:202/7294 · also Boma A. Adhi · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-8165-9792ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Exploration of Trade-offs Between General-Purpose and Specialized Processing Elements in HPC-Oriented CGRA
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable accelerators traditionally used in embedded computing. Recently, CGRA-like devices have gained traction for HPC and AI acceleration; however, typical HPC and AI workloads often require operations that current CGRAs cannot implement, such as complex mathematical calculations. In this work, we present a broad architectural study exploring potential heterogeneous computational resources in CGRA architectures for HPC, which are not commonly considered in typical CGRA architecture research. We first improved the general-purpose Processing Element (PE) of a baseline CGRA to optimize computational resources and then developed a new specialized PE for mathematical functions commonly found in HPC applications. Finally, we evaluated multiple CGRA configurations concerning floorplan, size, general-purpose/specialized PE ratio, and Power, Performance, and Area (PPA) results from hardware synthesis.
Emanuele Del Sozzo, Xinyuan Wang 0003, Boma Anantasatya Adhi, Carlos Cortes, Jason Helge Anderson, Kentaro Sano
IPDPS3
2023 Experimental Survey of FPGA-Based Monolithic Switches and a Novel Queue Balancer
abstract
This article studies small to medium-sized monolithic switches for FPGA implementation and presents a novel switch design that achieves high algorithmic performance and FPGA implementation efficiency. Crossbar switches based on virtual output queues (VOQs) and variations have been rather popular for implementing switches on FPGAs, with applications in network switches, memory interconnects, network-on-chip (NoC) routers etc. The implementation efficiency of crossbar-based switches is well-documented on ASICs, though we show that their disadvantages can outweigh their advantages on FPGAs. One of the most important challenges in such input-queued switches is the requirement for iterative scheduling algorithms. In contrast to ASICs, this is more harmful on FPGAs, as the reduced operating frequency and narrower packets cannot “hide” multiple iterations of scheduling that are required to achieve a modest scheduling performance. Our proposed design uses an output-queued switch internally for simplifying scheduling, and a queue balancing technique to avoid queue fragmentation and reduce the need for memory-sharing VOQs. Its implementation approaches the scheduling performance of a state-of-the-art FPGA-based switch, while requiring considerably fewer resources.
Philippos Papaphilippou, Kentaro Sano, Boma Anantasatya Adhi, Wayne Luk
IEEE Trans. Parallel Distributed Syst.3
2022 The Cost of Flexibility: Embedded versus Discrete Routers in CGRAs for HPC
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable architectures that inherit the performance and usability properties of Central Processing Units (CPUs) and the reconfigurability aspects of Field-Programmable Gate Arrays (FPGAs). Historically, CGRAs have been successfully used to accelerate embedded applications and are today also being considered to accelerate High-Performance Computing (HPC) applications in future supercomputers. However, embedded systems and supercomputers are two vastly different domains with different applications and constraints, and it is today not fully understood what CGRA design decisions adequately cater to the HPC market. One such unknown design decision is regarding the interconnect that facilitates intra-CGRA communication. Today, intra-CGRA communication comes in two flavors: using routers closely embedded into the compute units or using discrete routers outside the compute units. The former trades flexibility for a reduction in hardware cost, while the latter has greater flexibility but is more resource hungry. In this paper, we aspire to understand which of both designs best suits the CGRA HPC segment. We extend our previous methodology, which consists of both a parameterized CGRA design and an OpenMPcapable compiler, to accommodate both types of routing designs, including verification tests using RTL simulation. Our results show that the discrete router design can facilitate better use of processing elements (PEs) compared to embedded routers and can achieve up to 79.27% reduction in unnecessary PE occupancy for an aggressively unrolled stencil kernel on a 18 × 16 CGRA at a (estimated) hardware resource overhead cost of 6.3x. This reduction in PE occupancy can be used, for example, to exploit instruction-level parallelism (ILP) through even more aggressive unrolling.
Boma Anantasatya Adhi, Carlos Cortes, Yiyu Tan, Takuya Kojima, Artur Podobas, Kentaro Sano
CLUSTER1
2022 Exploring Inter-tile Connectivity for HPC-oriented CGRA with Lower Resource Usage
abstract
This research aims to explore the tradeoffs between routing flexibility and hardware resource usage, ultimately reducing the resource usage of our CGRA architecture while maintaining compute efficiency. we investigate statistics of connection usages among switch blocks for benchmark DFGs, propose several CGRA architecture with a reduced connection, and evaluate their hardware cost, routability of DFGs, and computational throughput for benchmarks. We found that the topology with horizontal plus diagonal connection saves about 30% of the resource usage while maintaining virtually the same routing flexibility as the full connectivity topology.
Boma Anantasatya Adhi, Carlos Cortes, Tomohiro Ueno, Yiyu Tan, Takuya Kojima, Artur Podobas, Kentaro Sano
FPT1
2021 Efficient Queue-Balancing Switch for FPGAs
abstract
This paper presents a novel FPGA-based switch design that achieves high algorithmic performance and an efficient FPGA implementation. Crossbar switches based on virtual output queues (VOQs) and variations have been rather popular for implementing switches on FPGAs, with applications to network-on-chip (NoC) routers and network switches. The efficiency of VOQs is well-documented on ASICs, though we show that their disadvantages can outweigh their advantages on FPGAs. Our proposed design uses an output-queued switch internally for simplifying scheduling, and a queue balancing technique to avoid queue fragmentation and reduce the need for memory-sharing VOQs. Our implementation approaches the scheduling performance of the state-of-the-art, while requiring considerably fewer FPGA resources.
Philippos Papaphilippou, Kentaro Sano, Boma Anantasatya Adhi, Wayne Luk
FPT3
2017 Multicore Cache Coherence Control by a Parallelizing Compiler
abstract
A recent development in multicore technology has enabled development of hundreds or thousands core processor. However, on such multicore processor, an efficient hardware cache coherence scheme will become very complex and expensive to develop. This paper proposes a parallelizing compiler directed software coherence scheme for shared memory multicore systems without hardware cache coherence control. The general idea of the proposed method is that an automatic parallelizing compiler analyzes the control dependency and data dependency among coarse grain task in the program. Then based on the obtained information, task parallelization, false sharing detection and data restructuration to prevent false sharing are performed. Next the compiler inserts cache control code to handle stale data problem. The proposed method is built on OSCAR automatic parallelizing compiler and evaluated on Renesas RP2 with 8 SH-4A cores processor. The hardware cache coherence scheme on the RP2 processor is only available for up to 4 cores and the hardware cache coherence can be completely turned off for non-coherence cache mode. Performance evaluation is performed using 10 benchmark program from SPEC2000, SPEC2006, NAS Parallel Benchmark (NPB) and Mediabench II. The proposed method performs as good as or better than hardware cache coherence scheme. For example, 4 cores with the hardware coherence mechanism gave us speed up of 2.52 times against 1 core for SPEC2000 "equake", 2.9 times for SPEC2006 "lbm", 3.34 times for NPB "cg", and 3.17 times for MediaBench II MPEG2 Encoder. The proposed software cache coherence control gave us 2.63 times for 4 cores and 4.37 for 8 cores for "equake", 3.28 times for 4 cores and 4.76 times for 8 cores for lbm, 3.71 times for 4 cores and 4.92 times for 8 cores for "MPEG2 Encoder".
Hironori Kasahara, Boma Anantasatya Adhi, Yuhei Hosokawa, Yohei Kishimoto, Masayoshi Mase
COMPSAC (1)3