Xinyuan Wang 0003

dblp:81/3778-3 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2024
0009-0003-9997-4023ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Raising Compute Density of Molecular Dynamics Simulation Through Approximate Memoization
abstract
Molecular dynamics (MD) simulation involves simulating the interactions of particles. MD has many applications in basic biological sciences, drug discovery, materials science, and other fields. Simulating$1\mu \mathrm{s}$of a 100K-atom system can take hours or days11https://www.bdr.riken.jp/en/research/labs/taiji-mlmdgrape4.html, where the compute-heavy aspect of MD is calculating the long-range forces between pairs of particles. In this paper, we explore the application of approximate computing in MD as a means to improve compute density. Specifically, we employ approximate memoization, where previously computed forces (and more) are stored in a table, and are retrieved in subsequent force calculations, provided the inputs to the force calculation are the same or similar. If the prior-computed table values can be used, significant computational work is avoided. In an experimental study, we apply software simulation to understand the degree to which approximation is feasible. We then propose a hardware implementation of memoization to be used within an ASIC MD simulator, MDGRAPE-4A [18]. We show that compute density, measured as$\text{pair-interactions}/(s\cdot\mu m^{2})$is improved substantially, between 40 % and 70 % for the studied cases. This is contingent on the particular system being simulated, the table size, and permitted level of approximation.
Salim Khemira, Xinyuan Wang 0003, Yutaka Tamiya, Makoto Taiji, Takahide Yoshikawa, Jason Helge Anderson
ASAP2
2024 Exploration of Trade-offs Between General-Purpose and Specialized Processing Elements in HPC-Oriented CGRA
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable accelerators traditionally used in embedded computing. Recently, CGRA-like devices have gained traction for HPC and AI acceleration; however, typical HPC and AI workloads often require operations that current CGRAs cannot implement, such as complex mathematical calculations. In this work, we present a broad architectural study exploring potential heterogeneous computational resources in CGRA architectures for HPC, which are not commonly considered in typical CGRA architecture research. We first improved the general-purpose Processing Element (PE) of a baseline CGRA to optimize computational resources and then developed a new specialized PE for mathematical functions commonly found in HPC applications. Finally, we evaluated multiple CGRA configurations concerning floorplan, size, general-purpose/specialized PE ratio, and Power, Performance, and Area (PPA) results from hardware synthesis.
Emanuele Del Sozzo, Xinyuan Wang 0003, Boma Anantasatya Adhi, Carlos Cortes, Jason Helge Anderson, Kentaro Sano
IPDPS2
2021 CGRA-ME: An Open-Source Framework for CGRA Architecture and CAD Research : (Invited Paper)
abstract
Coarse-grained reconfigurable arrays (CGRAs) are programmable hardware platforms that can be used to realize application-specific accelerators for higher performance and energy efficiency. A CGRA is a 2D array of configurable logic blocks & interconnect, where the logic blocks are typically large & ALU-like, and the interconnect is word-wide. CGRA-ME is a software framework that enables the modelling and exploration of CGRA architectures, as well as research on CGRA CAD algorithms. With CGRA-ME, an architect can specify a CGRA architecture at a high level of abstraction. A set of applications can be mapped onto the architecture to assess the mappability, power, performance and cost. CGRA-ME also allows one to generate synthesizable Verilog RTL for the modelled CGRA, permitting its implementation as an ASIC or FPGA overlay. In this paper, we describe the CGRA-ME framework [5] and overview its capabilities and current limitations. We discuss ongoing and prior research conducted with the framework, as well as outline future plans. We believe CGRA-ME will be a valuable contribution to the community, enabling new research on CGRA CAD & architectures.
Jason Helge Anderson, Rami Beidas, Vimal Chacko, Hsuan Hsiao, Xiaoyi Ling, Omar Ragheb, Xinyuan Wang 0003
ASAP7
2021 Double-Pumping the Interconnect for Area Reduction in Coarse-Grained Reconfigurable Arrays
abstract
We consider double-pumped interconnect as a means of area reduction in coarse-grained reconfigurable arrays (CGRAs). Interconnect multiplexers comprise a considerable portion of CGRA area. We apply double-pumping to halve the word-width of the interconnect multiplexers, saving area. The interconnect is operated at twice the system clock frequency, where the top and bottom half-words of a value are communicated in the first and second half of a clock cycle. Several circuit-level approaches for double-pumping are considered, and evaluated in different CGRA architectures with varied interconnect richness. Area and performance consequences are assessed through a 45nm standard-cell ASIC implementation. Overall CGRA area improvements of up to 16% are observed, depending on the CGRA architecture and double-pumping implementation.
Xinyuan Wang 0003, Hsuan Hsiao, Jason Helge Anderson
ASAP1