Grace Dinh

dblp:215/4425 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-9626-098XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Portable, High Performance Matrix Multiplication Micro-Kernels for RISC-V with ExO
abstract
The proliferation of RISC-V platforms and their use in a wide variety of scientific applications, including deep learning scenarios, has dramatically increased the interest to generate optimized code for them. In the field of HPC (High Performance Computing), the RISCV ISA (Instruction Set Architecture) has been adopted by a wide variety of designs with different micro-architecture; as a result, performance portability of existing codes is a major endeavor. Code generators and compilers such as Apache TVM, MLIR, or EXO provide a hardware abstraction for implementing optimized hardware-aware codes, thus reducing development time and potential errors. These generators can handle the full software stack, from basic micro-kernels to complex operations. In this work, we focus on the optimization of GEMM (general matrix-matrix multiplication), a key operation on top of which dense linear algebra libraries and deep learning frameworks are built. Specifically, we present an EXO-based GEMM microkernel generator for the RISC-V ISA with RVV vector extensions that addresses the lack of high-performance and portable GEMM micro-kernels. Our results demonstrate that, by generating a wide range of micro-kernels, one can obtain GEMM realizations that outperform those in the state-of-the-art high performance libraries.
Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Jie Lei 0007, Yuka Ikarashi, Grace Dinh, Francisco D. Igual, Enrique S. Quintana-Ortí
PDP6
2024 Tackling the Matrix Multiplication Micro-Kernel Generation with Exo
abstract
The optimization of the matrix multiplication (or GEMM) has been a need during the last decades. This operation is considered the flagship of current linear algebra libraries such as BLIS, OpenBLAS, or Intel OneAPI because of its widespread use in a large variety of scientific applications. The GEMM is usually implemented following the GotoBLAS philosophy, which tiles the GEMM operands and uses a series of nested loops for performance improvement. These approaches extract the maximum computational power of the architectures through small pieces of hardware-oriented, high-performance code called micro-kernel. However, this approach forces developers to generate, with a nonnegligible effort, a dedicated micro-kernel for each new hardware. In this work, we present a step-by-step procedure for generating micro-kernels with the Exo compiler that perform close to (or even better than) manually developed microkernels written with intrinsic functions or assembly language. Our solution also improves the portability of the generated code, since a hardware target is fully specified by a concise library-based description of its instructions.
Adrián Castelló 0001, Julian Bellavita, Grace Dinh, Yuka Ikarashi, Héctor Martínez 0002
CGO3
2023 DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
abstract
In the hardware design space exploration process, it is critical to optimize both hardware parameters and algorithm-to-hardware mappings. Previous work has largely approached this simultaneous optimization problem by separately exploring the hardware design space and the mapspace—both individually large and highly nonconvex spaces—independently. The resulting combinatorial explosion has created significant difficulties for optimizers.
Charles Hong, Qijing Huang 0001, Grace Dinh, Mahesh Subedar, Sophia Shao
MICRO3
2021 CoSA: Scheduling by Constrained Optimization for Spatial Accelerators
abstract
Recent advances in Deep Neural Networks (DNNs) have led to active development of specialized DNN accelerators, many of which feature a large number of processing elements laid out spatially, together with a multi-level memory hierarchy and flexible interconnect. While DNN accelerators can take advantage of data reuse and achieve high peak throughput, they also expose a large number of runtime parameters to the programmers who need to explicitly manage how computation is scheduled both spatially and temporally. In fact, different scheduling choices can lead to wide variations in performance and efficiency, motivating the need for a fast and efficient search strategy to navigate the vast scheduling space.To address this challenge, we present CoSA, a constrained-optimization-based approach for scheduling DNN accelerators. As opposed to existing approaches that either rely on designers’ heuristics or iterative methods to navigate the search space, CoSA expresses scheduling decisions as a constrained-optimization problem that can be deterministically solved using mathematical optimization techniques. Specifically, CoSA leverages the regularities in DNN operators and hardware to formulate the DNN scheduling space into a mixed-integer programming (MIP) problem with algorithmic and architectural constraints, which can be solved to automatically generate a highly efficient schedule in one shot. We demonstrate that CoSA-generated schedules significantly outperform state-of-the-art approaches by a geometric mean of up to 2.5× across a wide range of DNN networks while improving the time-to-solution by 90×.
Qijing Huang 0001, Aravind Kalaiah, James Demmel, Grace Dinh, John Wawrzynek, Thomas Norell, Sophia Shao
ISCA5
2020 Communication-Optimal Tilings for Projective Nested Loops with Arbitrary Bounds
abstract
Abstract Reducing communication - either between levels of a memory hierarchy or between processors over a network - is a key component of performance optimization (in both time and energy) for many nested loop problems, including dense linear algebra, particle interactions, and machine learning. Previous tiling based approaches for these problems have been used to find both lower bounds on the communication required to execute them and optimal rearrangements, or blockings, to attain such lower bounds. However, such general approaches have typically assumed the problem sizes are large, an assumption that is often not met in practice. In this paper, we provide an efficient way to both find and obtain, via an appropriate, efficiently constructible blocking, communication lower bounds and matching tilings which attain these lower bounds for nested loop programs with arbitrary loop bounds that operate on multidimensional arrays in the projective case, where the array indices are subsets of the loop indices. Our approach works on all such problems, regardless of dimensionality, size, memory access patterns, or number of arrays.
Grace Dinh, James Demmel
SPAA1