Kai-Ting Amy Wang

dblp:237/7599 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0003-2399-0844ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 PolyMorphous: An MLIR-Based Polyhedral Compiler with Loop Transformation Primitives
abstract
We present PolyMorphous, an MLIR-based polyhedral compiler that exposes a set of loop-based scheduling primitives, providing users with ample control over optimizations for input code. The primitives are expressed in a new Schedule dialect that is based on MLIR's Transform dialect. PolyMorphous' polyhedral engine collapses the primitives into a polyhedral schedule and checks for its legality. If necessary, and possible, the engine corrects an illegal schedule into a legal one. PolyMorphous is evaluated using the PolyBench suite. The evaluation validates PolyMorphous' approach in two ways. First, it shows that PolyMorphous enables users to explore the optimization space, and that it results in well-optimized code. Second, it shows that PolyMorphous' correction of illegal schedules enables users to optimize code with fewer primitives and improves their productivity by alleviating the need for manual correction. Specifically, the evaluation shows that PolyMorphous optimized code performs as well as code optimized by Pluto+ with its optimization flags tuned to achieve the best performance for each benchmark. Further, for several benchmarks, PolyMorphous optimized code performs better, confirming the value of providing users with control over optimizations. Averaged over all the benchmarks, PolyMorphous optimized code has a speedup of$1.24 \mathrm{x} / 1.25 \mathrm{x}$over that optimized by Pluto+ on Arm/X86 systems. PolyMorphous brings the approach of empowering users of polyhedral compilers with control over optimizations to a community-developed infrastructure, promoting the approach's adoptability. It does so while delivering performant code.
Jinman Zhao, Seyed Aryan Vahabpour, Xingyu Yue, Kai-Ting Amy Wang, Tarek S. Abdelrahman
IPDPS4
2024 On-the-Fly Data Layout Conversion for GEMM on AI Accelerators
abstract
GEMM accelerators used for AI typically require special layouts of their input and output data. Pre- and post-conversion to such layouts from and to standard row-major or column-major layouts degrades performance. This is particularly the case when conversion overhead cannot be amortized over multiple GEMM executions, as in, for example, LU factorization, a key computation in HPC.We propose a novel on-the-fly data layout conversion approach for GEMM used for LU factorization on Huawei’s DaVinci AI Core. The approach alleviates the bulk of conversion through the use of DMA and Vector engines to convert data layout as data moves through the memory hierarchy, and as it is processed. The approach reduces conversion overhead, but requires more use of the DMA and Vector engines, compared to when data is already in the accelerator’s required data layout.We experimentally evaluate the approach on an Ascend 910 AI processor with 32 DaVinci cores used to accelerate GEMM. We show that the approach reduces layout conversion time by 74% and that the additional use of the DMA/Vector engines reduces compute efficiency by no more than 15%. The result is an improvement in GEMM’s end-to-end performance— over explicit pre-/post-conversion—by up to ∼2X, and on average by 1.6X. In the context of LU factorization, where GEMM is repeatedly used for shrinking matrix sizes, our approach improves GEMM performance by up to 1.6X and on average by 1.4X.
Xingyu Yue, Chenchen Tang, Kai-Ting Amy Wang, Tarek S. Abdelrahman
ISPA5
2023 Tiling for DMA-Based Hardware Accelerators (WIP)
abstract
Many hardware accelerator architectures use DMA units to transfer memory which may be limited by the fixed-width size of the DMA transfer, and automatic loop tilers currently do not take the limitation of these DMA units into account. We present a compiler pass, implemented in MLIR, that uses polyhedral analysis on the memory access patterns in a loop nest and constrain the possible tile sizes based on the DMA chunk width. This allows the compiler to effectively tile loops for these architectures.
Alexandre Singer, Kai-Ting Amy Wang
LCTES2
2022 Extending SYCL's Programming Paradigm with Tensor-based SIMD Abstractions
abstract
Heterogeneous computing has emerged as an important method for supporting more than one kind of processors or accelerators in a program. There is generally a trade off between source code portability and device performance for heterogeneous programming. Thus, new programming abstractions to assist programmers to reduce their development efforts while minimizing performance penalties is extremely valuable.
Wilson Feng, Shucai Yao, Kai-Ting Amy Wang, Md Aamir Raihan, Laichun Feng, Chunrong Xu
ICPE3
2020 PSU: A Framework for Dynamic Software Updates in Multi-threaded C-Language Programs
abstract
A Dynamic Software Update (DSU) system enables an operator to modify a running program without interrupting its execution. However, creating a DSU system to allow programs written in the C programming language to be modified while they are executing is challenging. This paper presents the Portable Software Update (PSU) system, a new framework that allows the creation of C-language DSU programs. PSU offers a simple programming interface to build DSU versions of existing C programs. Once a program is built using PSU, updates can be applied by background threads that have negligible impact on the execution of the program. PSU supports multi-threaded and recursive programs without the use of safe points or thread blocking. PSU uses function indirection to redirect DSU functions calls to the newest version of the function code. Once a DSU function is invoked in a PSU program, it executes to completion using the version of the function that was active when it was invoked. However, if a new version is installed, any future calls to the same function always execute the newest version. This simple mechanism allows for quick loading of updates in PSU. PSU unloads obsolete version of DSU functions after they are no longer executing. This mechanism makes PSU the first DSU system for C-language programs that is able to unload older versions of code. This efficient use of resources enables many patches to be applied to a long-running application. A suite of specialized custom synthetic programs, and a DSU-enabled version of the MySQL database storage engine, are used to evaluate the overhead of the DSU-enabling features. The MySQL storage engine maintains over 95% of the performance of the non-DSU version and allows the entire storage engine to be updated while the database continues executing. PSU includes a simple and straightforward process for the modification of the storage engine that enables DSU.
Marcus Karpoff, José Nelson Amaral, Kai-Ting Amy Wang, Rayson Ho, Brice Dobry
SBAC-PAD3
2019 Replayable Execution Optimized for Page Sharing for a Managed Runtime Environment
abstract
We present Replayable Execution, a system for improving the efficiency of Function-as-a-Service (FaaS) frameworks. It takes advantage of standard kernel features to reduce memory usage and accelerate cold startup speed without changes to the OS kernel, language runtimes, and the surrounding FaaS deployment environment. Replayable Execution exploits the intensive-deflated execution characteristics of the majority of target applications. It uses checkpointing to save an image of an application, allowing this image to be shared across containers and resulting in speedy restoration at service startup. We apply Replayable Execution to a representative FaaS Java framework to create a ReplayableJVM execution, which together with benefits from deterministic execution of a warmed up runtime, offers 2X memory footprint reduction, and over 10X startup time improvement.
Kai-Ting Amy Wang, Rayson Ho, Peng Wu 0001
EuroSys1