EDBT 2026 Demo / reviewers in the wild / expert
Zhaoying Li 0004
dblp:218/1091-4
· DBLP profile ↗
19ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0001-7513-9494ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 5 first-author · 17 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Data-Driven Dynamic Execution Orchestration ArchitectureabstractDomain-specific accelerators deliver exceptional performance on their target workloads through fabrication-time orchestrated datapaths. However, such specialized architectures often exhibit performance fragility when exposed to new kernels or irregular input patterns. In contrast, programmable architectures like FPGAs, CGRAs, and GPUs rely on compile-time orchestration to support a broader range of applications; but they are typically less efficient under irregular or sparse data. Pushing the boundaries of programmable architectures requires designs that can achieve efficiency and high-performance on par with specialized accelerators while retaining the agility of general-purpose architectures. Zhenyu Bai, Pranav Dangi, Rohan Juneja, Zhaoying Li 0004, Zhanglu Yan, Huiying Lan, Tulika Mitra |
ASPLOS (1) | 4 |
| 2026 | HighP: In-Memory Acceleration of SpGEMM With High Bank-Level ParallelismabstractGeneralized sparse matrix-matrix multiplication (SpGEMM) is a critical computational primitive that is highly memory-bound due to its inherent irregular data-dependent access pattern. Near-bank processing-in-memory (PIM) is a promising technique to overcome the memory bottleneck of SpGEMM by performing computations near the bank where the data is stored. However, earlier PIM studies fail to fully utilize the high memory bandwidth when performing SpGEMM due to low bank-level parallelism. As a result, 80% memory bandwidth is wasted as observed in our in-depth experimental analysis.Our key insight in this paper is that non-conflicting matrix columns in SpGEMM, where each row of these columns has no more than one non-zero element, can be processed simultaneously in different banks. We hence propose HighP, a near-bank PIM accelerator for SpGEMM with high bank-level parallelism. We first propose a set-based search mechanism, which finds non-conflicting columns through set operations automatically. We then develop a DIMM-based PIM architecture with detailed hardware and workflow designs for SpGEMM. Set operation logic and unified scratchpad memory management are designed to perform set operations with high computational parallelism and to enhance data reuse, respectively. HighP provides up to 17.88× performance improvement compared to the state-of-th-eart SpGEMM accelerator and achieves up to 8.19× performance improvement over the state-of-the-art PIM solution. Dan Chen 0006, Huize Li, Huiying Lan, Zhaoying Li 0004, Pengcheng Yao, Tulika Mitra |
IEEE Trans. Computers | 4 |
| 2025 | Enhancing CGRA Efficiency Through Aligned Compute and Communication ProvisioningabstractCoarse-grained Reconfigurable Arrays (CGRAs) are domain-agnostic accelerators that enhance the energy efficiency of resource-constrained edge devices. The CGRA landscape is diverse, exhibiting trade-offs between performance, efficiency, and architectural specialization. However, CGRAs often overprovision communication resources relative to their modest computing capabilities. This occurs because the theoretically provisioned programmability for CGRAs often proves superfluous in practical implementations. Zhaoying Li 0004, Pranav Dangi, Chenyang Yin, Thilini Kaushalya Bandara, Rohan Juneja, Cheng Tan 0002, Zhenyu Bai, Tulika Mitra |
ASPLOS (1) | 1 |
| 2025 | Rewire: Advancing CGRA Mapping Through a Consolidated Routing ParadigmabstractCoarse-Grained Reconfigurable Arrays (CGRAs) balance the performance and power efficiency in computing systems. Effective compilers play a crucial role in fully realizing its potential. The compiler maps Data Flow Graphs (DFGs), which represent compute-intensive loop kernels, onto CGRAs. However, existing compilers often tackle DFG nodes individually, neglecting their intricate inter-dependencies. We introduce a novel mapping paradigm called Rewire that can place and route multiple nodes in one shot. Rewire first generates routing information that is shareable among multiple nodes via propagation. Then, Rewire intersects the routing information to generate individual placement candidates for each node. Finally, Rewire innovatively utilizes data dependencies as constraints to quickly find suitable placement for multiple nodes together. Our evaluation demonstrates that Rewire can generate more near-optimal mappings than prior works. Rewire achieves 2.1x and 1.3x performance improvement and 13.5x and 4.7x compilation time reduction, respectively, compared to two popular mappers. Zhaoying Li 0004, Dhananjaya Wijerathne, Dan Chen 0006, Huize Li, Cheng Tan 0002, Tulika Mitra |
DAC | 1 |
| 2025 | Building an Open CGRA Ecosystem for Agile InnovationabstractModern computing workloads, particularly in AI and edge applications, demand hardware-software co-design to meet aggressive performance and energy targets. Such co-design benefits from open and agile platforms that replace closed, vertically integrated development with modular, community-driven ecosystems. Coarse-Grained Reconfigurable Architectures (CGRAs), with their unique balance of flexibility and efficiency, are particularly well-suited for this paradigm. When built on open-source hardware generators and software toolchains, CGRAs provide a compelling foundation for architectural exploration, cross-layer optimization, and real-world deployment.In this paper, we will present an open CGRA ecosystem that we have developed to support agile innovation across the stack. Our contributions include HyCUBE, a CGRA with a reconfigurable single-cycle multi-hop interconnect for efficient data movement; PACE, which embeds a power-efficient HyCUBE within a RISC-V SoC targeting edge computing; and Morpher, a fully open-source, architecture-adaptive CGRA design framework that supports design space exploration, compilation, simulation, and validation. By embracing openness at every layer, we aim to lower barriers to innovation, enable reproducible research, and demonstrate how CGRAs can anchor the next wave of agile hardware development. We will conclude with a call for a unified abstraction layer for CGRAs and spatial accelerators, one that decouples hardware specialization from software development. Such a representation would unlock architectural portability, compiler innovation, and a scalable, open foundation for spatial computing. Rohan Juneja, Pranav Dangi, Thilini Kaushalya Bandara, Zhaoying Li 0004, Dhananjaya Wijerathne, Li-Shiuan Peh, Tulika Mitra |
ICCAD | 4 |
| 2025 | Inkstream: Instantaneous GNN Inference on Dynamic Graphs via Incremental UpdateabstractGraph Neural Network (GNN) on dynamic graphs that evolve with time necessitates constant updates. Current approaches aim to mitigate computational costs by limiting updates to the affected areas, essentially the$k$-hop neighborhood surrounding modified edges/vertices in$k$-layer GNNs. However, we identified that these strategies often involve unnecessary computation: (1) Within the$k$-hop neighborhood, a substantial number of nodes remain unaffected by changes in edges/vertices when GNN employs max or min as its aggregation function; (2) For certain model architectures, the node embeddings can be incrementally updated with minimal memory access and computation. In response to these observations, we developed InkStream, an innovative and general method for real-time GNN inference by avoiding unnecessary updates, significantly reducing inference time and energy cost. InkStream supports all common GNN aggregation functions while imposing minimal constraints on model architecture. It is grounded in the principle of minimalistic propagation and data retrieval, employing an event-based system to manage both the inter-layer propagation of effects and the intra-layer incremental updates of node embeddings. Additionally, InkStream offers remarkable extensibility and ease of configuration, making it adaptable to evolving GNN model structures. Our evaluation across three GNN models on six graph datasets reveals that InkStream significantly accelerates inference time from hours to mere milliseconds. The code is available at https://github.com/WuDan0399/InkStream. Zhaoying Li 0004, Tulika Mitra |
IPDPS | 2 |
| 2024 | FHE-CGRA: Enable Efficient Acceleration of Fully Homomorphic Encryption on CGRAsabstractFully Homomorphic Encryption (FHE) is an attractive privacy-preserving technique that allows computation directly on encrypted data without decryption. However, it incurs significant performance and memory costs due to intensive computations. In this work, we investigate the execution of FHE-enabled machine learning (ML) applications. We show that the runtime hardware reconfigurability of the underlying execution units of homomorphic operations is highly desirable for efficient hardware resource utilization during FHE-ML execution, due to the changing FHE encryption variants across different ML stages (e.g., the multiplicative level of the ciphertext) and corresponding optimal execution unit design. Based on the observation, we propose FHE-CGRA, a coarse-grained re-configurable architecture (CGRA) acceleration framework with an MLIR-based compiler toolchain for end-to-end homomorphic applications. The experiment shows that FHE-CGRA achieves up-to 8.15× speedup against a conventional CGRA baseline for accelerating the inference of FHE-encrypted convolution neural network (FHE-CNN) models, and up-to 16.48× power efficiency w.r.t. the state-of-the-art FPGA-based FHE-CNN accelerator design. Miaomiao Jiang, Yilan Zhu, Honghui You, Cheng Tan 0002, Zhaoying Li 0004, Jiming Xu, Lei Ju 0001 |
DAC | 5 |
| 2024 | PACE: A Scalable and Energy Efficient CGRA in a RISC-V SoC for Edge Computing Applicationsabstract▪Coarse-grained reconfigurable arrays (CGRAs) deliver high energy efficiency while maintaining the programmability advantages. ▪CGRA is the ideal candidate for efficiently handling loop kernels, which allows it to offload repetitive looping functions such as vector multiplication or hashing algorithms from CPUs. ▪It relies on a compiler to convert a given workload into a data flow graph (DFG) which is then mapped onto the hardware in a manner that achieves the highest possible energy efficiency. Vishnu P. Nambiar, Yi Sheng Chong, Thilini Kaushalya Bandara, Dhananjaya Wijerathne, Zhaoying Li 0004, Rohan Juneja, Li-Shiuan Peh, Tulika Mitra, Anh-Tuan Do |
HCS | 5 |
| 2024 | ASADI: Accelerating Sparse Attention Using Diagonal-based In-Situ ComputingabstractThe self-attention mechanism is the performance bottleneck of Transformer-based language models, particularly for long sequences. Researchers have proposed using sparse attention to speed up the Transformer. However, sparse attention introduces significant random access overhead, limiting computational efficiency. To mitigate this issue, researchers attempt to improve data reuse by utilizing row/column locality. Unfortunately, we find that sparse attention does not naturally exhibit strong row/column locality, but instead has excellent diagonal locality. Thus, it is worthwhile to use diagonal compression (DIA) format. However, existing sparse matrix computation paradigms struggle to efficiently support DIA format in attention computation. To address this problem, we propose ASADI, a novel software-hardware co-designed sparse attention accelerator. In the soft-ware side, we propose a new sparse matrix computation paradigm that directly supports the DIA format in self-attention computation. In the hardware side, we present a novel sparse attention accelerator that efficiently implements our computation paradigm using highly parallel in-situ computing. We thoroughly evaluate ASADI across various models and datasets. Our experimental results demonstrate an average performance improvement of 18.6 × and energy savings of 2.9× compared to a PIM-based baseline. Huize Li, Zhaoying Li 0004, Zhenyu Bai, Tulika Mitra |
HPCA | 2 |
| 2024 | ICED: An Integrated CGRA Framework Enabling DVFS-Aware AccelerationabstractCoarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. However, the relationships of voltage and frequency with the utilization of CGRA resources and the dynamic management of them are not well explored, leading to inefficient designs. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non-performance-constraining kernels. This paper proposes ICED - an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICED proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICED is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICED improves average utilization by$\mathbf{2}.\mathbf{3}\times$and energy-efficiency by$\mathbf{1}.\mathbf{32}\times$over a conventional CGRA. With streaming applications, ICED can achieve up to$\mathbf{1}.\mathbf{26}\times$energy-efficiency compared with a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput. Cheng Tan 0002, Miaomiao Jiang, Deepak Patil, Yanghui Ou, Zhaoying Li 0004, Lei Ju 0001, Tulika Mitra, Antonino Tumeo, Jeff Zhang 0001 |
MICRO | 5 |
| 2024 | Flip: Data-centric Edge CGRA AcceleratorabstractCoarse-Grained Reconfigurable Arrays (CGRA) are promising edge accelerators due to the outstanding balance in flexibility, performance, and energy efficiency. Classic CGRAs statically map compute operations onto the processing elements (PE) and route the data dependencies among the operations through the Network-on-Chip. However, CGRAs are designed for fine-grained static instruction-level parallelism and struggle to accelerate applications with dynamic and irregular data-level parallelism, such as graph processing. To address this limitation, we present Flip , a novel accelerator that enhances traditional CGRA architectures to boost the performance of graph applications. Flip retains the classic CGRA execution model while introducing a special data-centric mode for efficient graph processing. Specifically, it leverages the inherent data parallelism of graph algorithms by mapping graph vertices onto PEs rather than the operations and supporting dynamic routing of temporary data according to the runtime evolution of the graph frontier. Experimental results demonstrate that Flip achieves up to 36× speedup with merely 19% more area compared to classic CGRAs. Compared to state-of-the-art large-scale graph processors, Flip has similar energy efficiency and 2.2× better area efficiency at a much-reduced power/area budget. Peng Chen 0027, Thilini Kaushalya Bandara, Zhaoying Li 0004, Tulika Mitra |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2022 | PANORAMA: divide-and-conquer approach for mapping complex loop kernels on CGRAabstractCGRAs are well-suited as hardware accelerators due to power efficiency and reconfigurability. However, their potential is limited by the inability of the compiler to map complex loop kernels onto the architectures effectively. We propose PANORAMA, a fast and scalable compiler based on a divide-and-conquer approach to generate quality mapping for complex dataflow graphs (DFG) representing loop bodies onto larger CGRAs. PANORAMA improves the throughput of the mapped loops by up to 2.6x with 8.7x faster compilation time compared to the state-of-the-art techniques. Dhananjaya Wijerathne, Zhaoying Li 0004, Thilini Kaushalya Bandara, Tulika Mitra |
DAC | 2 |
| 2022 | LISA: Graph Neural Network based Portable Mapping on Spatial AcceleratorsabstractSpatial accelerators, such as Coarse-Grained Reconfigurable Arrays (CGRA), provide a promising pathway to scale the performance and power efficiency of computing systems. These accelerators depend on effective compilers to take advantage of the parallelism offered by the underlying architecture. Currently, the compilers are handcrafted for spatial accelerators, which is challenging from time to market perspective, especially with the rapid increase of diverse accelerators. In this paper, we present a portable compilation framework, called LISA, that can be tuned automatically to generate quality mapping for varied spatial accelerators. Our key contribution is to automatically identify the impact of the dataflow graph (DFG) structure characteristics (representing an application) on the mapping for a new accelerator. Towards this end, we abstract the DFG structure in graph attributes, use Graph Neural Network (GNN) to analyze the graph attributes, and identify the mapping impact for an accelerator architecture with an all-encompassing global view. Finally, we augment a simulated annealing-based mapping approach to take into account the impact of DFG structure in guiding the placement of the dataflow graph nodes and the routing of the dependencies on the accelerator. Our experimental evaluation concretely demonstrates the substantial benefit of our approach compared to the state-of-the-art solutions. Zhaoying Li 0004, Dhananjaya Wijerathne, Tulika Mitra |
HPCA | 1 |
| 2022 | Power-Performance Characterization of TinyML SystemsabstractTinyML systems are enabling machine learning (ML) inference at the edge. However, there exists little quantitative analysis of such systems. This paper presents a systematic performance and power characterization of diverse TinyML applications on micro-controllers (MCUs), spanning neural network models, software libraries, operating systems, and hardware architectures. We focus on the impact of the multiple layers of abstractions that provide higher programmability at the expense of performance and energy efficiency. We propose a model to estimate the costs of different abstraction layers and make recommendations for minimizing those costs. Our findings can help designers with Neural Architecture Search (NAS) and CNN inference optimization on edge devices. Yujie Zhang 0007, Dhananjaya Wijerathne, Zhaoying Li 0004, Tulika Mitra |
ICCD | 3 |
| 2022 | ChordMap: Automated Mapping of Streaming Applications Onto CGRAabstractStreaming applications, consisting of several communicating kernels, are ubiquitous in the embedded computing systems. The synchronous data flow (SDF) is commonly used to capture the complex communication patterns among the kernels. The general-purpose processors cannot meet the throughput requirement of the compute-intensive kernels in the current and emerging applications. The coarse-grained reconfigurable arrays (CGRAs) are well-suited to accelerate the individual kernel and the compiler technology is well-developed to support the mapping of a kernel onto a CGRA accelerator. However, the system-level mapping of the entire streaming application onto a resource-constrained CGRA to maximize throughput remains unexplored. We introduce a novel CGRA mapper, calledChordMap, to automatically generate a high-quality mapping of streaming applications represented as SDF onto CGRAs. We propose an optimized spatio-temporal mapping with modulo-scheduling that judiciously employs concurrent execution of multiple kernels to improve parallelism and thereby maximize throughput.ChordMapachieves, on average,$1.74\times $higher throughput across eight streaming applications compared to the state-of-the-art. Zhaoying Li 0004, Dhananjaya Wijerathne, Xianzhang Chen, Anuj Pathania, Tulika Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | HiMap: Fast and Scalable High-Quality Mapping on CGRA via Hierarchical AbstractionabstractCoarse-grained reconfigurable array (CGRA) has emerged as a promising hardware accelerator due to the excellent balance between reconfigurability, performance, and energy efficiency. The performance of a CGRA strongly depends on the existence of a high-quality compiler to map the application kernels on the architecture. Unfortunately, the state-of-the-art compiler technology falls short in generating high-performance mapping within an acceptable compilation time, especially with increasing CGRA size. We proposeHiMap—a fast and scalable CGRA mapping approach—that is also adept at producing close-to-optimal solutions for regular computational kernels prevalent in existing and emerging application domains. The key strategy behindHiMap’s efficiency and scalability is to exploit the regularity in the computation by employing a virtual systolic array (VSA) as an intermediate abstraction layer in a hierarchical mapping.HiMapfirst maps the loop iterations of the kernel onto a VSA and then distills out the unique patterns in the mapping. These unique patterns are subsequently mapped onto subspaces of the physical CGRA. They are arranged together according to the systolic array mapping to create a complete mapping of the kernel. Experimental results confirm thatHiMapcan generate application mappings that hit the performance envelope of the CGRA.HiMapoffers$17.3\times $and$5\times $improvement in performance and energy efficiency of the mappings compared to the state of the art. The compilation time ofHiMapfor near-optimal mappings is less than 15 min for 64$\times $64 CGRA while existing approaches take days to generate inferior mappings. Dhananjaya Wijerathne, Zhaoying Li 0004, Anuj Pathania, Tulika Mitra, Lothar Thiele |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | HiMap: Fast and Scalable High-Quality Mapping on CGRA via Hierarchical Abstractionabstract10.23919/DATE51398.2021.9473916 Dhananjaya Wijerathne, Zhaoying Li 0004, Anuj Pathania, Tulika Mitra, Lothar Thiele |
DATE | 2 |
| 2019 | CASCADE: High Throughput Data Streaming via Decoupled Access-Execute CGRAabstractA Coarse-Grained Reconfigurable Array (CGRA) is a promising high-performance low-power accelerator for compute-intensive loop kernels. While the mapping of the computations on the CGRA is a well-studied problem, bringing the data into the array at a high throughput remains a challenge. A conventional CGRA design involves on-array computations to generate memory addresses for data access undermining the attainable throughput. A decoupled access-execute architecture, on the other hand, isolates the memory access from the actual computations resulting in a significantly higher throughput. We propose a novel decoupled access-execute CGRA design called CASCADE with full architecture and compiler support for high-throughput data streaming from an on-chip multi-bank memory. CASCADE offloads the address computations for the multi-bank data memory access to a custom designed programmable hardware. An end-to-end fully-automated compiler synchronizes the conflict-free movement of data between the memory banks and the CGRA. Experimental evaluations show on average 3× performance benefit and 2.2× performance per watt improvement for CASCADE compared to an iso-area conventional CGRA with a bigger processing array in lieu of a dedicated hardware memory address generation logic. Dhananjaya Wijerathne, Zhaoying Li 0004, Manupa Karunarathne, Anuj Pathania, Tulika Mitra |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Set variation-aware shared LLC management for CPU-GPU heterogeneous architectureabstractHeterogeneous CPU-GPU multiprocessor systems-on-chip (HMPSoC) becomes a popular architecture choice for high performance embedded systems, where shared last-level cache (LLC) management becomes a critical design consideration. We observe that within a sampling period, CPU and GPU may have distinct access behaviors over various LLC sets. In this work, we propose a light-weighted and fined-grained cache management policy to cope with the CPU-GPU access behavior variation among cache sets. In particular, CPU and GPU requests are prioritized disparately in each LLC set during cache block insertion and promotion, based on the per-core utility behaviors and a per-set CPU-GPU miss counter. Experimental results show that our LLC management scheme outperforms the two state-of-the-art schemes TAP-RRIP and LSP by 12.6% and 10.01%, respectively. Zhaoying Li 0004, Lei Ju 0001, Hongjun Dai, Mengying Zhao, Zhiping Jia |
DATE | 1 |