Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yaqi Zhang 0001

dblp:76/10181-1 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
2since 2021 · last 2021
0000-0003-3882-282XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Reconfigurable computing and FPGAs · 62% Hardware accelerators and domain-specific architectures · 18% Electronic design automation · 8%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 13 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
1.242020
Gorgon: Accelerating Machine Learning from Relational Data · ISCA 2020
Scalable interconnects for reconfigurable spatial architectures · ISCA 2019
Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable dataflow accelerator
1.022021
Capstan: A Vector RDA for Sparsity · MICRO 2021
SARA: Scaling a Reconfigurable Dataflow Accelerator · ISCA 2021
Reconfigurable computing and FPGAs › FPGA compilation
compiler mapping
0.512021
SARA: Scaling a Reconfigurable Dataflow Accelerator · ISCA 2021
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation
0.512021
Capstan: A Vector RDA for Sparsity · MICRO 2021
Integrated circuit design
interconnect
0.412019
Scalable interconnects for reconfigurable spatial architectures · ISCA 2019
Compilers and program optimization › code generation › parallel code generation
accelerator code generation
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Compilers and program optimization › hardware compilation
high-level synthesis
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Reconfigurable computing and FPGAs
reconfigurable computing
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Electronic design automation
design space exploration
0.212016
Automatic Generation of Efficient Accelerators for Reconfigurable Hardware · ISCA 2016
Reconfigurable computing and FPGAs
FPGA accelerator
0.212016
Automatic Generation of Efficient Accelerators for Reconfigurable Hardware · ISCA 2016
Electronic design automation
high-level synthesis
0.212016
Automatic Generation of Efficient Accelerators for Reconfigurable Hardware · ISCA 2016
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.112021
SARA: Scaling a Reconfigurable Dataflow Accelerator · ISCA 2021
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112017
Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017

Methods — techniques the papers use, named apart from their topics

hyperblock scheduling · 1.0compiler-managed memory consistency · 1.0parallel pattern distillation · 0.9coarse-grained reconfiguration · 0.9out-of-order execution · 0.5declarative programming model · 0.5vector unit datapath design · 0.4parameterized templates · 0.2parallel patterns · 0.2artificial neural network · 0.2
YearPublicationVenuePosition
2021 SARA: Scaling a Reconfigurable Dataflow Accelerator
abstract
The need for speed in modern data-intensive work-loads and the rise of "dark silicon" in the semiconductor industry are pushing for larger, faster, and more energy and area-efficient architectures, such as Reconfigurable Dataflow Accelerators (RDAs). Nevertheless, challenges remain in developing mechanisms to effectively utilize the compute power of these large-scale RDAs. To address these challenges, we present SARA, a compiler that employs a novel mapping strategy to efficiently utilize large-scale RDAs. Starting from a single-threaded imperative abstraction, SARA spatially maps a program onto RDA's distributed resources, exploiting dataflow parallelism within and across hyperblocks to saturate the compute throughput of an RDA. SARA introduces (a) compiler-managed memory consistency (CMMC), a control paradigm that hierarchically pipelines a nested and data-dependent control-flow graph onto a dataflow architecture, and (b) a compilation flow that decomposes the program graph across distributed heterogeneous resources to hide low-level RDA constraints from programmers. Our evaluation shows that SARA achieves close to perfect performance scaling on a recently proposed RDA—Plasticine. Over a mix of deep-learning, graph-processing, and streaming applications, SARA achieves a 1.9× geo-mean speedup over a Tesla V100 GPU using only 12% of the silicon area.
Yaqi Zhang 0001, Nathan Zhang, Tian Zhao 0001, Matthew Vilim, Muhammad Shahbaz 0001, Kunle Olukotun
ISCA1
2021 Capstan: A Vector RDA for Sparsity
abstract
This paper proposes Capstan: a scalable, parallel-patterns-based, reconfigurable dataflow accelerator (RDA) for sparse and dense tensor applications. Instead of designing for one application, we start with common sparse data formats, each of which supports multiple applications. Using a declarative programming model, Capstan supports application-independent sparse iteration and memory primitives that can be mapped to vectorized, high-performance hardware. We optimize random-access sparse memories with configurable out-of-order execution to increase SRAM random-access throughput from 32% to 80%.
Alexander Rucker, Matthew Vilim, Tian Zhao 0001, Yaqi Zhang 0001, Raghu Prabhakar, Kunle Olukotun
MICRO4
2020 Gorgon: Accelerating Machine Learning from Relational Data
abstract
Accelerator deployment in data centers remains limited despite domain-specific architectures' promise of higher performance. Rapidly-changing applications and high nre cost make deploying fixed-function accelerators at scale untenable. More flexible than dsas, fpgas are gaining traction but remain hampered by cumbersome programming models, long synthesis times, and slow clocks. Coarse-grained reconfigurable architectures (cgra) are a compelling alternative and offer efficiency while retaining programmability-by providing general-purpose hardware and communication patterns, a single cgra targets multiple application domains.One emerging application is in-database machine learning: a high-performance, low-friction interface for analytics on large databases. We co-locate database and machine learning processing in a unified reconfigurable data analytics accelerator, Gorgon, which flexibly shares resources between db and ml without compromising performance or incurring excessive overheads in either domain. We distill and integrate database parallel patterns into an existing ML-focused cgra, increasing area by less than 4% while outperforming a multicore software baseline by 1500X. We also explore the performance impact of unifying db and ml in a single accelerator, showing up to 4x speedup over split accelerators.
Matthew Vilim, Alexander Rucker, Yaqi Zhang 0001, Sophia Liu, Kunle Olukotun
ISCA3
2019 Scalable interconnects for reconfigurable spatial architectures
abstract
Recent years have seen the increased adoption of Coarse-Grained Reconfigurable Architectures (CGRAs) as flexible, energy-efficient compute accelerators. Obtaining performance using spatial architectures while supporting diverse applications requires a flexible, high-bandwidth interconnect. Because modern CGRAs support vector units with wide datapaths, designing an interconnect that balances dynamism, communication granularity, and programmability is a challenging task.
Yaqi Zhang 0001, Alexander Rucker, Matthew Vilim, Raghu Prabhakar, William Hwang, Kunle Olukotun
ISCA1
2018 Spatial: a language and compiler for application accelerators
abstract
Industry is increasingly turning to reconfigurable architectures like FPGAs and CGRAs for improved performance and energy efficiency. Unfortunately, adoption of these architectures has been limited by their programming models. HDLs lack abstractions for productivity and are difficult to target from higher level languages. HLS tools are more productive, but offer an ad-hoc mix of software and hardware abstractions which make performance optimizations difficult.
David Koeplinger, Matthew Feldman, Raghu Prabhakar, Yaqi Zhang 0001, Stefan Hadjis, Ruben Fiszel, Tian Zhao 0001, Luigi Nardi, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
PLDI4
2017 Plasticine: A Reconfigurable Architecture For Parallel Paterns
abstract
Reconfigurable architectures have gained popularity in recent years as they allow the design of energy-efficient accelerators. Fine-grain fabrics (e.g. FPGAs) have traditionally suffered from performance and power inefficiencies due to bit-level reconfigurable abstractions. Both fine-grain and coarse-grain architectures (e.g. CGRAs) traditionally require low level programming and suffer from long compilation times. We address both challenges with Plasticine, a new spatially reconfigurable architecture designed to efficiently execute applications composed of parallel patterns. Parallel patterns have emerged from recent research on parallel programming as powerful, high-level abstractions that can elegantly capture data locality, memory access patterns, and parallelism across a wide range of dense and sparse applications.
Raghu Prabhakar, Yaqi Zhang 0001, David Koeplinger, Matthew Feldman, Tian Zhao 0001, Stefan Hadjis, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA2
2016 Automatic Generation of Efficient Accelerators for Reconfigurable Hardware
abstract
Acceleration in the form of customized datapaths offer large performance and energy improvements over general purpose processors. Reconfigurable fabrics such as FPGAs are gaining popularity for use in implementing application-specific accelerators, thereby increasing the importance of having good high-level FPGA design tools. However, current tools for targeting FPGAs offer inadequate support for high-level programming, resource estimation, and rapid and automatic design space exploration. We describe a design framework that addresses these challenges. We introduce a new representation of hardware using parameterized templates that captures locality and parallelism information at multiple levels of nesting. This representation is designed to be automatically generated from high-level languages based on parallel patterns. We describe a hybrid area estimation technique which uses template-level models and design-level artificial neural networks to account for effects from hardware place-and-route tools, including routing overheads, register and block RAM duplication, and LUT packing. Our runtime estimation accounts for off-chip memory accesses. We use our estimation capabilities to rapidly explore a large space of designs across tile sizes, parallelization factors, and optional coarse-grained pipelining, all at multiple loop levels. We show that estimates average 4.8% error for logic resources, 6.1% error for runtimes, and are 279 to 6533 times faster than a commercial high-level synthesis tool. We compare the best-performing designs to optimized CPU code running on a server-grade 6 core processor and show speedups of up to 16.7×.
David Koeplinger, Raghu Prabhakar, Yaqi Zhang 0001, Christina Delimitrou, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA3