Yaosheng Fu

dblp:147/5196 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Hardware accelerators and domain-specific architectures · 22% Performance modeling and evaluation · 20% Processor architecture and microarchitecture · 19%
Artificial intelligence
2 papers
Efficient and distributed learning · 36% Language models and text generation · 36% Deep learning architectures and training · 28%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › inference efficiency
inference optimization
0.912025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.912025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
KV cache compression
0.912025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Memory systems
cache coherence
0.722020
BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA Research · ASPLOS 2020
Coherence domain restriction on large scale systems · MICRO 2015
Processor architecture and microarchitecture
multicore design
0.632020
OpenPiton: An Open Source Manycore Research Framework · ASPLOS 2016
Coherence domain restriction on large scale systems · MICRO 2015
BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA Research · ASPLOS 2020
Processor architecture and microarchitecture
many-core architecture
0.622018
Power and Energy Characterization of an Open Source 25-Core Manycore Processor · HPCA 2018
OpenPiton: An Open Source Manycore Research Framework · ASPLOS 2016
GPUs and heterogeneous computing
GPU architecture
0.612022
GPU Domain Specialization via Composable On-Package Architecture · ACM Trans. Archit. Code Optim. 2022
Electronic design automation
design space exploration
0.512021
Need for Speed: Experiences Building a Trustworthy System-Level GPU Simulator · HPCA 2021
Performance modeling and evaluation › simulation › architectural simulation
GPU architecture simulation
0.512021
Need for Speed: Experiences Building a Trustworthy System-Level GPU Simulator · HPCA 2021
GPUs and heterogeneous computing
GPU microarchitecture
0.512021
Need for Speed: Experiences Building a Trustworthy System-Level GPU Simulator · HPCA 2021
Performance modeling and evaluation
simulation
0.512021
Need for Speed: Experiences Building a Trustworthy System-Level GPU Simulator · HPCA 2021
Performance modeling and evaluation
workload characterization
0.522021
Power and Energy Characterization of an Open Source 25-Core Manycore Processor · HPCA 2018
Need for Speed: Experiences Building a Trustworthy System-Level GPU Simulator · HPCA 2021
Processor architecture and microarchitecture › multiprocessor architecture
heterogeneous-ISA
0.412020
BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA Research · ASPLOS 2020
Energy-efficient computing
power characterization
0.312018
Power and Energy Characterization of an Open Source 25-Core Manycore Processor · HPCA 2018
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.312025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Machine learning › Deep learning architectures and training
transformer
0.312025
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression · ICML 2025
Performance modeling and evaluation › parallel performance evaluation
multicore scalability
0.212015
Coherence domain restriction on large scale systems · MICRO 2015
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.212022
GPU Domain Specialization via Composable On-Package Architecture · ACM Trans. Archit. Code Optim. 2022
Reconfigurable computing and FPGAs
FPGA prototyping
0.112016
OpenPiton: An Open Source Manycore Research Framework · ASPLOS 2016
Memory systems › cache coherence
directory-based coherence
0.112015
Coherence domain restriction on large scale systems · MICRO 2015

Methods — techniques the papers use, named apart from their topics

sparse attention · 1.7dimensionality reduction · 1.7KV cache eviction · 1.7trace-based simulation · 1.0component fidelity modeling · 1.0composable on-package architecture · 0.6FPGA prototyping · 0.4simulation · 0.3measurement · 0.3ASIC synthesis · 0.2
YearPublicationVenuePosition
2025 RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
abstract
Transformer-based Large Language Models rely critically on the KV cache to efficiently handle extended contexts during the decode phase. Yet, the size of the KV cache grows proportionally with the input length, burdening both memory bandwidth and capacity as decoding progresses. To address this challenge, we present RocketKV, a training-free KV cache compression strategy containing two consecutive stages. In the first stage, it performs coarse-grain permanent KV cache eviction on the input sequence tokens. In the second stage, it adopts a hybrid sparse attention method to conduct fine-grain top-k sparse attention, approximating the attention scores by leveraging both head and sequence dimensionality reductions. We show that RocketKV provides a compression ratio of up to 400×, end-to-end speedup of up to 3.7× as well as peak memory reduction of up to 32.6% in the decode phase on an NVIDIA A100 GPU compared to the full KV cache baseline, while achieving negligible accuracy loss on a variety of long-context tasks. We also propose a variant of RocketKV for multi-turn scenarios, which consistently outperforms other existing methods and achieves accuracy nearly on par with an oracle top-k attention scheme. The source code is available here: https://github.com/NVlabs/RocketKV.
Payman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai, Zhiding Yu, Alexey Tumanov
ICML2
2022 GPU Domain Specialization via Composable On-Package Architecture
abstract
As GPUs scale their low-precision matrix math throughput to boost deep learning (DL) performance, they upset the balance between math throughput and memory system capabilities. We demonstrate that a converged GPU design trying to address diverging architectural requirements between FP32 (or larger)-based HPC and FP16 (or smaller)-based DL workloads results in sub-optimal configurations for either of the application domains. We argue that a C omposable O n- PA ckage GPU (COPA-GPU) architecture to provide domain-specialized GPU products is the most practical solution to these diverging requirements. A COPA-GPU leverages multi-chip-module disaggregation to support maximal design reuse, along with memory system specialization per application domain. We show how a COPA-GPU enables DL-specialized products by modular augmentation of the baseline GPU architecture with up to 4× higher off-die bandwidth, 32× larger on-package cache, and 2.3× higher DRAM bandwidth and capacity, while conveniently supporting scaled-down HPC-oriented designs. This work explores the microarchitectural design necessary to enable composable GPUs and evaluates the benefits composability can provide to HPC, DL training, and DL inference. We show that when compared to a converged GPU design, a DL-optimized COPA-GPU featuring a combination of 16× larger cache capacity and 1.6× higher DRAM bandwidth scales per-GPU training and inference performance by 31% and 35%, respectively, and reduces the number of GPU instances by 50% in scale-out training scenarios.
Yaosheng Fu, Evgeny Bolotin, Niladrish Chatterjee, David W. Nellans, Stephen W. Keckler
ACM Trans. Archit. Code Optim.1
2021 Need for Speed: Experiences Building a Trustworthy System-Level GPU Simulator
abstract
The demands of high-performance computing (HPC) and machine learning (ML) workloads have resulted in the rapid architectural evolution of GPUs over the last decade. The growing memory footprint and diversity of data types in these workloads has required GPUs to embrace micro-architectural heterogeneity and increased memory system sophistication to scale performance. Effective simulation of new architectural features early in the design cycle enables quick and effective exploration of design trade-offs across this increasingly diverse set of workloads. This work provides a retrospective on the design and development of NVArchSim (NVAS), an architectural simulator used within NVIDIA to design and evaluate features that are difficult to appraise using other methodologies due to workload type, size, complexity, or lack of modeling flexibility. We argue that overly precise and/or overly slow architectural models hamper an architect's ability to evaluate new features within a reasonable time frame, hurting productivity. Because of its speed, NVAS is being used to trace and evaluate hundreds of HPC and state-of-the-art ML workloads on single-GPU or multi-GPU systems. By adding component fidelity only when necessary to improve system-level modeling accuracy, NVAS delivers simulation speed orders of magnitude higher than most publicly available GPU simulators while retaining high levels of accuracy and simulation flexibility. Building trustworthy high-level simulation platforms is a difficult exercise in balance and compromise; we share our experiences to help and encourage those in academia who take on the challenge of building GPU simulation platforms.
Oreste Villa, Daniel Lustig, Zi Yan, Evgeny Bolotin, Yaosheng Fu, Niladrish Chatterjee, Nan Jiang 0009, David W. Nellans
HPCA5
2020 BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA Research
abstract
Heterogeneous architectures and heterogeneous-ISA designs are growing areas of computer architecture and system software research. Unfortunately, this line of research is significantly hindered by the lack of experimental systems and modifiable hardware frameworks. This work proposes BYOC, a "Bring Your Own Core" framework that is specifically designed to enable heterogeneous-ISA and heterogeneous system research. BYOC is an open-source hardware framework that provides a scalable cache coherence system, that includes out-of-the-box support for four different ISAs (RISC-V 32-bit, RISC-V 64-bit, x86, and SPARCv9) and has been connected to ten different cores. The framework also supports multiple loosely coupled accelerators and is a fully working system supporting SMP Linux. The Transaction-Response Interface (TRI) introduced with BYOC has been specifically designed to make it easy to add in new cores with new ISAs and memory interfaces. This work demonstrates multiple multi-ISA designs running on FPGA and characterises the communication costs. This work describes many of the architectural design trade-offs for building such a flexible system. BYOC is well suited to be the premiere platform for heterogeneous-ISA architecture, system software, and compiler research.
Jonathan Balkind, Katie Lim, Michael Schaffner, Fei Gao 0016, Grigory Chirkov, Ang Li 0011, Alexey Lavrov, Tri M. Nguyen 0002, Yaosheng Fu, Florian Zaruba, Kunal Gulati, Luca Benini, David Wentzlaff
ASPLOS9
2018 Power and Energy Characterization of an Open Source 25-Core Manycore Processor
abstract
The end of Dennard's scaling and the looming power wall have made power and energy primary design goals for modern processors. Further, new applications such as cloud computing and Internet of Things (IoT) continue to necessitate increased performance and energy efficiency. Manycore processors show potential in addressing some of these issues. However, there is little detailed power and energy data on manycore processors. In this work, we carefully study detailed power and energy characteristics of Piton, a 25-core modern open source academic processor, including voltage versus frequency scaling, energy per instruction (EPI), memory system energy, network-on-chip (NoC) energy, thermal characteristics, and application performance and power consumption. This is the first detailed power and energy characterization of an open source manycore design implemented in silicon. The open source nature of the processor provides increased value, enabling detailed characterization verified against simulation and the ability to correlate results with the design and register transfer level (RTL) model. Additionally, this enables other researchers to utilize this work to build new power models, devise new research directions, and perform accurate power and energy research using the open source processor. The characterization data reveals a number of interesting insights, including that operand values have a large impact on EPI, recomputing data can be more energy efficient than loading it from memory, on-chip data transmission (NoC) energy is low, and insights on energy efficient multithreaded core design. All data collected and the hardware infrastructure used is open source and available for download at http://www.openpiton.org.
Michael McKeown, Alexey Lavrov, Mohammad Shahrad, Paul J. Jackson, Yaosheng Fu, Jonathan Balkind, Tri Minh Nguyen 0003, Katie Lim, Yanqi Zhou, David Wentzlaff
HPCA5
2016 OpenPiton: An Open Source Manycore Research Framework
abstract
Industry is building larger, more complex, manycore processors on the back of strong institutional knowledge, but academic projects face difficulties in replicating that scale. To alleviate these difficulties and to develop and share knowledge, the community needs open architecture frameworks for simulation, synthesis, and software exploration which support extensibility, scalability, and configurability, alongside an established base of verification tools and supported software. In this paper we present OpenPiton, an open source framework for building scalable architecture research prototypes from 1 core to 500 million cores. OpenPiton is the world's first open source, general-purpose, multithreaded manycore processor and framework. OpenPiton leverages the industry hardened OpenSPARC T1 core with modifications and builds upon it with a scratch-built, scalable uncore creating a flexible, modern manycore design. In addition, OpenPiton provides synthesis and backend scripts for ASIC and FPGA to enable other researchers to bring their designs to implementation. OpenPiton provides a complete verification infrastructure of over 8000 tests, is supported by mature software tools, runs full-stack multiuser Debian Linux, and is written in industry standard Verilog. Multiple implementations of OpenPiton have been created including a taped-out 25-core implementation in IBM's 32nm process and multiple Xilinx FPGA prototypes.
Jonathan Balkind, Michael McKeown, Yaosheng Fu, Tri Minh Nguyen 0003, Yanqi Zhou, Alexey Lavrov, Mohammad Shahrad, Adi Fuchs, Samuel Payne, Xiaohua Liang, Matthew Matl, David Wentzlaff
ASPLOS3
2016 Piton: A 25-core academic manycore research processor
abstract
Presents a collection of slides covering the following: many-core processors; cloud computing; data warehouses; and data centers.
Michael McKeown, Yaosheng Fu, Tri Minh Nguyen 0003, Yanqi Zhou, Jonathan Balkind, Alexey Lavrov, Mohammad Shahrad, Samuel Payne, David Wentzlaff
Hot Chips Symposium2
2015 Coherence domain restriction on large scale systems
abstract
Designing massive scale cache coherence systems has been an elusive goal. Whether it be on large-scale GPUs, future thousand-core chips, or across million-core warehouse scale computers, having shared memory, even to a limited extent, improves programmability. This work sidesteps the traditional challenges of creating massively scalable cache coherence by restricting coherence to flexible subsets (domains) of a system's total cores and home nodes. This paper proposes Coherence Domain Restriction (CDR), a novel coherence framework that enables the creation of thousand to million core systems that use shared memory while maintaining low storage and energy overhead. Inspired by the observation that the majority of cache lines are only shared by a subset of cores either due to limited application parallelism or limited page sharing, CDR restricts the coherence domain from global cache coherence to VM-level, application-level, or page-level. We explore two types of restriction, one which limits the total number of sharers that can access a coherence domain and one which limits the number and location of home nodes that partake in a coherence domain. Each independent coherence domain only tracks the cores in its domain instead of the whole system, thereby removing the need for a coherence scheme built on top of CDR to scale. Sharer Restriction achieves constant storage overhead as core count increases while Home Restriction provides localized communication enabling higher performance. Unlike previous systems, CDR is flexible and does not restrict the location of the home nodes or sharers within a domain. We evaluate CDR in the context of a 1024-core chip and in the novel application of shared memory to a 1,000,000-core warehouse scale computer. Sharer Restriction results in significant area savings, while Home Restriction in the 1024-core chip and 1,000,000-core system increases performance by 29% and 23.04x respectively when comparing with global home placement. We implemented the entire CDR framework in a 25-core processor taped out in IBM's 32nm SOI process and present a detailed area characterization.
Yaosheng Fu, Tri Minh Nguyen 0003, David Wentzlaff
MICRO1
2014 PriME: A parallel and distributed simulator for thousand-core chips
abstract
Modern processors are integrating an increasing number of cores, which brings new design challenges. However, mainstream architectural simulators primarily focus on unicore or multicore systems with small core counts. In order to simulate emerging manycore architectures, more simulators designed for thousand-core systems are needed. In this paper, we introduce the Princeton Manycore Executor (PriME), a parallelized, MPI-based, x86 manycore simulator. The primary goal of PriME is to provide high performance simulation for manycore architectures, allowing fast exploration of architectural ideas including cache hierarchies, coherence protocols and NoCs in thousand-core systems. PriME supports multi-threaded workloads as well as multi-programmed workloads. Furthermore, it parallelizes the simulation of multi-threaded workloads inside of a host machine and parallelizes the simulation of multi-programmed workloads across multiple host machines by utilizing MPI to communicate between different simulator modules. By using two levels of parallelization (within a host machine and across host machines), PriME can improve simulation performance which is especially useful when simulating thousands of cores. Prime is especially adept at executing simulations which have large memory requirements as previous simulators which use multiple machines are unable to simulate more memory than is available in any single host machine. We demonstrate PriME simulating 2000+ core machines and show near-linear scaling on up to 108 host processors split across 9 machines. Finally we validate PriME against a real-world 40-core machine and show the average error to be 12%.
Yaosheng Fu, David Wentzlaff
ISPASS1