Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Qunyou Liu

dblp:389/0331 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0002-7410-502XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 43% Hardware accelerators and domain-specific architectures · 21% Performance modeling and evaluation · 12%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
data movement
1.012026
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026
Hardware accelerators and domain-specific architectures
systolic array
1.012026
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
transformer inference accelerator
1.012026
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026
Performance modeling and evaluation
simulation
0.912025
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators · DAC 2025
Electronic design automation
system-level simulation
0.912025
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators · DAC 2025
Memory systems › memory management › virtual memory
address translation
0.812024
Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024
Memory systems › memory management › virtual memory › address translation
TLB
0.812024
Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024
Memory systems › memory management › virtual memory › address translation › TLB
TLB miss reduction
0.812024
Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024
Memory systems › memory management
virtual memory
0.812024
Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024
Performance modeling and evaluation › simulation › architectural simulation
full-system simulation
0.312026
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026
Integrated circuit design
interconnect
0.312025
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators · DAC 2025
GPUs and heterogeneous computing › heterogeneous architecture
accelerated processing unit
0.212024
Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024

Methods — techniques the papers use, named apart from their topics

gem5 simulation · 1.9DMA pipelining · 1.0matrix multiplication accelerator · 0.9simulation · 0.8
YearPublicationVenuePosition
2026 Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
abstract
Transformers have revolutionized AI in natural language processing and computer vision, but their enormous computation and memory demands pose significant challenges for hardware acceleration. In practice, end-to-end throughput is often limited by paged data movement and interconnect bandwidth, not just raw MAC count. This work proposes a unified system–accelerator co-design approach to efficiently accelerate transformer inference by jointly optimizing a novel hardware matrix accelerator and its system integration with paged, streaming dataflows and explicit overlap of compute and transfer. On the hardware side, we introduce MatrixFlow , a loosely‐coupled 16× 16 systolic‐array accelerator featuring a block‐based matrix multiplication method that is page-aligned (4 KB tiles), uses only a small (≈20 KB) on-chip buffer, and runs a pipelined schedule of DMA, compute, and DMA-out to fully utilize interconnect bandwidth, emphasizing standard DMA-driven streaming rather than large on-chip reuse. On the system side, we develop Gem5‐AcceSys , an extension of the gem5 full system simulator allowing exploration of standard interconnects (PCIe) and configurable memory hierarchies including Direct-Memory (DM), Direct-Cache (DC), and Device-Memory (DevMem) modes with SMMU/TLB effects. Through co‐design, MatrixFlow’s novel dataflow and the Gem5‐AcceSys platform are tuned in tandem to alleviate data‐movement bottlenecks without requiring specialized CPU‐instruction‐set modifications. We validate our approach with gem5 simulations on representative transformer models (BERT and ViT) across multiple data types and system setups. Results demonstrate up to 22× speed‐up in end‐to‐end inference over a CPU‐only baseline and performance gains of 5×–8× over state‐of‐the‐art loosely‐ and tightly‐coupled accelerators. Furthermore, we show that a standard PCIe-based host memory design can achieve ∼80% of the performance of on‐device HBM memory. Overall, paged streaming and pipeline overlap , not large local SRAMs, emerge as the most effective knobs for efficient transformer inference under realistic system constraints.
Qunyou Liu, Marina Zapater, David Atienza 0001
ACM Trans. Archit. Code Optim.1
2025 Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
abstract
The growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelerators (DSAs), increasingly essential. These accelerators enhance ML and image processing tasks by offloading computation from the CPU to dedicated hardware. These accelerators rely on interconnects for efficient data transfer, making interconnect design crucial for system-level performance. This paper introduces Gem5-AcceSys, an innovative framework for system-level exploration of standard interconnects and configurable memory hierarchies. Using a matrix multiplication accelerator tailored for transformer workloads as a case study, we evaluate PCIe performance across diverse memory types (DDR4, DDR5, GDDR6, HBM2) and configurations, including host-side and device-side memory. Our findings demonstrate that optimized interconnects can achieve up to $80 \%$ of device-side memory performance and, in some scenarios, even surpass it. These results offer actionable insights for system architects, enabling a balanced approach to performance and cost in next-generation accelerator design.
Qunyou Liu, Marina Zapater, David Atienza 0001
DAC1
2024 Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads
abstract
The increasing demand for computing power and the emergence of heterogeneous computing architectures have driven the exploration of innovative techniques to address current limitations in both the compute and memory subsystems. One such solution is the use of Accelerated Processing Units (APUs), processors that incorporate both a central processing unit (CPU) and an integrated graphics processing unit (iGPU). However, the performance of both APU and CPU systems can be significantly hampered by address translation overhead, leading to a decline in overall performance, especially for cache-resident workloads. To address this issue, we propose the introduction of a new intermediate address space (IAS) in both APU and CPU systems. IAS serves as a bridge between virtual address (VA) spaces and physical address (PA) spaces, optimizing the address translation process. In the case of APU systems, our research indicates that the iGPU suffers from significant translation look-aside buffer (TLB) misses in certain workload situations. Using an IAS, we can divide the initial address translation into front- and back-end phases, effectively shifting the bottleneck in address translation from the cache side to the memory controller side, a technique that proves to be effective for cache-resident workloads. Our simulations demonstrate that implementing IAS in the CPU system can boost performance by up to 40% compared to conventional CPU systems. Furthermore, we evaluate the effectiveness of APU systems, comparing the performance of IAS-based systems with traditional systems, showing up to a 185% improvement in APU system performance with our proposed IAS implementation. Furthermore, our analysis indicates that over 90% of TLB misses can be filtered by the cache, and employing a larger cache within the system could potentially result in even greater improvements. The proposed IAS offers a promising and practical solution to enhance the performance of both APU and CPU systems, contributing to state-of-the-art research in the field of computer architecture.
Qunyou Liu, Darong Huang 0003, Luis Costero, Marina Zapater, David Atienza 0001
ACM Trans. Archit. Code Optim.1