VLDB 2026 Research / reviewers in the wild / expert
Qunyou Liu
dblp:389/0331
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0002-7410-502XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 43% Hardware accelerators and domain-specific architectures · 21% Performance modeling and evaluation · 12% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
data movement |
1.0 | 1 | 2026 | Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026 |
Hardware accelerators and domain-specific architectures
systolic array |
1.0 | 1 | 2026 | Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
transformer inference accelerator |
1.0 | 1 | 2026 | Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026 |
Performance modeling and evaluation
simulation |
0.9 | 1 | 2025 | Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators · DAC 2025 |
Electronic design automation
system-level simulation |
0.9 | 1 | 2025 | Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators · DAC 2025 |
Memory systems › memory management › virtual memory
address translation |
0.8 | 1 | 2024 | Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.8 | 1 | 2024 | Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024 |
Memory systems › memory management › virtual memory › address translation › TLB
TLB miss reduction |
0.8 | 1 | 2024 | Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024 |
Memory systems › memory management
virtual memory |
0.8 | 1 | 2024 | Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024 |
Performance modeling and evaluation › simulation › architectural simulation
full-system simulation |
0.3 | 1 | 2026 | Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design · ACM Trans. Archit. Code Optim. 2026 |
Integrated circuit design
interconnect |
0.3 | 1 | 2025 | Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators · DAC 2025 |
GPUs and heterogeneous computing › heterogeneous architecture
accelerated processing unit |
0.2 | 1 | 2024 | Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloads · ACM Trans. Archit. Code Optim. 2024 |
Methods — techniques the papers use, named apart from their topics
gem5 simulation · 1.9DMA pipelining · 1.0matrix multiplication accelerator · 0.9simulation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-DesignabstractTransformers have revolutionized AI in natural language processing and computer vision, but their enormous computation and memory demands pose significant challenges for hardware acceleration. In practice, end-to-end throughput is often limited by paged data movement and interconnect bandwidth, not just raw MAC count. This work proposes a unified system–accelerator co-design approach to efficiently accelerate transformer inference by jointly optimizing a novel hardware matrix accelerator and its system integration with paged, streaming dataflows and explicit overlap of compute and transfer. On the hardware side, we introduce MatrixFlow , a loosely‐coupled 16× 16 systolic‐array accelerator featuring a block‐based matrix multiplication method that is page-aligned (4 KB tiles), uses only a small (≈20 KB) on-chip buffer, and runs a pipelined schedule of DMA, compute, and DMA-out to fully utilize interconnect bandwidth, emphasizing standard DMA-driven streaming rather than large on-chip reuse. On the system side, we develop Gem5‐AcceSys , an extension of the gem5 full system simulator allowing exploration of standard interconnects (PCIe) and configurable memory hierarchies including Direct-Memory (DM), Direct-Cache (DC), and Device-Memory (DevMem) modes with SMMU/TLB effects. Through co‐design, MatrixFlow’s novel dataflow and the Gem5‐AcceSys platform are tuned in tandem to alleviate data‐movement bottlenecks without requiring specialized CPU‐instruction‐set modifications. We validate our approach with gem5 simulations on representative transformer models (BERT and ViT) across multiple data types and system setups. Results demonstrate up to 22× speed‐up in end‐to‐end inference over a CPU‐only baseline and performance gains of 5×–8× over state‐of‐the‐art loosely‐ and tightly‐coupled accelerators. Furthermore, we show that a standard PCIe-based host memory design can achieve ∼80% of the performance of on‐device HBM memory. Overall, paged streaming and pipeline overlap , not large local SRAMs, emerge as the most effective knobs for efficient transformer inference under realistic system constraints. Qunyou Liu, Marina Zapater, David Atienza 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel AcceleratorsabstractThe growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelerators (DSAs), increasingly essential. These accelerators enhance ML and image processing tasks by offloading computation from the CPU to dedicated hardware. These accelerators rely on interconnects for efficient data transfer, making interconnect design crucial for system-level performance. This paper introduces Gem5-AcceSys, an innovative framework for system-level exploration of standard interconnects and configurable memory hierarchies. Using a matrix multiplication accelerator tailored for transformer workloads as a case study, we evaluate PCIe performance across diverse memory types (DDR4, DDR5, GDDR6, HBM2) and configurations, including host-side and device-side memory. Our findings demonstrate that optimized interconnects can achieve up to $80 \%$ of device-side memory performance and, in some scenarios, even surpass it. These results offer actionable insights for system architects, enabling a balanced approach to performance and cost in next-generation accelerator design. Qunyou Liu, Marina Zapater, David Atienza 0001 |
DAC | 1 |
| 2024 | Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloadsabstractThe increasing demand for computing power and the emergence of heterogeneous computing architectures have driven the exploration of innovative techniques to address current limitations in both the compute and memory subsystems. One such solution is the use of Accelerated Processing Units (APUs), processors that incorporate both a central processing unit (CPU) and an integrated graphics processing unit (iGPU). However, the performance of both APU and CPU systems can be significantly hampered by address translation overhead, leading to a decline in overall performance, especially for cache-resident workloads. To address this issue, we propose the introduction of a new intermediate address space (IAS) in both APU and CPU systems. IAS serves as a bridge between virtual address (VA) spaces and physical address (PA) spaces, optimizing the address translation process. In the case of APU systems, our research indicates that the iGPU suffers from significant translation look-aside buffer (TLB) misses in certain workload situations. Using an IAS, we can divide the initial address translation into front- and back-end phases, effectively shifting the bottleneck in address translation from the cache side to the memory controller side, a technique that proves to be effective for cache-resident workloads. Our simulations demonstrate that implementing IAS in the CPU system can boost performance by up to 40% compared to conventional CPU systems. Furthermore, we evaluate the effectiveness of APU systems, comparing the performance of IAS-based systems with traditional systems, showing up to a 185% improvement in APU system performance with our proposed IAS implementation. Furthermore, our analysis indicates that over 90% of TLB misses can be filtered by the cache, and employing a larger cache within the system could potentially result in even greater improvements. The proposed IAS offers a promising and practical solution to enhance the performance of both APU and CPU systems, contributing to state-of-the-art research in the field of computer architecture. Qunyou Liu, Darong Huang 0003, Luis Costero, Marina Zapater, David Atienza 0001 |
ACM Trans. Archit. Code Optim. | 1 |