Huanyu Qu

dblp:322/9419 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-0389-0032ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 36% Memory systems · 36% Hardware accelerators and domain-specific architectures · 23%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
1.022024
HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024
SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture · IEEE Trans. Computers 2024
Memory systems › processing-in-memory
near-memory processing
0.812024
SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture · IEEE Trans. Computers 2024
Parallel and multicore computing
parallel programming models and runtimes
0.812024
HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024
Memory systems
processing-in-memory
0.812024
SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture · IEEE Trans. Computers 2024
Parallel and multicore computing
task scheduling
0.812024
HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous multicore processors
0.212024
HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024

Methods — techniques the papers use, named apart from their topics

scheduling search · 1.5roofline model · 1.5just-in-time evaluation · 1.5prototype chip design · 0.8mapping strategy · 0.8
YearPublicationVenuePosition
2026 Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit Computation
abstract
Large Language Models (LLMs) have transformed society, but their computational and energy needs hinder efficient inference. The memory wall, the growing processor-memory speed disparity, remains a critical bottleneck for LLM. While Process-in-Memory (PIM) architectures address this challenge by co-locating computation with memory, achieving 5-20 × higher bandwidth than GPUs, existing scalable PIM solutions face critical trade-offs in flexibility, capacity, and efficiency when handling LLMs' dynamic memory-compute patterns and operator diversity. DRAM-PIM suffers from inter-bank communication overhead despite its vector parallelism. SRAM-PIM offers sub10ns latency for matrix operation but is constrained by limited capacity. This work introduces CompAir, a scalable PIM architecture that integrates DRAM-PIM and SRAM-PIM with hybrid bonding, enabling efficient linear computations while unlocking multi-granularity data pathways. We further develop CompAirNoC, an advanced network-on-chip (NoC) with an embedded arithmetic logic unit that performs non-linear operations during data movement. Such a design offloads the centralized communication bottleneck in the channel level to distributed banks, simultaneously reducing communication overhead and area cost for scalability. Finally, we develop a hierarchical Instruction Set Architecture that ensures both flexibility and programmability of the hybrid PIM. Experiments show CompAir delivers 1.83-7.98× faster prefill and 1.95 - 6.28 × faster decoding versus state-ofthe-art PIM designs, with 3.52 × lower energy than GPU-PIM hybrids. This work presents the first systematic exploration of hybrid DRAM-PIM and SRAM-PIM architectures with innetwork computation, paving the way towards a scalable PIM system for LLM inference.
Songchen Ma, Huanyu Qu, Jia Chen 0032, Junfeng Lin, Fengbin Tu
ISCA3
2024 HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks
abstract
The rapid development of deep learning has propelled many real-world artificial intelligence applications. Many of these applications integrate multiple neural networks (multi-NN) to cater to various functionalities. There are two challenges of multi-NN acceleration: (1) competition for shared resources becomes a bottleneck, and (2) heterogeneous workloads exhibit remarkably different computing-memory characteristics and various synchronization requirements. Therefore, resource isolation and fine-grained resource allocation for each task are two fundamental requirements for multi-NN computing systems. Although a number of multi-NN acceleration technologies have been explored, few can completely fulfill both of these requirements, especially for mobile scenarios. This paper reports a Hierarchical Asynchronous Parallel Model (HASP) to enhance multi-NN performance to meet both requirements. HASP can be implemented on a multicore processor that adopts Multiple Instruction Multiple Data (MIMD) or Single Instruction Multiple Thread (SIMT) architectures, with minor adaptive modification needed. Further, a prototype chip is developed to validate the hardware effectiveness of this design. A corresponding mapping strategy is also developed, allowing the proposed architecture to simultaneously promote resource utilization and throughput. With the same workload, the prototype chip demonstrates 3.62$\boldsymbol{\times}$, and 3.51$\boldsymbol{\times}$higher throughput over Planaria and 8.68$\boldsymbol{\times}$, 2.61$\boldsymbol{\times}$over Jetson AGX Orin for MobileNet-V1 and ResNet50, respectively.
Songchen Ma, Taoyi Wang, Guanrui Wang, Chenhang Song, Huanyu Qu, Junfeng Lin, Jing Pei
IEEE Trans. Computers7
2024 SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture
abstract
Building hybrid systems that incorporate various processing-in-memory (PIM) devices and processing-near-memory (PNM) technologies can offer complementary advantages in both efficiency and flexibility, while many-core architectures show great potential in deploying data-centric parallel applications with high performance. Compilers for the hybrid PN/IM architecture are critical for enabling such computing systems to be put into practical use. However, most of the existing neural network compilers for PIM or PNM are optimized from the perspective of an operator, and cannot effectively take advantage of a decentralized core-level dataflow with large on-chip memory access bandwidth. Here, we propose a full-stack System-on-graph Compiler (SongC) framework for many-core architecture, which optimizes the efficiency of the PIM devices and leverages the flexibility of the PNM architectures. SongC establishes multi-level graph abstractions to clarify the critical deployment challenges at different levels and generalizes the standard optimizations, decoupling versatile algorithms and diverse types of hardware. To handle the complexity of many-core resource utilization, we also establish a simulation-compilation interaction flow, including a just-in-time evaluator to boost the scheduling search and an extended Roofline model, referred to as the Palace model, to guide the search. Experiments demonstrate the various optimizations and overall performance of SongC and reveal the capability of strategy exploration.
Junfeng Lin, Huanyu Qu, Songchen Ma, Xinglong Ji, Xiaochuan Li 0003, Chenhang Song
IEEE Trans. Computers2