EDBT 2026 Demo / reviewers in the wild / expert
Huanyu Qu
dblp:322/9419
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-0389-0032ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 36% Memory systems · 36% Hardware accelerators and domain-specific architectures · 23% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
1.0 | 2 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture · IEEE Trans. Computers 2024 |
Memory systems › processing-in-memory
near-memory processing |
0.8 | 1 | 2024 | SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture · IEEE Trans. Computers 2024 |
Parallel and multicore computing
parallel programming models and runtimes |
0.8 | 1 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 |
Memory systems
processing-in-memory |
0.8 | 1 | 2024 | SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture · IEEE Trans. Computers 2024 |
Parallel and multicore computing
task scheduling |
0.8 | 1 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous multicore processors |
0.2 | 1 | 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks · IEEE Trans. Computers 2024 |
Methods — techniques the papers use, named apart from their topics
scheduling search · 1.5roofline model · 1.5just-in-time evaluation · 1.5prototype chip design · 0.8mapping strategy · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit ComputationabstractLarge Language Models (LLMs) have transformed society, but their computational and energy needs hinder efficient inference. The memory wall, the growing processor-memory speed disparity, remains a critical bottleneck for LLM. While Process-in-Memory (PIM) architectures address this challenge by co-locating computation with memory, achieving 5-20 × higher bandwidth than GPUs, existing scalable PIM solutions face critical trade-offs in flexibility, capacity, and efficiency when handling LLMs' dynamic memory-compute patterns and operator diversity. DRAM-PIM suffers from inter-bank communication overhead despite its vector parallelism. SRAM-PIM offers sub10ns latency for matrix operation but is constrained by limited capacity. This work introduces CompAir, a scalable PIM architecture that integrates DRAM-PIM and SRAM-PIM with hybrid bonding, enabling efficient linear computations while unlocking multi-granularity data pathways. We further develop CompAirNoC, an advanced network-on-chip (NoC) with an embedded arithmetic logic unit that performs non-linear operations during data movement. Such a design offloads the centralized communication bottleneck in the channel level to distributed banks, simultaneously reducing communication overhead and area cost for scalability. Finally, we develop a hierarchical Instruction Set Architecture that ensures both flexibility and programmability of the hybrid PIM. Experiments show CompAir delivers 1.83-7.98× faster prefill and 1.95 - 6.28 × faster decoding versus state-ofthe-art PIM designs, with 3.52 × lower energy than GPU-PIM hybrids. This work presents the first systematic exploration of hybrid DRAM-PIM and SRAM-PIM architectures with innetwork computation, paving the way towards a scalable PIM system for LLM inference. Songchen Ma, Huanyu Qu, Jia Chen 0032, Junfeng Lin, Fengbin Tu |
ISCA | 3 |
| 2024 | HASP: Hierarchical Asynchronous Parallelism for Multi-NN TasksabstractThe rapid development of deep learning has propelled many real-world artificial intelligence applications. Many of these applications integrate multiple neural networks (multi-NN) to cater to various functionalities. There are two challenges of multi-NN acceleration: (1) competition for shared resources becomes a bottleneck, and (2) heterogeneous workloads exhibit remarkably different computing-memory characteristics and various synchronization requirements. Therefore, resource isolation and fine-grained resource allocation for each task are two fundamental requirements for multi-NN computing systems. Although a number of multi-NN acceleration technologies have been explored, few can completely fulfill both of these requirements, especially for mobile scenarios. This paper reports a Hierarchical Asynchronous Parallel Model (HASP) to enhance multi-NN performance to meet both requirements. HASP can be implemented on a multicore processor that adopts Multiple Instruction Multiple Data (MIMD) or Single Instruction Multiple Thread (SIMT) architectures, with minor adaptive modification needed. Further, a prototype chip is developed to validate the hardware effectiveness of this design. A corresponding mapping strategy is also developed, allowing the proposed architecture to simultaneously promote resource utilization and throughput. With the same workload, the prototype chip demonstrates 3.62$\boldsymbol{\times}$, and 3.51$\boldsymbol{\times}$higher throughput over Planaria and 8.68$\boldsymbol{\times}$, 2.61$\boldsymbol{\times}$over Jetson AGX Orin for MobileNet-V1 and ResNet50, respectively. Songchen Ma, Taoyi Wang, Guanrui Wang, Chenhang Song, Huanyu Qu, Junfeng Lin, Jing Pei |
IEEE Trans. Computers | 7 |
| 2024 | SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core ArchitectureabstractBuilding hybrid systems that incorporate various processing-in-memory (PIM) devices and processing-near-memory (PNM) technologies can offer complementary advantages in both efficiency and flexibility, while many-core architectures show great potential in deploying data-centric parallel applications with high performance. Compilers for the hybrid PN/IM architecture are critical for enabling such computing systems to be put into practical use. However, most of the existing neural network compilers for PIM or PNM are optimized from the perspective of an operator, and cannot effectively take advantage of a decentralized core-level dataflow with large on-chip memory access bandwidth. Here, we propose a full-stack System-on-graph Compiler (SongC) framework for many-core architecture, which optimizes the efficiency of the PIM devices and leverages the flexibility of the PNM architectures. SongC establishes multi-level graph abstractions to clarify the critical deployment challenges at different levels and generalizes the standard optimizations, decoupling versatile algorithms and diverse types of hardware. To handle the complexity of many-core resource utilization, we also establish a simulation-compilation interaction flow, including a just-in-time evaluator to boost the scheduling search and an extended Roofline model, referred to as the Palace model, to guide the search. Experiments demonstrate the various optimizations and overall performance of SongC and reveal the capability of strategy exploration. Junfeng Lin, Huanyu Qu, Songchen Ma, Xinglong Ji, Xiaochuan Li 0003, Chenhang Song |
IEEE Trans. Computers | 2 |