EDBT 2026 Demo / reviewers in the wild / expert
Yikan Qiu
dblp:336/0625
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quartet: A Digital Compute-in-Memory Versatile AI Accelerator With Heterogeneous Tensor Engines and Off-Chip-Less DataflowabstractAlthough most AI core operations can be formulated as matrix multiplications (MMs), their characteristics are quite different. In some cases, one input may be a constant weight matrix while the other is a dynamic feature matrix, i.e., FWMM, or both inputs may be dynamic features as FFMM. Furthermore, the data involved can also be sparse, leading to variations like SpFWMM and SpFFMM. To address this challenge, this paper investigates a versatile accelerator architecture for AI algorithms based on heterogeneous tensor engines. For MM operators with varying characteristics, this paper proposes four-quadrant heterogeneous tensor engines to handle FWMM, SpFWMM, FFMM, and SpFFMM, respectively. These four tensor engines are comprised of SRAM-based single address digital compute-in-memory (CIM) array, SRAM-based multi-address digital CIM array, systolic array, and multi-SIMD array, respectively. In addition, to improve the AI execution efficiency, this paper proposes a dual-level multi-issue mechanism to achieve inter-operator and inter-block parallelization along with an off-chip-less dataflow enabled by on-chip unified memory pool. Thanks to the integration of aforementioned innovations, this paper develops a versatile AI acceleration chip Quartet, which achieves exceptional energy efficiency. Specifically, for graph convolutional network on PubMed, it demonstrates a$19.56\times $and$3.47\times $improvement in energy efficiency compared to similar works, ReDCIM and TensorCIM, respectively. Yikan Qiu, Guoxiang Li, Meng Wu 0005, Yifan Jia 0009, Le Ye, Yufei Ma 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | 3D-SubG: A 3D Stacked Hybrid Processing Near/In-Memory Accelerator for Subgraph GNNsabstractSubgraph Graph Neural Networks (GNNs) are emerging as a promising approach to enhance GNN expressiveness, but their more complex graph structures with numerous independent and irregular subgraphs pose significant hardware deployment challenges. In this work, we propose 3D-SubG, a 3D stacked hybrid processing-near/in-memory accelerator for subgraph GNNs. With hybrid bonding packaging technology, a logic die is 3D stacked with a DRAM die for highly parallel memory accesses. The logic die employs digital SRAM-based processing-in-memory (PIM) macros to boost computation density and minimize data transfer. We further propose a bit-level non-zero gathering method to exploit graph sparsity for PIM, a workloadbalanced mapping strategy for subgraph allocation onto different logic-to-DRAM blocks, and a distributed global pooling approach to reduce inter-block data movements. Experimental results show that 3D-SubG achieves average improvements of $146.11 \times$ in performance, $934.18 \times$ in area efficiency, and $1171.80 \times$ in energy efficiency compared to RTX 3090Ti. Guoxiang Li, Runnan Xu, Ruohang Xu, Yikan Qiu, Renati Tuerhong, Muhan Zhang, Le Ye, Yufei Ma 0002 |
DAC | 4 |
| 2024 | DCIM-GCN: Digital Computing-in-Memory Accelerator for Graph Convolutional NetworkabstractGraph convolutional network (GCN) has gained great success in a diverse range of intelligent tasks. However, the hardware performance of GCNs is often bounded by random and non-continuous memory accesses due to the sparse graph data, which incur high latency and high power consumption. The emerging computing-in-memory (CIM) architecture significantly reduces the overhead of data movements, which is suitable for memory-intensive GCN acceleration. Existing analog-based CIM solutions require a large amount of analog-to-digital (AD) and digital-to-analog (DA) conversions, which dominate the overall area and power consumption. Furthermore, the analog non-ideality can degrade accuracy and reliability of CIM. To address these challenges, this work proposes a digital CIM accelerator based on SRAM, called DCIM-GCN, to accelerate GCN algorithm. DCIM-GCN introduces innovations on three levels: circuit, architecture, and algorithm. At the circuit level, digital CIM is proposed with SRAM sub-arrays to eliminate the power and area expensive AD/DA converters. Furthermore, we have incorporated the multi-address feature into the digital CIM, thereby leveraging its ability to efficiently process sparse matrix multiplication. At the architecture level, the sparsity-aware computation engine takes advantage of sparsity in GCNs and leverages CIM to minimize memory accesses and data movements. Finally, at the algorithm level, the balance mapping algorithm tackles workload imbalance issues, while the vertex reorder algorithm reduces idle states for aggregation engines, resulting in increased hardware utilization. Our DCIM-GCN achieves 1.89$\times$and 2.42$\times$speedup and 4.58$\times$and 9.46$\times$energy efficiency improvement on average over other CIM-based graph accelerators, e.g., PASGCN and PIM-GCN, respectively. Yufei Ma 0002, Yikan Qiu, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | DCIM-GCN: Digital Computing-in-Memory to Efficiently Accelerate Graph Convolutional NetworksabstractComputing-in-memory (CIM) is emerging as a promising architecture to accelerate graph convolutional networks (GCNs) normally bounded by redundant and irregular memory transactions. Current analog based CIM requires frequent analog and digital conversions (AD/DA) that dominate the overall area and power consumption. Furthermore, the analog non-ideality degrades the accuracy and reliability of CIM. In this work, an SRAM based digital CIM system is proposed to accelerate memory intensive GCNs, namely DCIM-GCN, which covers innovations from CIM circuit level eliminating costly AD/DA converters to architecture level addressing irregularity and sparsity of graph data. DCIM-GCN achieves 2.07X, 1.76X, and 1.89× speedup and 29.98×, 1.29×, and 3.73× energy efficiency improvement on average over CIM based PIMGCN, TARe, and PIM-GCN, respectively. Yikan Qiu, Yufei Ma 0002, Meng Wu 0005, Le Ye, Ru Huang 0001 |
ICCAD | 1 |