EDBT 2026 Demo / reviewers in the wild / expert
Quan Deng 0003
dblp:09/11365-3
· DBLP profile ↗
11ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-1454-8486ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vector Value Prediction with Element-wise Stride CompressionabstractThe increasing emphasis on vectorization and Single Instruction, Multiple Data (SIMD) processing reflects their central role in modern processors. However, as workloads in data processing, multimedia, and algorithmic operations grow in complexity, they introduce more pronounced data dependencies, leading to longer execution times compared to scalar instructions. To address these evolving challenges, we present the Vector Value TAGE predictor (VVTAGE), a novel value predictor specifically designed for vector instructions. Although value prediction has been proposed as a fundamental strategy to enhance processor performance, it has traditionally focused on predicting 64-bit scalar values to mitigate data dependencies and improve pipeline throughput. VVTAGE extends the prediction capabilities of existing scalar predictors to accommodate the wide vector registers used in contemporary processors. Our research demonstrates that VVTAGE can significantly improve processor performance, achieving performance gains of up to 20.1% and an average increase of 4.53% in the evaluated SIMD benchmarks. This innovative approach and surprising results represent a significant advancement in optimizing the performance of SIMD processors. Furthermore, to enhance the scalability of VVTAGE, we propose an element-wise stride compression method to reduce its storage overhead. Experimental results show that VVTAGE still retains 64% performance gain while reducing 15.4KB overhead. Yanmeng Huang, Ling Yang 0008, Yuanhu Cheng, Quan Deng 0003, Junbo Tie, Yongwen Wang, Hai Zhong, Libo Huang 0002 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | PFP: Parallel Floating-Point Vector Multiplication Acceleration in MAGIC ReRAMabstractEmerging applications, e.g., machine learning, large language models (LLMs), and graphic processing, are rapidly developing and are both compute-intensive and memory-intensive. Computing in Memory (CIM) is a promising architecture that accelerates these applications by eliminating the data movement between memory and processing units. Memristor-aided logic (MAGIC) CIM achieves massive parallelism, flexible computing, and non-volatility. However, MAGIC ReRAM performs floating-point (FP) vector multiplication sequentially, which wastes parallel computing resources and is limited by the array size. To solve this issue, we propose a parallel floating-point vector multiplication accelerator in MAGIC ReRAM. We exploit three levels of parallelism during the calculation of FP vector multiplication, referred to as PFP. First, we leverage the parallelism of MAGIC ReRAM. Second, we bring forward the final exponent to make the exponent calculations parallel. Third, we decouple the calculation of exponent, mantissa, and sign, which allows parallel calculation across accumulation. The experimental results show that PFP achieves a performance speedup of 2.51× and 15% energy savings compared to AritPIM when performing FP32 vector multiplication with a vector length of 512. Quan Deng 0003 |
DATE | 3 |
| 2024 | ImSPU: Implicit Sharing of Computation Resources Between Vector and Scalar Processing Units
Hongbing Tan, Guichu Sun, Liquan Xiao, Yuanhu Cheng, Quan Deng 0003, Bingcai Sui, Yongwen Wang, Libo Huang 0002 |
Euro-Par (2) | 8 |
| 2024 | SSC: An SRAM-Based Silence Computing Design for On-chip Memory
Quan Deng 0003, Yiyue Hu, Libo Huang 0002, Yongwen Wang |
ICA3PP (4) | 2 |
| 2023 | Fast Approximate LUT-based Vector Multiplication in DRAMabstractVector multiplication is widely used in real-world applications. To accelerate vector multiplication, processing-in-memory-based domain-specific architectures leverage lookup tables (LUTs) to decrease the computational complexity of multi-plication. However, the overhead of LUTs increases exponentially with the result space, which throttles the system performance.To decouple the LUT size and the performance improvement, we make a trade-off between accuracy and performance. We propose fast approximate LUT-based vector multiplication in DRAM, which builds the partial LUT of higher-bit locations and reuses the LUT to calculate the lower-bit data. We propose an operand reorganizing optimization to group more zero, which does not affect the results. To reduce the LUT and pre-calculation cost further, we propose a value encoding. Our experiment shows that the performance of the proposed design can be improved by 4x and 1.25x compared with LAcc in AlexNet and MobileNetV2 without any accuracy loss, respectively. Besides, the performance of the proposed design improves by up to 3.74x compared with pLUTo in multi-vector multiplication. The area and power of the proposed design are 54.8mm2and 5.35W, respectively. Quan Deng 0003, Yongwen Wang |
ICPADS | 2 |
| 2020 | FRF: Toward Warp-Scheduler Friendly STT-RAM/SRAM Fine-Grained Hybrid GPGPU Register File DesignabstractModern graphics processing units (GPUs) exhibit increasing demands for register files (RFs) with larger capacity and bank sizes, which jeopardize the traditional SRAM-based RF designs due to their large die area and long access latency. Recent hybrid RF designs, e.g., SRAM and spin-transfer torque random access memory (STT-RAM)-based RFs, mitigate the issue by exploiting the density and performance advantages in STT-RAM and SRAM, respectively. However, existing hybrid RF designs adopt coarse integration that has limited write bandwidth between SRAM and STT-RAM, which restricts the adoption of different warp schedulers at runtime. In this article, we propose FRF, a warp-scheduler friendly fine-grained hybrid RF design using SRAM/STT-RAM hybrid cell (HC) structures. By integrating one SRAM cell and N STT-RAM cells as one HC, FRF exploits internal write paths to enlarge the access bandwidth between SRAM and STT-RAM and thus greatly optimizes the area and performance. FRF enables the concurrent context-switching such that different warp schedulers may be adopted at runtime. FRF adopts interleaved register mapping (IRM) and on-demand register remapping to further improve the utilization of SRAM in each HC. Our experimental results show that, on average, FRF achieves 50% performance improvement and 40% energy consumption reduction over the coarse-grained hybrid design when adopting loose round-robin (LRR), and achieves 159% efficiency improvement over pure STT-RAM-based RF. Quan Deng 0003, Youtao Zhang, Shuzheng Zhang, Minxuan Zhang, Jun Yang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | LAcc: Exploiting Lookup Table-based Fast and Accurate Vector Multiplication in DRAM-based CNN AcceleratorabstractPIM (Processing-in-memory)-based CNN (Convolutional neural network) accelerators leverage the characteristics of basic memory cells to enable simple logic and arithmetic operations so that the bandwidth constraint can be effectively alleviated. However, it remains a major challenge to support multiplication operations efficiently on PIM accelerators, in particular, DRAM-based PIM accelerators. This has prevented PIM-based accelerators from being immediately adopted for accurate CNN inference. Quan Deng 0003, Youtao Zhang, Minxuan Zhang, Jun Yang 0002 |
DAC | 1 |
| 2019 | RFAcc: a 3D ReRAM associative array based random forest acceleratorabstractRandom forest (RF) is a widely adopted machine learning method for solving classification and regression problems. Training a random forest demands a large number of relational comparison and data movement operations, which take long time when using modern CPUs. Accelerating random forest training using either GPUs or FPGAs achieves only modest speedups. Quan Deng 0003, Youtao Zhang, Jun Yang 0002 |
ICS | 2 |
| 2019 | DWMAcc: Accelerating Shift-based CNNs with Domain Wall MemoriesabstractPIM (processing-in-memory) based hardware accelerators have shown great potentials in addressing the computation and memory access intensity of modern CNNs (convolutional neural networks). While adopting NVM (non-volatile memory) helps to further mitigate the storage and energy consumption overhead, adopting quantization, e.g., shift-based quantization, helps to tradeoff the computation overhead and the accuracy loss, integrating both NVM and quantization in hardware accelerators leads to sub-optimal acceleration. In this paper, we exploit the natural shift property of DWM (domain wall memory) to devise DWMAcc, a DWM-based accelerator with asymmetrical storage of weight and input data, to speed up the inference phase of shift-based CNNs. DWMAcc supports flexible shift operations to enable fast processing with low performance and area overhead. We then optimize it with zero-sharing , input-reuse , and weight-share schemes. Our experimental results show that, on average, DWMAcc achieves 16.6× performance improvement and 85.6× energy consumption reduction over a state-of-the-art SRAM based design. Zhengguo Chen, Quan Deng 0003, Nong Xiao 0001, Kirk Pruhs, Youtao Zhang |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | DrAcc: a DRAM based accelerator for accurate CNN inferenceabstractModern Convolutional Neural Networks (CNNs) are computation and memory intensive. Thus it is crucial to develop hardware accelerators to achieve high performance as well as power/energy-efficiency on resource limited embedded systems. DRAM-based CNN accelerators exhibit great potentials but face inference accuracy and area overhead challenges. Quan Deng 0003, Lei Jiang 0001, Youtao Zhang, Minxuan Zhang, Jun Yang 0002 |
DAC | 1 |
| 2017 | Towards warp-scheduler friendly STT-RAM/SRAM hybrid GPGPU register file designabstractModern Graphics Processing Units (GPUs) widely adopt large SRAM based register file (RF) to enable fast context-switch. A large SRAM RF may consume 20% to 40% GPU power, which has become one of the major design challenges for GPUs. Recent studies mitigate the issue through hybrid RF designs that architect a large STT-RAM (Spin Transfer Torque Magnetic memory) RF and a small SRAM buffer. However, the long STT-RAM write latency throttles the data exchange between STT-RAM and SRAM, which deprecates warp scheduler with frequent context switches, e.g., round robin scheduler. In this paper, we propose HC-RF, a warp-scheduler friendly hybrid RF design using novel SRAM/STT-RAM hybrid cell (HC) structure. HC-RF exploits cell level integration to improve the effective bandwidth between STT-RAM and SRAM. By enabling silent data transfer from SRAM to STT-RAM without blocking RF banks, HC-RF supports concurrent context-switching and decouples its dependency on warp scheduler. Our experimental results show that, on average, HC-RF achieves 50% performance improvement and 44% energy consumption reduction over the coarse-grained hybrid design when adopting LRR(Loose Round Robin) warp scheduler. Quan Deng 0003, Youtao Zhang, Minxuan Zhang, Jun Yang 0002 |
ICCAD | 1 |